Video generation is the forefront of artificial intelligence's multimodal capabilities and one of the most rapidly iterating fields in the past two years. From OpenAI's Sora to domestic products such as Kelin, PixVerse, Vidu, etc., Wensheng Video Technology is rapidly moving from the laboratory to mass applications.
In early 2024, OpenAI's Sora model shocked the global AI community. Based on the Diffusion Transformer architecture, Sora is able to generate high-definition videos of up to 60 seconds based on text descriptions, achieving unprecedented levels of picture coherence, physical rule simulation, and camera motion. Although Sora has not yet been fully opened to the public, it has established a technological benchmark for cultural and creative videos - and video generation has entered the era of "second level creation".
The progress of domestic video generation models is also remarkable. The Kling model launched by Kwai supports the generation of 2K resolution videos, and performs well in the smoothness of characters' movements and the consistency of scenes. The Vidu model jointly released by Shengshu Technology and Tsinghua University adopts a unified Diffusion architecture to achieve end-to-end generation from text to video, achieving international leading levels in semantic understanding accuracy and picture detail richness.
From a technical perspective, current AI video generation can be divided into three main directions: frame by frame generation based on diffusion models, sequence modeling based on Transformers, and hybrid architectures that combine the two. The diffusion model has significant advantages in image quality, but there are challenges in terms of temporal consistency - flicker and jitter between adjacent frames are common issues. The Transformer approach has more advantages in long-term modeling and can better maintain the cross frame coherence of videos.
The latest progress worth paying attention to is the integration of multimodal video understanding and generation. Traditional video generation only focuses on the one-way conversion from "text to video", while the new generation of systems now supports richer interactive modes such as "image+text → video" and "video+text → editing". The video editing function launched by Pika Labs allows users to delineate specific areas in the video and input modification instructions, and AI automatically replaces and adapts the content of the areas.
In practical applications, AI video generation is penetrating from short video creation to professional film and television production. Film and television professionals have started using AI tools for rapid production of concept videos (Pre viz), auxiliary design of special effects scenes, and mass production of advertising creative materials. However, AI video generation still has a long way to go in terms of long video storytelling, character consistency, and emotional expression.
Another important trend is the rise of end-to-end video generation capabilities. With the significant improvement in NPU performance of chips such as Qualcomm Snapdragon 8 Gen 4 and Apple M4, some lightweight video generation models can now run on mobile devices, achieving real-time video effects and personalized video creation. This will further lower the threshold for video creation and allow more ordinary users to participate in AI video creation.
[Reference source] The content of this article is comprehensively collated from the official OpenAI blog, the Kwai Kering technical report and the publicly released technical information in the field of AI video generation.