In February 2024, OpenAI released a Sora trailer, stunning the world with realistic videos generated from text, and pushing "Wensheng Video" to the forefront of AI competitions. More than two years have passed, and AI video generation has evolved from a few second clip of "picture and music" to an industrialized creative tool that can control camera movements, match characters, and synchronize sound and image.
1、 Pattern: Parallel development of domestic and international lines
On the overseas track, OpenAI's Sora will be officially opened to the public at the end of 2024, supporting text and image generated videos, as well as video stitching and looping; Google DeepMind's Veo series continues to iterate, emphasizing native audio and cinematic quality; Runway, Luma and other start-up companies are constantly increasing their efforts in stylization and fine control.
China is also unwilling to be outdone: Kwai can be flexible, ByteDance is a dream, bio digital technology Vidu, MiniMax Conch AI and other products are intensively released, and each has its own strengths in Chinese scene understanding, character consistency, controllable operation, etc. Wensheng Video has become one of the most fiercely competitive tracks for domestic large-scale models.
2、 Technical keywords: consistency, controllability, narrative
The biggest pain point of early Wensheng videos was "uncontrollability" - the generated results were like opening a blind box, with inconsistent character images and frequent drifting of image details. In the past two years, the industry's efforts have almost all revolved around three keywords:
Consistency: By using techniques such as reference images, character locking, and face preservation, the same character can "look the same" in different shots, which is a prerequisite for moving from "fragments" to "works".
Controllability: The control of camera language is becoming increasingly refined, from simple "zoom in" to multi shot script arrangement, allowing creators to plan storyboards like directors.
Narrative: Multi shot generation, storyboard driven, voice over and subtitle automatic generation, AI is moving from "generating images" to "generating stories". Short dramas, advertisements, and e-commerce marketing have become the first scenarios to benefit.
3、 Landing: Who is actually using it?
At present, the most mature landing scenarios for AI videos are concentrated in three categories: advertising and marketing, using AI to quickly generate multiple versions of creative materials, significantly reducing shooting costs; The second is short dramas and videos, with AI assisted generation of storyboard previews and special effects shots; Thirdly, in the early stage of film and television production, AI is used for concept rehearsals to help directors "see" the final film effect before actual filming.
For enterprises, the cost-effectiveness advantage of AI video is obvious: traditional TVC shooting costs are high, while AI generated materials can be iterated on demand, producing a version in a few minutes. For individual creators, adding a subscription to a computer can give them the production capabilities that used to require a team to complete.
4、 Calm reminder
Under the trend, we also need to be clear headed: AI videos can still make mistakes in complex actions, physical laws, and long shot logic; The issue of copyright and deepfakes is receiving increasing regulatory attention, and many countries have introduced or are considering labeling obligations for AI generated content. Creators should comply with platform rules and local laws and regulations when using AI video tools, and clearly label the generated content.
From a few seconds of stunning moments to a complete work, AI video is going through the path of "from toys to tools" that all new technologies will experience. It can be foreseen that in the next two to three years, "AI directors" will become the infrastructure of the content industry, and what is truly scarce are still creativity, aesthetics, and narrative abilities.
[Reference source] The content of this article is comprehensively summarized from the industry information published by OpenAI official blog, Google DeepMind official blog and Kwai Technology.