Back to Home

Latest progress in AI video generation technology: from Sora to Ke Ling

June 1, 2026 at 08:04 AMSource: RunByAI0 comment(s)TechReview

In 2024, OpenAI's Sora model brought AI video generation from the laboratory to the public eye. A year has passed, and this technology has not been a flash in the pan, but has evolved at an astonishing pace, forming a new track of global competition. Chinese science and technology enterprises also quickly caught up, with Kwai Kering, ByteDance is a dream, Alibaba Tongyi Wan and other products coming on the scene, jointly promoting the leapfrog development of AI video generation technology.

Sora's successor versions have achieved significant improvements in multiple dimensions. The most impressive improvement is its consistency. Although the videos generated by the first generation Sora have stunning visual effects, there are still significant shortcomings in object tracking and spatiotemporal coherence - objects may appear and disappear in the frame, and facial features of characters will change with camera switching. The new generation model has significantly improved these shortcomings by introducing a 3D spatiotemporal attention mechanism, resulting in highly consistent videos in one minute long scenes.

Kwai Keling has gone out of a differentiation route. While maintaining video quality, Ke Ling has focused on optimizing its ability to understand Chinese scenes. Whether it's the streetscape of Chinese cities, traditional cultural elements, or Chinese character recognition, Ke Ling's performance is superior to international competitors. Ke Ling was also the first to launch a photo generated video function, where users only need to upload one image and AI can generate dynamic videos based on its composition and style, greatly reducing the threshold for creation.

The ByteDance Dreamina is particularly outstanding on the mobile end. Thanks to ByteDance's profound accumulation in the field of short videos, Dream has a unique grasp of video dynamics, and the generated dance and sports scenes are smooth and natural, especially suitable for social media content creation.

At the technical principle level, AI video generation is undergoing a transition from diffusion models to diffusion Transformer hybrid architectures. This new architecture combines the image generation capabilities of diffusion models with the sequence modeling advantages of Transformers, enabling it to handle longer spatiotemporal sequences. In addition, the resolution of video generation is also moving from 720p to 4K, and the frame rate is gradually approaching the movie standard of 24fps.

It can be foreseen that AI video generation will create enormous value in fields such as film and television production, advertising creativity, education and training, and social entertainment. In the coming year, real-time generation and high-definition long videos will become the focus of technological breakthroughs.

video generationSoraAI videoClever
Discussion

Comments (0)

No comments yet. Be the first!

Leave a Comment