Back to Home

Video Generation Enters the DiT Era: Architecture Evolution from Sora to Open Source Models

August 24, 2026 at 03:11 PMSource: RunByAI0 comment(s)TechView

In early 2024, OpenAI released Sora, demonstrating its ability to generate realistic minute level videos and bringing a technical term into the public eye - DiT (Diffusion Transformer). Previous video generation models were mostly based on the U-Net architecture, but Sora proved that combining diffusion processes with Transformers can achieve significant performance improvements in video generation tasks.

The core idea of DiT is not complicated: to divide an image or video into patches (blocks), which are processed by a Transformer like text tokens, and then gradually denoised using a diffusion model to generate content. Compared to U-Net, Transformer architecture has stronger scalability - there is a smoother scaling law relationship between parameter quantity, data quantity, and generation quality. The DiT paper (Scalable Diffusion Models with Transformers, published in ICCV 2023) systematically validated this trend and became the foundational architecture for many subsequent video generation models.

Sora's technical report points out that it adopts the route of "video compression network+DiT": first, the encoder compresses the video into latent space, then models the spatiotemporal patch with DiT in latent space, and finally decodes and restores it. This design allows the model to handle videos of different resolutions and aspect ratios, and naturally supports operations such as "extending videos" and "splicing videos". The report also emphasizes that the training data comes from publicly available and authorized video materials, which has become a focus of industry discussion.

After Sora was released, the open source community quickly followed suit. Domestic and foreign research teams have successively open-source multiple DiT based video generation models, such as Open Sora, which significantly reduces the training and inference threshold of DiT routes. These open-source projects not only reproduce basic capabilities, but also continue to explore directions such as controllable generation, camera motion, and audio synchronization. At the same time, DiT has also begun to have a reverse impact on image generation - the new generation of image models generally adopt the DiT architecture, which is superior to early U-Net schemes in terms of detail representation and semantic consistency.

From an industry perspective, DiT has shifted video generation from "laboratory demonstrations" to "engineering deployment". The inference cost remains the main bottleneck - video generation requires many orders of magnitude more computation than image generation, and technologies such as KV caching, parallel sampling, and model distillation are becoming new research hotspots. It can be foreseen that in the next year, video generation models will compete fiercely around efficiency, controllability, and multimodal fusion on the basis of architectural convergence.

【 Reference sources 】 OpenAI Sora Technical Report (Video Generation Models as World Simulators), DiT Paper (Scalable Diffusion Models with Transformers, ICCV 2023), and project documents publicly released by the open source community.

AI videoDiTvideo generation
Discussion

Comments (0)

No comments yet. Be the first!

Leave a Comment