AA-Video-T2V v2.0基准发布
新视频生成基准出炉,Wan 3.0夺冠,Seedance 2.5最贵但人体表现最佳,MiniMax H3性价比突出。
Artificial Analysis发布AA-Video-T2V v2.0视频生成基准,包含68,000个人类偏好投票。该基准在1080p分辨率下评估10种使用场景和10种能力,Wan 3.0排名第一,Dreamina Seedance 2.5在人体解剖和对话同步方面领先。基准还包含AA-Video-T2V-Silent v2.0版本,专注于无音频视频生成。
We are launching AA-Video-T2V v2.0, our new benchmark for evaluating text to video models, alongside AA-Video-T2V-Silent v2.0 for video generation without audio. Built on a new methodology, it judges every model at 1080p on a regularly refreshed prompt set, and ranks them across 10 use cases, 10 capabilities and a wide range of styles.
Video models are being adopted across more industries and workflows, from film studios to advertising agencies. Our new benchmark not only ranks models overall, but also shows which model is best for specific use case and video generation capability. Use cases are grounded in how consumers and enterprises use video generation. Capabilities draw on lab and academic research, and on how creators and businesses push video models today. We tag every prompt by use case (such as Live-Action Film and Marketing & Advertising) and by the capability it tests (such as Text Rendering and Audio Synchronization), and the overall benchmark samples evenly across both. Because of this, the overall ranking reflects a model's versatility across use cases and well-roundedness across capabilities. We also tag each prompt by visual style, such as photorealistic, 3D render, cartoon and anime, and hand-drawn illustration.
AI video is also moving onto bigger screens and into production, from microdramas to movie theaters, while low barrier to generate is resulting in a proliferation of low quality AI video content. The quality bar keeps rising, so we now judge every clip at 1080p and high bitrate.
We are launching AA-Video-T2V v2.0 with more than 68,000 high quality human preference votes from private evaluators based in US/UK over 1,000 prompts, and AA-Video-T2V-Silent v2.0 with more than 47,000 votes over 500 prompts.
Initial insights from an in-depth analysis of the 10 highest ranking models on the Artificial Analysis AA-Video-T2V v2.0 Leaderboard:
➤ Wan 3.0 ranks #1 overall and leads 10 of the 20 category boards, including Cartoon and Anime style and Animation & Gaming use case, at $12 per minute of video.
➤ Dreamina Seedance 2.5 ranks #2 and is the human performance specialist, #1 on both Human Anatomy and Dialogue & Lip Sync. At $34.12/min it is the most expensive model in the top 10.
➤ MiniMax H3 (768p) ranks #3, statistically tied with Seedance 2.5 at $4.80/min, about 1/7 of the price. It is also #1 on Text Rendering.
➤ FLUX 3 ranks #4, with its strongest results on Text Rendering (#3) and Dialogue & Lip Sync (#2).
➤ Gemini Omni Flash 1.1 ranks #5 and is the graphic 2D and audio specialist, #1 on UI/UX & Motion Design use case, Flat Design style and Audio Synchronization capability.
See below for the use case, capability and style breakdowns 🧵