fal 开源视频模型 H3 Max 实现 35 倍加速,可生成 60 分钟连续可控视频
fal's Gorkem Yurtseven and Batuhan Taskaya on how faster-than-real-time generation unlocked continuo...
fal 的工程师用 H3 Max 在 Twitch 上直播生成连续视频,这比之前模型快很多,现在好莱坞公司也喜欢用它来快速调整镜头和场景。
fal 团队通过优化代码和提升 GPU 利用率,将开源视频模型 H3 Max 加速 35 倍,使其生成速度超过实时拍摄。该模型能连续生成长达 60 分钟的互动视频,用户可通过提示词实时控制场景和动作,且保持场景一致性。
fal's Gorkem Yurtseven and Batuhan Taskaya on how faster-than-real-time generation unlocked continuo...
fal's Gorkem Yurtseven and Batuhan Taskaya on how faster-than-real-time generation unlocked continuous, interactive AI video: "One of our engineers, Rehan, started streaming a live stream of continuous generations of H3 Max from his laptop. He was doing some prompt tricks, trying to keep a coherent story, and then he started livestreaming that on Twitch." "Our ML team was essentially trying to take every single video model and apply these optimizations and tricks... You never could generate five seconds under five seconds. Once H3 Max unlocked it, the ML team was like, 'This is insane.'" "You can essentially stream infinitely. We capped it at an hour... I think it's the only model that can generate up to 60 minutes of continuous video that is action-controlled." "You can start with a prompt, say there's an office setting and someone is working, and then 30 seconds later it just imagines by itself. 30 seconds later, you can say, 'A woman walks in through the door.' It can take the prompt and reflect it immediately, which is the most fun part." "And the office is still the same office. The camera can pan back to the original person, and the original person is still there in the same state." @gorkem @isidentical @fal Your browser does not support the video tag. 🔗 View on Twitter a16z @a16z . @fal 's Gorkem Yurtseven and Batuhan Taskaya on making an open source video model 35x faster, and what Hollywood wanted after they built it: Last month, MiniMax released H3, an open source video model. fal rebuilt it - they cut down the steps the model takes to make a video, rewrote the code under each stage, and got the GPUs to 70-80% of their theoretical ceiling instead of their usual 30-40%. No quality loss. Video now generates faster than you can film it. The models have gotten so cheap and fast that end users aren't even asking for improvements in either category anymore. The gap has moved to quality, or how closely the model follows the prompt. Hollywood wasn't a customer a year ago and is now fal's fastest growing segment. Studios love it for the little things - extending a shot, moving the camera, changing the lighting. Those edits land 80-90% of the time. fal is chasing 99.9%. In this conversation with a16z's Jennifer Li: 00:00 Intro 01:45 The first open model worth rebuilding 03:45 The industry ran out of compute in April 06:25 35x faster without losing quality 11:05 5s of video generated in 1.5s 14:25 The weekend fal dropped everything 16:15 3 viral projects nobody planned 18:20 Teaching a video model to remember 20:25 An hour of video that's consistent 23:30 2x the usage of other models in 3 weeks 26:35 The case for giving video away for free 29:00 Blender sketch in, finished shot out 30:40 Chasing 99.9% reliability 34:15 Hollywood, fal's fastest-growing customer 36:50 Why studios wouldn't touch it until now YouTube: youtu.be/SDbRJXQrYGY @gorkem @isidentical @fal @JenniferHli Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 3 🔄 1 ❤️ 4 👀 3897 📊 2 ⚡