Atlas用三台iPhone重现《黑客帝国》子弹时间
World Labs co-founder Justin Johnson says Atlas can recreate the famous "Bullet Time" shot from The ...
World Labs的Atlas模型用三台iPhone就能完成原来需要数百台相机才能完成的子弹时间效果,大幅降低3D空间捕获成本。
World Labs联合创始人Justin Johnson表示,Atlas模型能重现《黑客帝国》中著名的"子弹时间"镜头。原拍摄需要数百台相机和绿幕,而Atlas仅需三台iPhone。Atlas通过新视角预测技术,将数字空间3D表示的捕获需求减少了50到100倍。
World Labs co-founder Justin Johnson says Atlas can recreate the famous "Bullet Time" shot from The ...
World Labs co-founder Justin Johnson says Atlas can recreate the famous "Bullet Time" shot from The Matrix with iPhones: "The way they did that shot is they had a ring of hundreds of cameras. [Neo] fell over in the studio, they had hundreds of cameras viewing that angle on a green screen, and then they used those hundreds and hundreds of cameras to make that famous shot in The Matrix." "Now with Atlas, we can do this with as few as three cameras. No studio capture, no green screen, no expensive calibration. We can literally stick three iPhones on tripods and use these to take video of something happening." "From those three iPhone videos, we can then reframe the shot and imagine, like, freeze time, have the camera fly in as the milk is splashing up, and get these amazing frozen-time views. We can do this with just a couple cameras." @jcjohnss Your browser does not support the video tag. 🔗 View on Twitter a16z @a16z World Labs co-founders Fei-Fei Li, Justin Johnson, Ben Mildenhall, and a16z's Martin Casado on Atlas, a world model for spatial intelligence: LLMs are built on next token prediction. Video models are built on next frame prediction. Atlas is built on new view prediction, and it's the first model to unify pixel generation and pixel reconstruction, two problems computer vision has kept in separate tracks for half a century. The practical result is a 50 to 100x reduction in what it takes to digitally capture a 3D representation of a space. Previously, you needed 100 to 300 photos of a single room. Atlas can work from just three. In this conversation, they get into the slow motion shot from The Matrix that took hundreds of cameras and now takes three iPhones, the overnight Slack message that made them bet the company in five seconds, why robotics is bottlenecked on data rather than chips, and the case that new view prediction is AI-complete. 00:00 Intro 01:50 The Matrix slow motion scene now takes three iPhones 02:48 Why new view prediction is the primitive 07:10 Unifying generation and reconstruction 11:15 Gaussian splats became the bottleneck 14:17 Dense capture used to mean 300 photos 17:30 Why reconstruction needs generation to fill the gaps 18:44 The LLM lesson image models missed 23:39 The video that made them go all in 28:04 3D design is 95% revisions 30:50 The problem in robotics is data, not chips 32:48 Why a robot policy can't be trained like an image model 34:44 When the simulator becomes the planner 36:45 Frozen time required footage full of movement 40:57 Why new view prediction is AI-complete 42:43 Nature gave animals eyes but not trees YouTube: youtube.com/watch?v=qn1QDD… @drfeifei @jcjohnss @BenMildenhall @theworldlabs @martin_casado Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 2 🔄 2 ❤️ 22 👀 7633 📊 4 ⚡