WorldSonus:给世界模型配上实时空间立体声
WorldSonus: Bringing Sound to Worlds
给世界模型生成的视频实时配音,延迟低到 RTF 0.41,还能中途改音效,声音还会随镜头位置变化,搞世界模型的可以看看。
arXiv 论文 WorldSonus 是一个面向世界模型的交互式视频转音频框架,解决三个问题:实时生成、中途指令控制、空间立体声对齐。它采用流式因果自回归扩散架构,以 0.41 的低实时因子(RTF)合成音频块。通过基于音频的描述管线和按块索引的提示调度,用户可以在生成过程中动态修改声音事件。团队用立体声和 ambisonic 数据做监督训练,实验显示其在开放域视频转音频基准上与双向双向模型持平或更优。
WorldSonus: Bringing Sound to Worlds
Recent advances in world models have enabled increasingly realistic visual synthesis. However, these generated environments remain largely silent. Bringing sound to world models poses three core challenges: real-time generation to keep pace with interactive video streams, interactive control to respond to mid-stream sound instructions, and spatially aligned stereo to reflect scene geometry and camera motion. To address these demands, we introduce WorldSonus, an interactive video-to-audio framework designed for real-time spatial sound synthesis in world models. For real-time generation, WorldSonus employs a streaming causal autoregressive diffusion architecture that synthesizes audio chunks at a low real-time factor (RTF) of 0.41. For interactive control, we incorporate an audio-centric captioning pipeline with chunk-indexed prompt scheduling, enabling dynamic manipulation of sound events during generation. For spatial alignment, we leverage high-quality stereo supervision curated from diverse stereo and ambisonic data. Extensive experiments demonstrate that while tailored for world models, WorldSonus generalizes effectively to open-domain video-to-audio benchmarks, matching or outperforming state-of-the-art bidirectional models in both acoustic quality and spatial alignment. Project page: https://noizai.github.io/WorldSonus/