模型精选73°

vLLM-Omni 技术报告发布:统一多模态生成服务运行时

精选理由

vLLM 出了个统一运行时,把语音、图像扩散、世界模型、机器人回环这些异构请求塞进一套服务框架,做全模态应用的可以看看

vLLM 团队发布 vLLM-Omni 技术报告,定位为全模态生成的统一服务运行时。核心问题在于语音助手、视觉生成、世界模型和机器人回环的执行模式各不相同,多阶段自回归、迭代扩散和带状态的会话超出了单一文本解码循环的范畴。vLLM-Omni 用一个 orchestrator 推进请求跨阶段流转,专用引擎负责计算,connector 传递负载,同一会话路径支撑双工、世界模型和机器人回环。论文和代码仓库已公开,团队接受社区贡献。

原文 · vLLM

🚀 Excited to share the vLLM-Omni technical report: a unified serving runtime for omni-modality generation. 📄 Paper: https://t.co/b0KvTmrEF3 🔗 Repo: https://t.co/I1wL6iqO0r

Speech assistants, visual generation, world models, and robot loops have pushed serving past a single text decode loop. The execution patterns diverge with the output: multi-stage autoregressive pipelines, iterative diffusion, and sessions that carry state from step to step. LLM servers and diffusion stacks each go deep on only one of these, so deployments fall back to stitching disjoint runtimes together. vLLM-Omni is the shared control plane for that mix. An orchestrator advances each request across stages; specialized engines run the compute; a connector carries the payloads; and the same session path keeps duplex, world-model, and robot loops on one runtime.

Really grateful to the vLLM and vLLM-Omni teams for the collaboration and support. Contributions are welcome. 🙏

  • arXiv: OpenAI10-07 02:04原文