vLLM-Omni技术报告:统一多模态生成运行时
vLLM-Omni Technical Report: A Unified Serving Runtime for Omni-Modality Generation
vLLM团队发布了统一多模态运行时vLLM-Omni,解决了不同模态模型各自为政的问题,让语音、视觉和机器人系统可以协同工作。
vLLM-Omni是一个统一的多模态生成运行时系统,专为处理文本、音频、图像、视频和动作等不同模态的模型而设计。该系统采用多阶段流水线架构,在H100和H200硬件上进行了测试,特别针对Qwen3-Omni模型进行了评估。vLLM-Omni提供了OpenAI兼容和OpenPI API,支持语音合成、图像/视频生成、世界模型、机器人交互和双向工作负载。
vLLM-Omni Technical Report: A Unified Serving Runtime for Omni-Modality Generation
Interaction with intelligent systems is expanding beyond text-centric chatbots and coding agents. Speech-native assistants, visual generation and editing, world-model environments, and robot action loops require models that emit text, audio, images, video, and actions. These models differ in execution pattern: multi-stage autoregressive omni and TTS pipelines, iterative diffusion or flow-matching generators, and longer-lived world-model or robot loops that carry state across steps. As a result, serving is no longer a single text decode loop, but a heterogeneous multi-stage workflow with cross-stage transfer, streaming, and session-shaped interaction. Existing inference stacks are typically optimized for one architecture family. LLM servers deepen autoregressive scheduling and KV management, while diffusion stacks deepen denoising and parallel generation. Neither provides a shared control plane for pipelines that emit speech, pixels, or actions through separate generators, so production deployments often fall back to ad-hoc composition across disjoint runtimes. We present vLLM-Omni, a unified serving runtime for omni-modality generation. vLLM-Omni organizes each workload as a multi-stage pipeline under a single orchestrator that admits requests, advances them across stages, and demultiplexes streaming outputs. Specialized engines and stage replicas provide compute; a connector carries heavy payloads on the data plane; and session-oriented control supports long-lived duplex, world-model, and robot workloads. This report covers the architecture (stage-level KV paths, replica pools, multi-hardware platforms, and efficiency stack) and OpenAI-compatible and OpenPI APIs for omni, TTS, image/video, world-model, robot, and duplex workloads. We evaluate on the multimodal nightly CI on H100 (TTS and MiniCPM-o on H200), focused on Qwen3-Omni.