PyTorch推出原生TPU后端引擎vLLM和SGLang
Google和Meta工程师将分享如何让vLLM和SGLang在TPU上原生运行,保留GPU用户熟悉的API和工作流。
Google和Meta工程师将在PyTorch Conference North America介绍基于TorchTPU的新vLLM和SGLang引擎。这些引擎使现有基础设施能在TPU上原生运行,保留GPU用户熟悉的调度器、批处理系统和OpenAI兼容API。新模型和功能可在TPU上运行,无需单独的TPU实现,生产环境运行经验包括编译时间、设备放置和KV缓存布局。
Want to hear more about new vLLM and SGLang engines built on a new PyTorch-native TPU backend, TorchTPU?
Find Qi Zhou of @Google, Colin Taylor and Angela Yi of @Meta at PyTorch Conference North America in San Jose where they will discuss how they focused on making the existing @sgl_project & @vllm_project infrastructure work natively on TPU, preserving the schedulers, batching systems, OpenAI-compatible APIs, and torch.compile workflows already familiar to GPU users.
Because the serving engines remain upstream, new models and features can run on TPU without requiring separate TPU-specific implementations. Learn more about how it works across the stacks & what they learned running this in production, including compile time, device placement under tracing, KV cache layout, and cross-chip collectives.
Register for #PyTorchCon NA now: https://t.co/jBApW8ocHQ