vLLM将在PyTorchCon展示LLM推理引擎更新
vLLM团队将在PyTorchCon分享LLM推理引擎的最新优化,包括KV缓存管理和GPU内核改进。
vLLM将在PyTorchCon北美大会上展示其高吞吐量、内存高效的LLM推理和 serving引擎。Simon Mo将发表主题演讲,介绍vLLM核心架构改进和KV缓存管理优化。技术会议将探讨注意力机制、KV缓存传输、分布式服务等内容。
vLLM (@vllm_project) is a high-throughput, memory-efficient inference and serving engine for LLMs. At #PyTorchCon North America, you’ll find vLLM-related work throughout the program, from Simon Mo’s keynote to technical sessions and posters covering LLM inference and serving.
@simon_mo_ will present the keynote “vLLM Update: Scaling Open Frontier Inference Infrastructure,” covering improvements to vLLM’s core architecture and major optimizations in KV cache management and GPU kernels.
Across #PyTorchCon, technical sessions dig into vLLM-related inference, including attention, KV cache management and transfer, disaggregated serving, elastic expert parallelism, scaling across hardware without forks, and more.
"Meet the Developers of vLLM" will feature George Novack and Nick Hill. Posters cover additional vLLM-related work on custom accelerators, Trainium, expert parallelism and RDMA KV-cache transfer, multi-chip KV-cache transfer, tier-aware routing, and more.
Explore the vLLM sessions: https://t.co/kvOrhsTPkm
PyTorch conferences are the open source AI community’s town square, where what’s next gets decided. Register for PyTorch Conference North America, October 20–21 in San Jose: https://t.co/cQzFxEziBX