vLLM 核心维护者将于 10 月 19 日在 San Jose 谈推理经济账
vLLM 的核心维护者亲自讲推理服务的成本和延迟怎么权衡,做部署的人去听听能省不少踩坑时间。
vLLM 项目宣布,核心维护者 Tyler Michael Smith 将于 10 月 19 日在 San Jose 的活动上讲解大规模推理服务中的成本、延迟与容量权衡。该分享聚焦推理经济学问题,即在不同负载下如何平衡 GPU 开销与响应速度。活动由 Red Hat、NVIDIA 和 IBM 联合支持举办。
Serving AI at scale brings real tradeoffs in cost, latency, and capacity. Hear from vLLM core maintainer Tyler Michael Smith on inference economics in San Jose, Oct 19. Thanks @RedHat, @NVIDIAAI, and @IBM for bringing the community together!