vLLM 三周优化 DeepSeek-V4.1-Flash:低并发提速 1.9 倍
vLLM 花三周把 DeepSeek-V4.1-Flash 跑快了 1.9 倍,AgentX 上吞吐 5.3 倍,还附交互图表拆解过程,做部署的可以看看。
vLLM 团队从零开始用了三周时间完成对 DeepSeek-V4.1-Flash 的推理优化。在低并发场景下,推理速度提升 1.9 倍。在 SemiAnalysis_ 的 AgentX 基准上,单个用户达到 150 TPS 时吞吐量达到 5.3 倍。推文附带了可逐步查看的交互式图表,展示具体优化手段。
1/ Three weeks from day 0, DeepSeek-V4.1-Flash on vLLM runs 1.9× faster at low concurrency and delivers 5.3× the throughput at 150 TPS per user on @SemiAnalysis_ AgentX.
Here is how, with interactive figures you can step through 🧵 https://t.co/tytotfLOts