模型精选

vLLM 长上下文基准:双 RTX 5090 跑出 110 tok/s

精选理由

vLLM 官方转发的长上下文基准,双张 RTX 5090 就能跑到 110 tok/s,想自己搭推理服务的话很有参考价值。

NicW_AI 公布了一份针对 vLLM 的长上下文推理基准测试。结果显示双卡 RTX 5090 配置下可达到 110 tok/s 的生成速度。vLLM 团队转发了该结果,并指出长上下文场景主要压测显存与执行路径,服务引擎架构对性能的影响不亚于硬件本身。

原文 · vLLM

Appreciate the benchmark from @NicW_AI. Long-context serving heavily stresses memory and execution paths—getting 110 tok/s on dual RTX 5090s shows how much the serving engine architecture matters alongside raw silicon.

https://t.co/LpDjhMM0fd