论文精选

清华发布TokenRouter大模型服务系统

精选理由

清华团队新系统让大模型服务效率提升64倍,小大模型协同处理更高效。

清华大学研究团队提出TokenRouter服务系统,实现了令牌级别的小模型和大模型路由。该系统在5种路由方法测试中,吞吐量达到现有系统的2.01至64.15倍。TokenRouter为每个模型分配独立服务器,允许模型间传递工作并保留文本记忆缓存。

原文 · rohanpaul_ai

New Tsinghua paper builds TokenRouter, a serving system that runs per-token small-and-large model routing at upto 64.15X the throughput of existing setups.

Current popular serving frameworks (like vLLM and SGLang) run one model per request, so when two models share an answer, every step waits for the slower one.

TokenRouter gives each model its own server and lets them pass work back and forth. So, it hands a half-written answer between them while keeping the model’s memory of the text so far (the KV cache), and holds requests for a moment so each model works on bigger batches.

Across 5 routing methods, throughput rose 2.01 to 64.15 times over the stronger existing setup.

– arxiv. org/abs/2610.12242

Title: "TokenRouter: Efficient Serving System for Token-Level LLM Routing"