论文

DuplexCadence:让全双工语音模型运行提速 2.85 倍、峰值内存降低 38.8%

DuplexCadence: Exact State and Execution from a Speech Model's Declared Timelines

精选理由

跑全双工语音模型的开发者可以看,它把推理速度提到原来的 2.85 倍还省内存,输出结果完全一致。

全双工语音模型按一秒节拍流式推进对话,token 转音频的合成尾段常因 GPU 等待海量小操作而超时。DuplexCadence 把模型原生的区域时钟计数器显式声明给运行时,据此实现按需状态分配与精确形状的图回放。在四个已发布模型、三种解码器架构上测试,输出逐位一致,速度达原运行时的 2.85 倍,峰值内存降低 38.8%。实时双工路径上,SPEAK 平均耗时从超出 1 秒节拍的 14% 降到节拍内的 2%。

原文 · arXiv cs.AI

DuplexCadence: Exact State and Execution from a Speech Model's Declared Timelines

Full-duplex speech models support streaming interaction that listens and speaks at the same time. Serving them is governed by a strict, repeating deadline: conversation advances on a one-second cadence, and every second of input must be turned into a second of speech before the next second arrives. Because stages within a session run in strict sequence, per-invocation overhead cannot be batched away. Profiling reveals that the autoregressive stages of a duplex second already fit within the period, whereas the token-to-audio synthesis tail is what causes overruns. This tail stage suffers from orchestration slack where the GPU is left waiting as thousands of tiny, regular operations are issued one by one, while also wasting substantial memory by over-provisioning state at static implementation constants. Existing remedies, such as graph recording and demand-sized allocation, fail because streaming state dynamics violate their prerequisites. The root cause is that the runtime lacks the model's native clocks: the per-region counters that govern advancement rates and retention policies. We propose DuplexCadence, which explicitly declares native clocks to the runtime and derives two mutually enabling rules: demand-sized state allocation at a stable address, and exact-shape graph replay without padding. The former eliminates idle memory and stabilizes tensor pointers, while the latter removes orchestration slack without padding overhead. Evaluated on four released models across three decoder architectures with bit-for-bit identical output, DuplexCadence reaches $2.85\times$ the stock runtime's speed at $38.8\%$ lower peak memory. On the live duplex path, mean SPEAK time falls from $14\%$ over the one-second cadence to $2\%$ under it, enabling models to reliably keep up with interactive speech while markedly expanding multi-