低差异抖动优化量化循环状态缓存
Low-Discrepancy Dither for Quantized Recurrent State Caches
论文提出了一种确定性抖动方法,能显著提升量化语言模型在长文本生成中的性能,无需额外计算成本。
研究人员提出了一种基于黄金比例Weyl抖动的确定性舍入方法,用于Mamba风格和混合语言模型的量化循环状态缓存。该方法在纯模型和混合模型、不同存储格式以及长解码场景下表现优于随机舍入,且无需随机数生成。研究还发现舍入到最近值的方法在短评估中看似最佳,但在长文本生成中误差会持续增长。
Low-Discrepancy Dither for Quantized Recurrent State Caches
Mamba-style and hybrid language models compress their past into a fixed-size recurrent state that is rewritten at every generated token. Storing this state in low precision saves memory bandwidth, but every rounding error is fed back into the next update and can accumulate over long generations. Production systems round the state stochastically; we ask which rounding rule such caches should use. We find that a deterministic golden-ratio Weyl dither, which needs no random numbers, consistently brings the quantized model closer to the full-precision one than stochastic rounding, across pure and hybrid models, storage formats, and long decoding horizons, at no extra cost. Round-to-nearest behaves differently: because it discards small updates, its error keeps growing, so it can look best in short evaluations yet falls far behind over long generations. A discrepancy analysis explains this ordering, and we document implementation pitfalls that silently remove the benefit.