论文

WUSH-KV: KV缓存量化新方法

WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms

精选理由

WUSH-KV用数据自适应变换降低KV缓存量化误差,2-bit性能超越OSCAR,适合长上下文推理场景。

WUSH-KV是一种低比特KV缓存量化技术,通过校准数据构建独立的键值变换。该方法在2-bit量化下,在SGLang框架中实现了与OSCAR变换相当或更优的性能。WUSH-KV在层级重建误差和端到端困惑度指标上均优于其他测试变换。

原文 · arXiv cs.LG

WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms

KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which constructs a data-aware transform from the second-order statistics of both factors in a matrix product to reduce quantization error. WUSH-KV uses calibration data to construct separate key and value transforms, with the value transform folded into the model weights and the key transform applied after RoPE. The transforms can be paired with clipped quantizers. For one such quantizer, QuEST INT, we show that, under mild assumptions, the WUSH transform is near-optimal. With this quantizer, WUSH-KV reduces layerwise reconstruction error and achieves the lowest end-to-end perplexity among other tested transforms. For end-to-end evaluation, we integrate WUSH-KV into SGLang using OSCAR-style percentile-clipped affine quantization. At 2-bit, WUSH-KV performs comparably to or outperforms the OSCAR transform across all evaluated models and downstream tasks.