论文

Softmax 重参数化方法降低输出头量化损失,Phi-4-mini KL 误差降至 0.256

Softmax Reparameterization for Output-Head Quantization

精选理由

量化小模型的人可以看看:一个一行搜索就能让输出头 W4 量化的 KL 误差降七成的方法,还附 10.8% 延迟实测。

一篇 arXiv 论文提出 softmax 重参数化,一种面向输出头的训练后量化方法,核心是在量化前从每个输出行减去词表行均值的标量倍数,并通过验证 KL 在 RTN、AW-MSE 和 GPTQ 三种量化器下分别搜索系数。在七个输出头的 W4 实验中,Phi-4-mini 的 AW-MSE KL 误差从 0.936 降到 0.256。在 WikiText 上选出的冻结系数迁移到 C4 和 OpenWebMath 时,在 24 个对比中 18 个胜过均值中心化。W2 压力测试下收益覆盖几乎所有模型-量化器组合,且对 Phi 输出头量化可将 batch-one 生成延迟降低 10.8%。

原文 · arXiv cs.LG

Softmax Reparameterization for Output-Head Quantization

Large vocabularies make output heads a substantial inference cost in small language models. We propose softmax reparameterization, a post-training method that selects a functionally equivalent output head before quantization. The method subtracts a scalar multiple of the vocabulary-row mean from every output row and selects the coefficient by validation KL separately for RTN, activation-weighted MSE, and full-Hessian GPTQ. This one-dimensional search includes the original head and fixed mean-centering, preserves the full-precision softmax distribution, and leaves the trained decoder unchanged; a rank-one correction handles nonlinear logit paths such as soft-capping. Across seven heads, W4 gains concentrate where baseline quantization substantially distorts predictions: on Phi-4-mini, AW-MSE KL falls from 0.936 to 0.256. The gains survive stronger GPTQ calibration and remain complementary to exact per-channel scaling and affine quantization. Across four heads and three W4 quantizers, frozen WikiText-selected coefficients also transfer to C4 and OpenWebMath, outperforming mean-centering in all 18 comparisons where the frozen coefficient differs from $1$ and matching it in the remaining six. At W2, used as a compression stress test, benefits broaden across nearly the full model--quantizer matrix. Matched residual analysis shows that improved fidelity can accompany greater logit reconstruction error while reducing the residual's Fisher-weighted cost. For shift-compatible heads, reparameterization adds no inference operation and preserves packed W4 execution: with the decoder held in BF16, quantizing the Phi output head reduces batch-one generation latency by 10.8% relative to the BF16-head baseline.