Telescopic Language Model:一次训练在任意深度截断都能用
Telescopic Language Models
训练一次就能按算力预算随意截断的模型,论文里比固定出口方法省 43% 的曲线面积,做模型压缩的可以看看。
arXiv 论文提出 Telescopic Language Model(TLM),用随机前缀监督加全容量锚点训练嵌套容量 Transformer,每步只需两次前向-反向传播,不改架构、推理时零额外开销。在 200M 参数代理套件、20B FineWeb-Edu token 的实验中,单次 TLM 训练让全部二十个层级前缀都是有效语言模型。相比 MLMS 这类固定出口套件,TLM 将质量-预算曲线下面积降低 43-44%,满容量时表现持平,每次运行 GPU 成本约低 12%。论文还指出固定出口方案在未监督深度上困惑度会掉到 10^2-10^5。
Telescopic Language Models
One deployed language model must often serve many compute budgets, yet serving each budget still means a separate training or compression run per point. We train a Telescopic Language Model (TLM) to be that continuum: a nested-capacity Transformer supervised by stochastic prefix supervision with a full anchor. At every step, one randomly truncated prefix of the capacity axis is trained against the full next-token target, alongside one full-capacity pass, so the trained artifact is a valid language model at every depth. Two forward-backward passes per step, no architectural change, nothing extra at inference. Fixed-exit suites such as Matryoshka Language Model Suites (MLMS) occupy one point in this design space, and the point has a cost: supervising only a few fixed exits leaves the nested model at chance level everywhere else (perplexity 10^2-10^5 in our baselines). On a 200M proxy suite (20B FineWeb-Edu tokens, identical data stream for all methods), a single TLM run is a valid language model at every one of its twenty layer prefixes, in perplexity and on perplexity-sensitive downstream tasks, reducing the area under the quality-budget curve by 43-44% relative to the fixed-exit suites while matching them at full capacity, at ~12% lower GPU cost per run. The prefix sampling density is a dial: concentrating it on a few depths recovers fixed-exit quality there at the price of the continuum, so the operating points become a training-time choice rather than an architectural one. These results indicate that the training objective, not the nesting itself, is what makes a model elastic.