论文精选

Looped Transformer 三种复用方案对比:Nanbeige 4.2 实测解析

精选理由

三条推文都叫 looped transformer 但完全是三种方案,选错能烧掉一周训练,这篇用 Nanbeige 4.2 实测数据帮你分清。

Looped transformer 一词在过去一个月被用来指代三种不同的权重复用方案。文章拆解了开源模型 Nanbeige 4.2 的做法:22 层网络循环两遍,同一组权重被应用 44 次,块计算量约为标准 Transformer 的两倍,KV cache 也随之翻倍。该设计保留了标准 Transformer 约 75% 的 token 效率;若共享 cache 可将显存减半,但质量增益持续偏低。另一种方案让状态在 token 间传递,使序列长度充当深度,其训练仅用 5 亿 token,却耗费约 20 倍的 GPU 时。

原文 · AlphaSignal

You’ve probably heard the buzz around Looped transformers, but which one?

Over the past month, the same word covered three reuse tricks.

Implementing the wrong one burns a training week. It also budgets a 3B file like a 3B decode.

We opened a looped transformer open model, Nanbeige 4.2:

> 22 layers, two passes > 44 applications, same weights > About twice the block compute > KV cache doubles with it

It kept about 75% of a standard Transformer's token efficiency.

Sharing the cache halved its memory. And quality gains were consistently lower.

A second design carries state token to token. Sequence length is the depth.

Training used 500 million tokens. That run cost about 20 times the GPU-hours.

So, Which reuse trick is the viral tweet? If extra loops do not raise quality, what are you paying? Does the file still need a second cache?

Full Breakdown ↓↓