Looped Transformer 三种复用方案对比:Nanbeige 4.2 实测解析
三条推文都叫 looped transformer 但完全是三种方案,选错能烧掉一周训练,这篇用 Nanbeige 4.2 实测数据帮你分清。
Looped transformer 一词在过去一个月被用来指代三种不同的权重复用方案。文章拆解了开源模型 Nanbeige 4.2 的做法:22 层网络循环两遍,同一组权重被应用 44 次,块计算量约为标准 Transformer 的两倍,KV cache 也随之翻倍。该设计保留了标准 Transformer 约 75% 的 token 效率;若共享 cache 可将显存减半,但质量增益持续偏低。另一种方案让状态在 token 间传递,使序列长度充当深度,其训练仅用 5 亿 token,却耗费约 20 倍的 GPU 时。
You’ve probably heard the buzz around Looped transformers, but which one?
Over the past month, the same word covered three reuse tricks.
Implementing the wrong one burns a training week. It also budgets a 3B file like a 3B decode.
We opened a looped transformer open model, Nanbeige 4.2:
> 22 layers, two passes > 44 applications, same weights > About twice the block compute > KV cache doubles with it
It kept about 75% of a standard Transformer's token efficiency.
Sharing the cache halved its memory. And quality gains were consistently lower.
A second design carries state token to token. Sequence length is the depth.
Training used 500 million tokens. That run cost about 20 times the GPU-hours.
So, Which reuse trick is the viral tweet? If extra loops do not raise quality, what are you paying? Does the file still need a second cache?
Full Breakdown ↓↓