评论员谈 Beam:不靠 Claude 蒸馏,用 RL 压缩轨迹
有人聊 Beam 这模型:猜它没用 Claude 蒸馏,纯靠 RL 炼出来的,看看不带 Claude 血统能打几分。
评论员 teortaxesTex 发推称,只要算力充足,用更多 RL 算力配合长度惩罚来压缩轨迹(trajectories)是相对直接的做法。他认为 Beam 的看点在于推测完全没有经过蒸馏。他提出的问题是:不依赖 Claude 蒸馏数据,模型能竞争到什么程度。这条推文属于观点判断,Beam 尚未公布基准成绩。
I think burning more RL compute to compress trajectories with length penalty is ≈straightforward if you have the compute. What makes me excited about Beam is that I presume it wasn’t distilled at all. Time to see how competitive you can get without the taint of Claude.