论文73°

基础LLM与LRM范式转变研究

精选理由

François Chollet分享LLM到LRM的范式转变,解释了为什么LRM拥有基础LLM缺乏的流体智能。

研究指出2024年前基础LLM与现代LRM的关键区别在于从归纳范式到演绎范式的转变。基础LLM在ARC 1基准测试上表现仅约10-15%,即使规模扩大10万倍也仅从0%提升到10%。而LRM在2025年已在该基准上达到饱和,展现出显著的流体智能。

原文 · François Chollet

The critical distinction between base LLMs (2024 and earlier) and modern LRMs is not symbolic tool use. It's the switch from a transductive paradigm (intuit the answer to the query) to an inductive paradigm (intuit the program/instructions that produce the answer to the query).

They're trained to be inductive, and they perform test-time induction, i.e. test-time prediction of a NL program / reasoning chain. This unlocks entirely new capabilities -- in particular fluid intelligence. Base LLMs, to this day, have ~0 fluid intelligence. LRMs have substantial levels of fluid intelligence.

The performance of LLMs on ARC 1 (a benchmark from 2019) remains ~10-15% today. Scaling them up by a factor ~100,000x got them from 0% to 10%. Meanwhile LRMs the same size or smaller saturated ARC 1 in 2025.