MoE推理群体路由有效秩研究
From Concentration to Differentiation and Back: Routing Effective Rank in MoE Reasoning Cohorts
这篇论文揭示了MoE模型推理过程中路由相似性的集中-分化-再集中规律,为理解MoE内部计算提供了新视角。
研究人员提出路由有效秩deff指标,用于分析MoE推理群体的内部计算重组。在10种MoE配置和5个数学/科学基准测试中,deff表现出可重复的低-高-低轨迹,在98.5%的3,105个模型-问题组中呈现明显的内部最大值。分解显示共同模式重分配约占轨迹的三分之二,剩余谱贡献约四分之一。推理 effort增加会使最大值延迟2.59个八度,并延长高秩期。
From Concentration to Differentiation and Back: Routing Effective Rank in MoE Reasoning Cohorts
Test-time scaling produces cohorts of reasoning rollouts, yet there is no standard label-free account of how their internal computation reorganizes as inference unfolds. We introduce routing effective rank deff, the entropy-effective dimensionality of a cross-rollout graph built from MoE expert-routing similarity. Across ten MoE configurations and five math/science benchmarks, deff exhibits a reproducible low-high-low trajectory, with a prominent interior maximum in 98.5% of 3,105 model-question cohorts: routing similarity is concentrated early, maximally differentiated at intermediate budgets, and reconcentrated later, and the timing of this maximum varies systematically with architecture and reasoning effort. An exact decomposition separates cohort-wide common-mode mass from residual spectral dimensionality: common-mode reallocation accounts for about two thirds of the trajectory, while the residual spectrum contributes about one quarter and retains substantial variation beyond the common mode. The decomposition further localizes behavior: among non-unanimous cohorts, increases in common-mode concentration strongly predict same-answer recoverability, and higher reasoning effort delays the maximum by 2.59 octaves (doublings of the token budget) and consistently expands the high-rank period across all four tested architectures, locating the effort effect in timing and duration rather than peak amplitude. Correctness comparisons separate structural monitoring from answer selection, positioning routing effective rank as a decomposable, label-free diagnostic of cohort organization - a principled spectral lens on how MoE reasoning cohorts differentiate and reconcentrate over inference time.