论文

PaSS:用迭代零空间投影提取人格子空间,实现 LLM 行为控制

Extracting Persona Subspaces Through Iterative Nullspace Projection For Modulation

精选理由

PaSS 不训练模型,用 INLP 在推理时把人格拆成多维子空间来放大或抑制某种特质,在 MATH-500、GSM8K 上验证比单方向激活导向更细。

论文提出 PaSS,一种推理时控制范式,将 LLM 的人格建模为潜在空间中的多维子空间,而非单一主方向。该方法借助 Iterative Nullspace Projections(INLP)通过迭代概念擦除线性分离人格相关方向,无需对比监督样本,也无需重训练。研究在 MATH-500、GSM8K、IFEval、TinyAlpaca 等任务上对六种人格进行因果评估,结果显示其调制强度优于单方向加性方法,同时保持内容保真度。论文还逐一分析子空间内被剥离出的方向,揭示其各自编码的人格行为侧面。

原文 · arXiv cs.AI

Extracting Persona Subspaces Through Iterative Nullspace Projection For Modulation

Large Language Models (LLMs) can adopt distinct personas to tune their semantics, expertise, and perspective to different users and tasks. Precise control over these traits is critical to ensure safety and reliability in model behavior. Existing methods like activation steering and prompt-based persona induction reduce a persona to a single dominant direction, missing the finer, nested traits that emerge only once that dominant signal is factored out. We introduce modulation as a setting where the persona context is already embedded in the content being manipulated, requiring control methods to amplify or suppress a trait already present rather than inject it from scratch. PaSS is an inference-time control paradigm that models personas as multi-dimensional subspaces in a model's latent space without supervised contrastive examples. The persona subspaces are extracted via iterative concept erasure and applied to modulate persona-guided generation without retraining. To extract this subspace, we use Iterative Nullspace Projections (INLP) to linearly and iteratively isolate persona-specific directions. We causally evaluate six personas against diverse tasks like MATH-500, TinyAlpaca, GSM8K, and IFEval, showing that discriminative, iterative subspace extraction captures diverse traits underlying a given persona, enabling stronger and larger modulation than single-direction additive methods, while maintaining content fidelity. We further study individual peeled directions within each subspace to uncover the distinct aspects of persona behavior they encode. Overall, we show that persona subspaces offer a controllable, interpretable, and generalizable framework for modulating LLM behavior without sacrificing task performance.