论文

Internalizer 超网络可将文档编译为 284B 参数模型的 LoRA 适配器

Internalizer: Portable Context-to-Parameter Mapping for Very Large Language Models

精选理由

把一份文档直接编译成 LoRA 权重,284B 的 DeepSeek v4 Flash 也能用,top-1 准确率从 63.4% 拉到 84.9%

论文提出 Internalizer,一种 Context-to-Parameter Mapping 超网络,能把文档直接映射为 LoRA 适配器。它为冻结的 284B 参数 DeepSeek v4 Flash 生成文档专属适配器,比此前 14B 参数的演示规模大两个数量级。其参数大部分位于模型无关的 trunk 中,可在小模型上低成本训练后迁移到大模型。在最长 4096 token 的未见文档上,生成的适配器达到 84.9% top-1 和 97.8% top-5 教师强制准确率,而基座模型仅为 63.4% 和 83.5%。训练完成后,一次前向传播即可将任意文档转为适配器。

原文 · arXiv: DeepSeek

Internalizer: Portable Context-to-Parameter Mapping for Very Large Language Models

Hypernetworks that map a context directly to a LoRA adapter let a large language model carry that context in its weights, but prior work has demonstrated them only on base models of up to 14 billion parameters. We present the Internalizer, a state-of-the-art, portable Context-to-Parameter Mapping hypernetwork that generates document-specific LoRA adapters for the frozen 284B-parameter DeepSeek v4 Flash, a target two orders of magnitude larger than in any previous work. Most of its parameters live in a model-agnostic trunk with only thin entry and exit layers per base model, so it trains cheaply against small models before being ported to the large one. On unseen documents of up to 4096 tokens, the generated adapters reach 84.9% top-1 and 97.8% top-5 teacher-forced accuracy against 63.4% and 83.5% for the base model, with nothing in the context window but a three-word instruction. Once the hypernetwork is trained, a single forward pass turns any document into an adapter for such a model, which could be served alone for speed or alongside the document in the window to raise accuracy further.