论文

Prompt Minimization:压缩提示词长度而不损失输出质量

Prompt Minimization: Reducing Input Redundancy Without Sacrificing Output Fidelity

精选理由

arXiv 上这篇论文教你把提示词砍到最短还不掉输出质量,还给了三个可操作的压缩框架,写 prompt 的人可以看看怎么省 token。

arXiv 论文提出 prompt minimization 方法,目标是将提示词压缩到最小信息密度形式,同时保持输出保真度。论文指出冗长提示会损害 LLM 的推理和准确率,并增加推理延迟与计算开销,尤其当上下文包含整份文档或代码库时更明显。作者提出三个变体框架来识别和评估最小化提示词,实验显示压缩后的提示词输出与原长提示词相当。研究还表明输入空间存在大量冗余,多个不同提示词可产生等价输出。

原文 · arXiv cs.AI

Prompt Minimization: Reducing Input Redundancy Without Sacrificing Output Fidelity

Despite the growing capabilities of large language models (LLMs), prompt design remains largely heuristic and ad hoc. This project will explore $\textit{prompt minimization}$, the process of reducing prompts to their smallest, most information-dense form while preserving output fidelity. Practically, shorter prompts reduce computational overhead and inference latency, especially when large contexts, such as entire documents or codebases, are included unnecessarily. Further, longer prompts can damage LLM reasoning and accuracy. Theoretically, the existence of multiple prompts yielding equivalent outputs suggests a high degree of redundancy in the input space, raising fundamental questions about what information is essential to elicit specific model behaviors. We propose three variant frameworks to identify and evaluate minimal prompts and demonstrate that minimal prompts often produce outputs comparable to those of their longer counterparts. These findings suggest new directions for efficient prompt engineering and deepen our understanding of input compression in LLMs.