论文

后处理算法实现黑盒生成模型输出的统计属性对齐

Statistical attribute alignment for black-box generative AI via output post-processing

精选理由

一篇偏理论的论文:不用改模型内部,只在输出端做后处理就能让性别、种族等属性分布对齐目标,还证明了查询次数最优,做公平性或合成数据的可以看看。

arXiv 论文 2609.31607 研究黑盒访问场景下生成式 AI 的统计属性对齐问题,目标是让 m≥1 个输出的属性分布贴近用户指定目标分布。应用场景包括性别、种族、年龄等受保护属性的公平性控制,以及合成数据生成中的分布代表性。作者提出针对精确对齐与近似对齐两类任务的查询次数最小化算法,并证明在 m→∞ 时算法的最优性。在文生图和带地理编码的 persona 生成任务上,该后处理算法相比基于提示词的干预方式改善了属性对齐效果。

原文 · arXiv cs.LG

Statistical attribute alignment for black-box generative AI via output post-processing

Generative AI systems are increasingly used, but aligning their outputs with user requirements poses a continuing challenge. Here, we aim to ensure that the distribution of an attribute of an AI-generated output aligns with a user-specified target. This is motivated by examples such as fairness, where we want to ensure that a protected attribute (e.g., gender, race, or age categories) follows a desired distribution, and synthetic data generation, where we want the generated data to be representative of a target distribution. We study the practically important black-box access setting, where a user can repeatedly query a generative AI model. The goal is to return $m\ge 1$ outputs whose joint attribute distribution is as close as possible to this target. For both exact and approximate alignment, we develop algorithms that minimize the expected number of queries to the generator, and we further demonstrate their optimality as the number of requested outputs $m \rightarrow \infty$. Experiments on text-to-image generation and geocoded persona generation tasks show that our post-processing algorithms improve statistical attribute alignment, complementing prompting-based interventions.