FinAutoRubric:专家引导的金融研究 Agent 自动评分标准生成
FinAutoRubric: Expert-Guided Automatic Rubric Generation for Evaluating Financial Research Agents
一篇讲怎么给金融 Agent 打分的论文:专家定规则,Agent 生成评分标准,还附带 100 个查询的基准,做金融评估的可以看看。
FinAutoRubric 针对金融研究 Agent 的评估问题,让专家提供可复用的评估指南,由 Agent 和代码完成具体任务的评分标准(rubric)生成、审查与校验。流程中 writer agent 负责查询每个预期值,reviewer agent 负责核实,失败案例升级给人工处理。在三个专家编写的金融基准上,其评分标准与人类打分一致程度达到最强生成器水平,内部分析师在盲评中更偏好它。配套发布的 FinAutoRubric Benchmark 包含 100 个查询、78 个任务、覆盖八类资产,并显示早期模型的评分标准对更新一代模型仍有提升空间。
FinAutoRubric: Expert-Guided Automatic Rubric Generation for Evaluating Financial Research Agents
Evaluating finance research agents requires rubrics that reflect expert standards and fix the values correct as of an information cutoff. Expert-reviewed finance benchmarks rely on fixed, per-item rubrics, which are costly to extend and cannot encode each institution's own standard. In FinAutoRubric, experts specify reusable evaluation guidance, while agents and code carry out query-specific rubric generation, review, and validation. This expert guidance governs every agent, as prompts and as rules that code enforces, and a Task Bank of reusable criteria carries it across tasks. In long-horizon loops that follow the expert guidance, a writer agent researches every expected value and a reviewer agent verifies it, and failures escalate to a human. On three expert-authored finance benchmarks, its rubrics track expert scoring as closely as the strongest evaluated generator while stating the expert rubric's expected value for more criteria, their scores agree with human grading, and in-house analysts prefer them in a blind review. The released 100-query FinAutoRubric Benchmark, built from in-house analysts' key questions across 78 tasks and eight asset classes, shows that rubrics from an earlier model generation still leave headroom for a later one.