技巧

Harrison Chase 谈 Goodhart 定律:自主智能体的评估困境

精选理由

Harrison Chase 转发了一条很扎心的观察:eval 驱动开发对窄任务好用,自主智能体就失灵了,看看这个死循环怎么破。

LangChain 创始人 Harrison Chase 转发 Hunter Gerlach 的观点,引用 Goodhart 定律——当指标成为目标,它就不再是好指标。他指出 eval 驱动开发适用于范围窄的任务,但对更自主的智能体不适用。Gerlach 提出的循环是:先创建高价值 eval,再对抗由此产生的 Goodhart 效应,然后重复这一过程。

原文 · Harrison Chase

This is great point Goodharts law: when a measure becomes a target, it ceases to be a good measure Initial take: eval driven development works for narrowly scoped things, but for more autonomous agents it doesn’t So then - how do you hill climb those? Hunter Gerlach @HunterGerlach 1. Create new high-value evals 2. Realize you need to fight against Goodhart's Law 3. Go to 1 🔗 View Quoted Tweet 💬 3 🔄 0 ❤️ 3 👀 566 📊 3 ⚡