技巧精选

实测个人智能体记忆并非越多越好:10 行左右最佳

精选理由

做个人智能体的可以看看:Haiku 4.5 实测记忆超过 10 行反而更糟,计数类任务该写代码而不是写记忆规则。

实验用 Claude Haiku 4.5 测试记忆条数与规则违规率的关系:无记忆时违规率 77%,10 行记忆降到 20%,更长后回升到约 25%。需要计数的任务(如累计消费总额)用文字规则会失败 44%,改用代码记录则失败 0%。靠智能体自己改写记忆的方法最好也只能到 48% 违规率,而直接告知全部偏好可低至 7.1%。

原文 · rohanpaul_ai

More memory does not keep helping personal agents, because relevant notes help up to a point and then extra lines start making the agent miss rules.

A personal agent can't learn every user preference through memory notes, and more notes eventually make it worse, so use code for anything it must count or track and keep memory short.

Written rules work for style, like signing texts with the user's first name. For a running spending total, a stated rule still failed 44% of the time, while code that kept the total failed 0%.

Memory size has a sweet spot. With Claude Haiku 4.5, violations dropped from 77% with no memory to 20% at 10 lines, then rose to about 25% with longer memories.

Agents that rewrite their own memory from user complaints improve early, then stall. The best methods ended near 48% violations, against 7.1% when the agent was simply told every preference.