ScienceBuddy通过交替优化提示和模型提升性能
ScienceBuddy研究教你如何交替优化提示和模型,比单独改进任一方法效果更好。
ScienceBuddy研究显示,交替优化提示和模型可将科学代理准确率从42.2%提升至73.3%。仅优化提示和技能,4B模型在生物任务上的准确率从31.1%提升至51.1%。仅重新训练模型,4次内解决问题的比例从48.3%提升至67.8%。
Don't pick between better prompts and a better model: improving both in turns lifted a science agent from 42.2% to 73.3% accuracy.
Researchers correct AI agents all the time, but fixing an answer in chat doesn't make the agent better at the next task.
Correcting an AI agent in chat, the fix usually dies with the conversation.
ScienceBuddy, an AI assistant for scientists, turns that feedback into scored test tasks. Then it takes turns: rewrite the agent's prompts and skills, retrain the model, and repeat.
With a small 4B model on biology tasks, prompt and skill changes alone raised accuracy from 31.1% to 51.1%. Retraining alone also helped, raising the share of problems solved within 4 tries from 48.3% to 67.8%.
If you build agents, save every user correction as a test, and keep upgrading your prompts and your model in turns.
– arxiv. org/abs/2609.17523v1
Title: "ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents"