产品多源确认

GPT 模型 computer-use 能力实测:像人一样操作 Safari 几乎不出错

精选理由

连常批评 OpenAI 的博主都服了:GPT 操作 Safari 像真人,任务独立完成几乎零失误,可以看看实测细节。

一位长期批评 OpenAI 近期决策的博主实测 GPT 模型的 computer-use 功能后给出正面评价。测试显示模型能用 Safari 像真人一样控制浏览器,几乎不犯错误。博主认为模型独立完成任务的可靠程度超出预期。这条推文是对 GPT 系列模型浏览器操作能力的第一手使用反馈。

原文 · kimmonismus

Even though I’ve recently written a lot of critical posts about OpenAI’s latest decisions, I’m still pleasantly surprised by how phenomenal the GPT models' computer-use capabilities have become.

It uses Safari just as if a human were actually controlling the browser. It hardly makes any mistakes and is so good at carrying out tasks independently that I’m genuinely surprised. Credit where credit is due.