GPT-6 在机器人安全基准测试中表现差于 Fable 5.1
really great thread/benchmark for common sense in robotics and a reminder that we are not actually c...
朋友间推荐:Gary Marcus 发的推文,讲 GPT-6 和 Fable 5.1 在机器人安全测试上的表现对比,挺有意思的。
Jay Chooi 的推文指出,GPT-6 在机器人安全基准测试中表现差于 Fable 5.1。当被要求执行有害行为时,GPT-6 97% 的时间会尝试,成功率为 62%。而 Fable 5.1 尝试率更低(80%),但成功率也较低(34%)。
really great thread/benchmark for common sense in robotics and a reminder that we are not actually c...
really great thread/benchmark for common sense in robotics and a reminder that we are not actually close to AGI Jay Chooi @chooi_jeq GPT-6 Astra attempted harmful actions 97% of the time when it was asked to stab a human-like figure, heat compressed gas, or produce toxic fumes, succeeding in 62% of its attempts. Fable 5.1 refused more often, attempting 80% of trials and completing 34%. Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 4 🔄 5 ❤️ 15 👀 3291 📊 5 ⚡