模型73°

TypeSafe CEO批评公共基准测试

STOP making "Jevbench"es, stop asking for public benchmarks, they completely miss the point of Jev a...

精选理由

TypeSafe CEO谈为什么AI能解决难题却难做简单工作,解释Jev模型设计理念,直言公共基准测试毫无意义。

TypeSafe CEO @CompleteSkeptic 表示,Jev基准测试完全忽视了Jev模型的核心价值。他认为选择正确的任务比一切都重要,并指出所有基准测试都容易被操纵。TypeSafe拒绝在API层面使用公共基准测试和拒绝机制,专注于在软件中实现可靠决策而非聊天功能。

图片来源 · Latent.Space
原文 · Latent.Space

STOP making "Jevbench"es, stop asking for public benchmarks, they completely miss the point of Jev a...

STOP making "Jevbench"es, stop asking for public benchmarks, they completely miss the point of Jev and you won't believe how easy it is to game every benchmark you hold dear This is @CompleteSkeptic 's bitterest lesson of all: picking the right task beats everything Your browser does not support the video tag. 🔗 View on Twitter Latent.Space @latentspacepod Jev and the System One Model: RLCD, intelligence/$, reliable AI, & the end of chat-first AI latent.space/p/jev @typesafeai CEO @CompleteSkeptic explains why AI can solve extraordinarily hard problems yet still fail to automate basic work, why Jev is built for reliable decisions inside software instead of chat, why TypeSafe rejects public benchmarks and refusals at the API layer, why data and the right task matter more than brute-force compute, how System One Models could reshape coding agents and software, and why even with $1 billion he wouldn’t pre-train a model from scratch. Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 0 👀 92 ⚡