模型精选

Grok 4.7在FrontierCode 1.1基准上得分略低于Grok 4.6

On FrontierCode 1.1, our benchmark for real-world engineering tasks, Grok 4.7 scores slightly below ...

精选理由

Grok 4.7在Cognition的工程基准上比上代4.6还低,改代码爱超范围是主因

Cognition在其面向真实工程任务的基准FrontierCode 1.1上测得,Grok 4.7总分略低于Grok 4.6。Grok 4.7在不少高难度任务上表现不错,但会在部分任务中超范围改动(over scope)。这一倾向把它的聚合分数拖到低于上一代Grok 4.6。详细分析见devin.ai/blog/grok-4-7。

图片来源 · Cognition
原文 · Cognition

On FrontierCode 1.1, our benchmark for real-world engineering tasks, Grok 4.7 scores slightly below ...

On FrontierCode 1.1, our benchmark for real-world engineering tasks, Grok 4.7 scores slightly below Grok 4.6. While strong on many hard tasks, it tends to over scope on others, leading it to trail behind Grok 4.6 in the aggregate. Read more: devin.ai/blog/grok-4-7 💬 1 🔄 0 ❤️ 2 👀 152 📊 1 ⚡