70位从业者调研:各家实验室模型水平分层
阅读原文: https://t.co/XHn3uau4Y1
有 guy 找了 70 个从业者给各家模型排座次,Anthropic 和 OpenAI 并列第一梯队,DeepSeek 落后一代,结果挺真实
Chris Barber 调研了 70 位 AI 从业者,让他们给各实验室的模型代际水平分层。Anthropic 和 OpenAI 被一致认为处于 Frontier 梯队,Google DeepMind 和 xAI 被一致归为落后一代(n-1)。Moonshot(Kimi)、智谱(GLM)、DeepSeek 获过半投票归入 n-1,而阿里(Qwen)、字节(Seed)、MiniMax、腾讯(Hy)多被认为落后两代(n-2)。Apple 和 Amazon(Nova)则被一致认为落后三代以上。
阅读原文: https://t.co/XHn3uau4Y1
阅读原文: x.com/chrisbarber/st… Chris Barber @chrisbarber I asked people which tier they'd put each lab in: Frontier: - Anthropic: consensus - OpenAI: consensus n-1 (1 generation behind): - Google DeepMind: consensus - xAI: consensus - Meta (Muse Spark): over half say n-1 - Moonshot (Kimi): over half say n-1, some n-2 - Zhipu (GLM): over half say n-1, some n-2 - DeepSeek: over half say n-1, some n-2 - Thinking Machines (Inkling): leans n-1, some put n-2 n-2 (2 generations behind): - Alibaba (Qwen): leans n-2, some put n-1 - ByteDance (Seed): over half say n-2 - Xiaomi (MiMo): leans n-2, some put n-1 - MiniMax: over half say n-2, some n-1 - Tencent (Hy): over half say n-2 - NVIDIA (Nemotron): leans n-2, some put n-3+ n-3+ (3+ generations behind): - Microsoft (MAI): over half say n-3+, some n-2 - Cohere: over half say n-3+, some n-2 - Mistral: over half say n-3+ - Amazon (Nova): consensus - Apple: consensus Consensus = 75%+, over half = 51%-74%, leans = the most common pick but under half, some = 20%+. I asked 70 people. Engineers/researchers/founders at AI data cos, people at labs (I didn't count votes for their own lab), and founders/engineers at other startups. I didn't vote. 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 0 👀 173 ⚡