模型精选73°

Vals AI发布递归自改进指数评估模型

Vals AI CEO Rayan Krishnan on how they score a model's ability to build its own successor: "One ben...

精选理由

Vals AI CEO详解如何评估模型自我改进能力,解决了行业缺乏统一标准的问题。

Vals AI推出了递归自改进指数(RSI),用于评估模型构建自身后续版本的能力。该指数通过预训练、后训练、工程级别等多个维度的代理指标进行评估。目前大模型实验室开始在其模型卡片中报告RSI潜力,但缺乏统一标准。这一评估方法旨在提供跨模型的公平比较基准。

图片来源 · a16z
原文 · a16z

Vals AI CEO Rayan Krishnan on how they score a model's ability to build its own successor: "One ben...

Vals AI CEO Rayan Krishnan on how they score a model's ability to build its own successor: "One benchmark we released recently is our Recursive Self-Improvement Index. It's a topic which a lot of the big labs have been talking about and starting to report on in their model cards." "There isn't a shared language to talk about the RSI potential of models. So we created this as an apples-to-apples way to actually benchmark across the models." "In an ideal world, what you want to do is take a frontier model and have it train the next version of itself and see where the delta comes from. But obviously that's very expensive and slow." "What we're doing is forming a set of proxies for every part of the process it takes to build the next version of the model." "There's work around pre-training, post-training, harness-level engineering, and then seeing in which mechanisms and behaviors the models are able to do very good research work and build something new, and where they're struggling." @RayanKrishnan @JenniferHli Your browser does not support the video tag. 🔗 View on Twitter a16z @a16z Vals AI co-founder and CEO Rayan Krishnan with a16z's Ben Horowitz and Jennifer Li on grading AI, what it costs, and who gets to make the rules: Every big industry eventually grows an independent testing layer. AI has credit ratings to learn from and Enron to avoid. Model capability today is still mostly self-reported. As public benchmarks saturate and models get better at optimizing for the tests themselves, Rayan makes the case for independent, continuously evolving evaluations. The harder problem is geopolitical. Reagan's "trust but verify" worked during the Cold War because you could fly over and count the missiles. No simple equivalent for AI models exists. In this conversation with Erik Torenberg, they get into how you measure a model's ability to improve itself, why every good benchmark eventually has to be retired, and what happens when token spend begins to rival employee salaries. 00:00 Intro 02:20 Llama 4 on public vs private benchmarks 05:24 Nobody agreed how to test humans either 06:55 What movie ratings teach us about AI 08:55 The Enron problem in benchmarking 11:36 Why a good benchmark has to be retired 13:22 Evals that run for weeks, not seconds 16:20 Where the real workday starts at 4pm 18:08 A firm really is just its evals 20:35 Why Sonnet can cost more than Opus 22:42 One engineer, 6 billion tokens in a day 25:05 Who should set the rules for models 28:55 Public sector enforces, private verifies 33:32 Why sovereign AI is inefficient and happening anyway 35:00 The AI version of trust-but-verify 37:15 Where cyber evals have to go next YouTube: youtube.com/watch?v=WO9c9q… @RayanKrishnan @ValsAI @bhorowitz @JenniferHli @eriktorenberg Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 2 🔄 0 ❤️ 10 👀 5730 📊 2 ⚡