行业

Vals AI CEO:企业需可防御的AI评估

Vals AI CEO Rayan Krishnan on why defensible evals are critical for enterprises' survival: "The lab...

精选理由

Vals AI CEO谈企业如何通过独立评估体系 justify AI投资ROI,当token spend可能超过salary spend时,评估体系决定企业生死。

Vals AI CEO Rayan Krishnan指出,企业面临AI投资回报不明确的问题,代币支出可能超过薪资支出。他强调,企业需要清晰展示AI投资的ROI,而独立、持续演进的评估体系将成为企业长期竞争的关键。随着公共基准饱和,模型优化测试本身的能力越来越重要。

图片来源 · a16z
原文 · a16z

Vals AI CEO Rayan Krishnan on why defensible evals are critical for enterprises' survival: "The lab...

Vals AI CEO Rayan Krishnan on why defensible evals are critical for enterprises' survival: "The labs side of it is very clear. If you're raising lots of money, investing heavily in building models, it's essential for you to show why your model is getting better and why the customer should pay a premium for them." "But what I think is still underappreciated is, on the enterprise side, this is turning out to be existential as well." "We're in this world where it is still very unclear what ROI looks like and how to value this intelligence that's being used... Token spend may start to eclipse salary spend." "If this is such a meaningful line item in your costs, you have to justify the ROI much more cleanly... Over time, I think a firm really is just its evals." "The ability for a company to make its evals legible in order to solve this ROI calculus is going to be the reason why that company wins out over its competitors in the long term." @RayanKrishnan @eriktorenberg Your browser does not support the video tag. 🔗 View on Twitter a16z @a16z Vals AI co-founder and CEO Rayan Krishnan with a16z's Ben Horowitz and Jennifer Li on grading AI, what it costs, and who gets to make the rules: Every big industry eventually grows an independent testing layer. AI has credit ratings to learn from and Enron to avoid. Model capability today is still mostly self-reported. As public benchmarks saturate and models get better at optimizing for the tests themselves, Rayan makes the case for independent, continuously evolving evaluations. The harder problem is geopolitical. Reagan's "trust but verify" worked during the Cold War because you could fly over and count the missiles. No simple equivalent for AI models exists. In this conversation with Erik Torenberg, they get into how you measure a model's ability to improve itself, why every good benchmark eventually has to be retired, and what happens when token spend begins to rival employee salaries. 00:00 Intro 02:20 Llama 4 on public vs private benchmarks 05:24 Nobody agreed how to test humans either 06:55 What movie ratings teach us about AI 08:55 The Enron problem in benchmarking 11:36 Why a good benchmark has to be retired 13:22 Evals that run for weeks, not seconds 16:20 Where the real workday starts at 4pm 18:08 A firm really is just its evals 20:35 Why Sonnet can cost more than Opus 22:42 One engineer, 6 billion tokens in a day 25:05 Who should set the rules for models 28:55 Public sector enforces, private verifies 33:32 Why sovereign AI is inefficient and happening anyway 35:00 The AI version of trust-but-verify 37:15 Where cyber evals have to go next YouTube: youtube.com/watch?v=WO9c9q… @RayanKrishnan @ValsAI @bhorowitz @JenniferHli @eriktorenberg Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 2 🔄 1 ❤️ 12 👀 6989 📊 3 ⚡