Vals AI进行无限代币实验
Vals AI CEO Rayan Krishnan on the unlimited tokens experiment that revealed an anxious obligation fo...
Vals AI花了150万测试AI工具使用,发现Devin很省代币,教你如何智能选择AI工具而不破产。
Vals AI CEO Rayan Krishnan进行了一项为期一个月的代币最大化实验,团队工程师每天消耗10亿至20亿代币,总花费约150万美元。这一支出相当于当月员工薪资的10倍。实验发现Cognition Devin工具代币效率很高,公司已决定更多采用。订阅模式相比代币计价能获得更好的价格模型。
Vals AI CEO Rayan Krishnan on the unlimited tokens experiment that revealed an anxious obligation fo...
Vals AI CEO Rayan Krishnan on the unlimited tokens experiment that revealed an anxious obligation for engineers to use models all the time, everywhere: "I wanted to do a tokenmaxxing experiment. I was able to get unlimited access for our team for a month for some of the coding tools. We had a lot of engineers spending between one to two billion tokens a day." "I went back and did some math, and it looked like in that month, we spent roughly $1.5 million worth of tokens... It was actually 10x more we were spending in tokens than employee salary for that month." "We cannot continue with this mode of operation for the next month. How do we intelligently figure out what are the right tools we should use, for what teams, and what projects?" "We found some pretty surprising insights. The Cognition Devin tool is actually very token-efficient, and so that's a place we've chosen to adopt more." "There are a lot of places where you can get better pricing models out of subscriptions as opposed to token-based pricing." "It's actually informed our strategy for how we can effectively tokenmaxx without spending $1.5 million per month." @RayanKrishnan @JenniferHli Your browser does not support the video tag. 🔗 View on Twitter a16z @a16z Vals AI co-founder and CEO Rayan Krishnan with a16z's Ben Horowitz and Jennifer Li on grading AI, what it costs, and who gets to make the rules: Every big industry eventually grows an independent testing layer. AI has credit ratings to learn from and Enron to avoid. Model capability today is still mostly self-reported. As public benchmarks saturate and models get better at optimizing for the tests themselves, Rayan makes the case for independent, continuously evolving evaluations. The harder problem is geopolitical. Reagan's "trust but verify" worked during the Cold War because you could fly over and count the missiles. No simple equivalent for AI models exists. In this conversation with Erik Torenberg, they get into how you measure a model's ability to improve itself, why every good benchmark eventually has to be retired, and what happens when token spend begins to rival employee salaries. 00:00 Intro 02:20 Llama 4 on public vs private benchmarks 05:24 Nobody agreed how to test humans either 06:55 What movie ratings teach us about AI 08:55 The Enron problem in benchmarking 11:36 Why a good benchmark has to be retired 13:22 Evals that run for weeks, not seconds 16:20 Where the real workday starts at 4pm 18:08 A firm really is just its evals 20:35 Why Sonnet can cost more than Opus 22:42 One engineer, 6 billion tokens in a day 25:05 Who should set the rules for models 28:55 Public sector enforces, private verifies 33:32 Why sovereign AI is inefficient and happening anyway 35:00 The AI version of trust-but-verify 37:15 Where cyber evals have to go next YouTube: youtube.com/watch?v=WO9c9q… @RayanKrishnan @ValsAI @bhorowitz @JenniferHli @eriktorenberg Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 5 🔄 2 ❤️ 16 👀 8250 📊 5 ⚡