Gary Marcus 转发观点:当前 agentic 框架完全不够安全,需推倒重建
Gary Marcus 转发了一位 20 年经验工程师的长文,观点很尖锐:90% 成本该花在测试和安全上,agent 框架得推倒重来。
Romeo Lupascu 在评论采访中提出,AI 测试环境可以做到安全且可被监控,但监控必须由确定性软件而非另一个 AI 来完成。他结合 20 多年软件开发经验指出,即便传统软件也长期测试不足,此前 bot 攻破 HuggingFace 等系统正说明这一点。他估算约 90% 的成本应投入测试与安全而非模型本身,并断言当前 agentic 框架必须基于不同基础原理从头重新设计。Gary Marcus 在 X 上转发并连用三个💯表示赞同。
💯💯”The current agentic framework is fully unsafe and must be redesigned from scratch based on different fundamental principles.” Romeo Lupascu @RomeoLupascu Some observations around the interview Q:Can be AI test environments be safe? IMO: yes, but it needs effort in being built Q: Can AI test environments be monitored safely? IMO: yes, but it needs effort to be correctly setup and tested then it needs also a big chunk of compute to do the actual monitoring. And no, monitoring can't be none by another AI but deterministic software Q: Can models be fully tested? IMO: no, not in a practical way as deterministic software (classic/non AI) can (and yet it isn't) @JensenHuang Mr. Huang is correct in his evaluation but we need to realize what are we dealing with in this AI tech. Investors push the labs for results fast and they will always do as they want that profit yesterday if possible. Dario and Sam are probably begging for help (to be regulated exteranlly) precisely because they took too much debt and spending any of that money in "testing & security" related work that is invisible for anyone and most important to investors, is impossible. I worked for more than 20 years in software development and I can assure you that even the non AI software is under tested because doing full testing is too expensive. The fact that some bots were able to hack huggingface and orher systems means not only that the labs were sloppy with their own security but that the classic software they "cracked" wasn't tested sufficiently to clear out all the defects. If a software has zero defects then it is unhackabke. We were swimming before AI in an unsafe sea of classic software that could have been made safe and then we rushed to add even more unsafe software (AI) on top. What can go wrong when our society is totally dependent on these systems? Slowing down in AI and everything in software engineering doesn't mean to work less. It means to divert most if the investment funds in parts that aren't part of the product but of the tooling that produces the product. You must convince all the investors that pushing for "fast mirackes" in AI is a very bad mistake and they need to allow the labs to spend most of their money in indirect costs for testing and monitoring. That is what "AI slowdown" actually means. My gut guess is that around 90% of the cost must go in testing and safety than in building the models themselves. Last but not least. The current agentic framework is fully unsafe and must be redesigned from scratch based on different fundamental principles. @GaryMarcus @Grady_Booch @burkov @ChombaBupe @ezraklein @andersoncooper youtu.be/TxyayEjTiZQ?si… 🔗 View Quoted Tweet 💬 13 🔄 15 ❤️ 72 👀 7366 📊 18 ⚡