21名顶尖学者:LLM是通往AGI的死胡同
Y’all gonna get mad if I said I told you so …
GaryMarcus分享21名顶尖学者观点,LLM无法实现AGI,真正的智能需要视觉基础。
斯坦福、牛津、DeepMind等机构的21名研究人员发表论文《视觉通用智能》,称LLM是死胡同。论文指出文本缺乏物理和几何特性,无法构建真正的通用智能。研究者认为下一代模型应从原始视觉经验、图像和视频开始学习,而非仅依赖文本。
Y’all gonna get mad if I said I told you so …
Y’all gonna get mad if I said I told you so … Superman @thesupermanmx Researchers argues OpenAI and Anthropic will never get us to AGI. 21 top researchers from Stanford, Oxford, DeepMind, CMU & Meta released a paper saying LLMs are a dead end. It’s called “Visual General Intelligence”. And it completely flips how we look at artificial intelligence. For years, the playbook has been simple: feed mountains of web text into a Transformer, scale up the parameters, and watch reasoning emerge. GPT proved language can take you far. But text is fundamentally limited. It’s a compressed, human-abstracted symbol system. It lacks physics. It lacks geometry. It lacks the raw, unadulterated reality of the physical world. The paper argues that true general intelligence cannot be built on words alone. It requires a vision-centered foundation. Instead of starting with language and translating pixels into text, the next generation of models must start with raw visual experience, images, spatial geometry, and continuous video. Think about how humans learn. A baby understands gravity, permanence, and spatial reasoning long before it ever learns to string a sentence together. Vision isn't just an input modality. It is the core operating system of physical reality. When models learn natively from visual streams and video dynamics, they don't just memorize text patterns. They learn physics. They learn cause and effect. They build a true internal model of the world. 🔗 View Quoted Tweet 💬 3 🔄 1 ❤️ 8 👀 632 📊 3 ⚡