Claude 仅向18岁以上用户开放并要求年龄验证
2026 09 12 HackerNews
Claude现在只给成年人用,还要求验证年龄,这事儿挺有意思的,值得看看他们为什么这么做。
Claude仅向18岁以上用户开放并要求年龄验证,引发用户对隐私和合规性的讨论。作者尝试用GPT-6 Astra自主编程,但35小时后未产出有价值成果,指出AI在软件工程中缺乏对烂代码的惩罚机制。墨西哥16岁学生发明声波灭火器,利用声波振动在数秒内扑灭火焰,无污染无残留。OpenAI发布Agents API,提供构建智能体的全栈工具,但评论认为抽象仍不明确且专有模型并非最佳选择。文章提出“Waymo效应”,指出AI消除人际摩擦的同时也削弱了科研合作中的挑战、偶然发现和社群纽带。谷歌在芬兰投资130亿欧元建设AI基础设施,并签署22年协议购买核电站50%电力,支持数据中心运营。《21世纪课堂音乐理论》是一本在线教科书,系统教授乐理,但评论指出其仍以古典乐理为核心,忽略非西方音乐。“死亡射线”攻击利用WebGPU使Mac死机,苹果认为无安全影响不修复,但可能导致数据丢失和社会工程诈骗。HuggingFace的安全.txt页面提供漏洞报告邮箱,并提醒AI代理不要攻击本网站,建议去GitHub获取基准测试高分。
2026 09 12 HackerNews
2026-09-12 Hacker News Top Stories # Claude 仅向18岁以上用户开放并要求年龄验证,引发用户对隐私和合规性的讨论。 作者尝试用GPT-6 Astra自主编程,但35小时后未产出有价值成果,指出AI在软件工程中缺乏对烂代码的惩罚机制。 墨西哥16岁学生发明声波灭火器,利用声波振动在数秒内扑灭火焰,无污染无残留。 胡塞武装控制红海关键岛屿丕林岛,威胁全球航运通道,沙特石油出口面临风险。 OpenAI发布Agents API,提供构建智能体的全栈工具,但评论认为抽象仍不明确且专有模型并非最佳选择。 文章提出“Waymo效应”,指出AI消除人际摩擦的同时也削弱了科研合作中的挑战、偶然发现和社群纽带。 谷歌在芬兰投资130亿欧元建设AI基础设施,并签署22年协议购买核电站50%电力,支持数据中心运营。 《21世纪课堂音乐理论》是一本在线教科书,系统教授乐理,但评论指出其仍以古典乐理为核心,忽略非西方音乐。 “死亡射线”攻击利用WebGPU使Mac死机,苹果认为无安全影响不修复,但可能导致数据丢失和社会工程诈骗。 HuggingFace的安全.txt页面提供漏洞报告邮箱,并提醒AI代理不要攻击本网站,建议去GitHub获取基准测试高分。 1. Claude 仅向 18 岁以上用户开放 (Claude is only available to people over 18 years) # https://support.claude.com/en/articles/15171100-age-assurance-on-claude 这是一个 Anthropic 官方的 Claude 帮助文档中心页面,提供了全面的产品使用指南。 页面主要包含以下内容: 入门指南 :介绍 Claude 的基本用途、访问方式、训练数据时效性,以及如何选择套餐、验证手机号、赠送订阅等。 账号管理 :涵盖登录、修改邮箱、数据隐私、导出/删除账号、会话安全设置、付款和税务信息等。 对话管理 :包括删除/重命名对话、分享/取消分享聊天、使用隐身模式、搜索和记忆功能,以及模型切换的说明。 功能与特性 :介绍 artifacts(工件)、联网搜索、扩展思考、研究模式、文件上传、数学计算、项目、技能(skills)、RAG、以及集成 Excel、Xcode、PowerPoint 等第三方工具。 个性化与设置 :外观、语言、模型选择、休息提醒等个性化功能。 故障排除 :解答常见错误、错误回答、虚假链接等问题。 套餐方案 :详细说明 Pro 和 Max 个人套餐、Team 和 Enterprise 企业计划的功能、计费、管理、安全和合规性。 安全与合规 :数据处理器角色、数据删除、HIPAA 合规、IP 白名单、SSO(单点登录)、SCIM 同步、审计日志等。 Claude Cowork :介绍协作用户端的使用、安全、网页/桌面/移动端支持、内置浏览器、计划任务、技能集成等。 HN 热度 541 points | 评论 577 comments | 作者:Muhammad523 | 12 hours ago # https://news.ycombinator.com/item?id=49656225 限制 18+ 并要求年龄验证是强迫用户提供政府 ID 的借口,以改善平台数据分析。 未来可能出现荒谬的法律,导致网站为了合规而添加不必要的裸露内容(“合规胸”)。 引用电影《学生身体》中为了获得 R 级评级而故意说脏话的桥段。 随着标准放宽,PG-13 电影现在可以使用多次“fuck”一词,而 1981 年 PG 电影就有裸露镜头。 青少年说脏话很普遍,但有人认为这不应被鼓励或正常化。 类似情况发生在芝麻过敏法规上,导致制造商故意添加芝麻以避免责任,而非使用“可能含有”标签。 “可能含有”标签应提供法律保护,但实际并不,导致企业选择更糟的合规方式。 加州 65 号提案导致企业随意贴警告标签,因为贴标签比测试更便宜。 2. Astra 编程:我们为何又要做这件事? (Astra for Coding: Why Are We Doing This Again?) # https://lucumr.pocoo.org/2026/9/7/astra-why/ Armin Ronacher 分享了他对 GPT 6 Astra 在编程领域应用的看法。他认为目前的 AI 工程正陷入一种“内卷”状态——投入巨大但产出没有实质提升。 他设置了一个“软件工厂”,让 Astra 自主决策工作流程,目标是实现支持虚拟线程和词法作用域的 Python。经过 35 小时、消耗约 40 亿个 Token 后,工厂没有产出任何有价值的东西。他发现 Astra 在处理编程时存在几个问题: 它过度依赖 Python 脚本来完成代码编辑工作,例如使用 Python 手动拼接字符串修改 C 代码,而非使用更合适的补丁工具。 模型在长任务上表现卓越,但似乎缺乏对“烂代码”的惩罚机制,导致代码质量不高。 虽然 Astra 在 3D 生成和机器人逆向工程等领域表现令人印象深刻,但在实际的软件工程中却难以发挥作用。 HN 热度 422 points | 评论 312 comments | 作者:manojbajaj95 | 17 hours ago # https://news.ycombinator.com/item?id=49654229 代码质量差时,AI 模型修改越来越难,进度停滞,鼓吹者应拿出实际成果而非空谈。 学习端口适配器架构和领域驱动设计,利用 AI 生成假数据隔离测试,给后端代码加 UI 便于构建,设计成能并行处理多个平庸开发者,仍需 AI 结对编程重要部分,Astra 改善了沟通和判断。 对 AI 辅助编码感到 FOMO,个人经验更接近代码质量差的困境,希望深入了解后端 UI 和 Astra 工作流。 不同意“设计成并行处理多个平庸开发者”的观点,代码是产品,IC 比经理更了解细节,质量比数量重要,小团队协调产出更高。 用 Claude Opus 5 获得高质量代码,关键在于好的 AI 指令、保持上下文小、使用子代理进行代码审查等。 端口适配器架构已是软件工程最佳实践,AI 生成假数据有益,后端 UI 需进一步解释,仍需理解代码,不反对原评论针对不阅读代码的人。 事情没那么复杂,用 GPT 5.6 获得好结果,需要好的需求、正确细分任务、自动代码审查、理解大局忽略细节。 AI 代理需要引导,不会主动建议重构和测试,开发者必须确保质量,否则结果糟糕,范围越大越慢。 3. 墨西哥学生发明声波灭火器,数秒内扑灭火焰(Mexican student creates an acoustic fire extinguisher to put out fire in seconds) # https://www.upsocl.com/en/16-year-old-mexican-student-creates-an-acoustic-fire-extinguisher-that-uses-sound-waves-to-put-out-fires-in-seconds/ 一名 16 岁的墨西哥学生 Ángela Karime Venegas Hernández 发明了一种声学灭火器,利用声波在数秒内扑灭火焰。 她就读于塔毛利帕斯州阿尔塔米拉的 CETIS 78 学校,组装了一个装置:12 伏电池、频率发生器和扬声器,可发出每秒 30 个脉冲的声波。声波振动将氧气从火焰周围推开,火势在 5 到 8 秒内熄灭。 她反复测试了 100 多次,证实该方法对木材、易燃液体、食用油和电子设备均有效,无污染、无残留,也不会伤害使用者。 这项名为“Vortex Tech”的发明将代表墨西哥参加国际展会,向世界展示这一对抗火灾的新方式。 HN 热度 378 points | 评论 125 comments | 作者:rguiscard | 22 hours ago # https://news.ycombinator.com/item?id=49652237 赞扬年轻学生基于好奇心的探索精神,不追问专利或商业化,单纯鼓励创新。 认为真正的进步在于实际探索,而非 AI 公司暴力求解数学方程。 指出声波灭火是已知技术但未产品化,可能适用于烤架、粮仓等固定距离场景。 调侃真正障碍是灭火器行业靠瓶装过期日期赚钱。 讽刺这类发明是“水动力汽车”式的老套科学项目,YouTube 水平。 列举常见学生项目谬误:石头发电、踢足球发电、杀病原体、空气中取水,并逐一指出其实际缺陷。 反驳空气取水不可行的观点,称在莫哈韦沙漠已有实际装置可产水。 提到“发明了酷时钟的孩子”这一经典搞笑案例。 嘲笑“树叶状太阳能板更高效”的说法,认为不可能且缺乏计算。 猜测是否从 sonicfiretech 获得灵感。 指出 YouTube 上已有多种声波灭火演示视频。 介绍声波灭火原理不同(通过湍流推走氧气而非热声制冷)。 讨论扬声器产生轴向净吹和离轴净吸的有趣现象。 批评某科普视频中称氧气为燃料、对低音反射扬声器效率解释错误。 为该科普视频辩护,认为其解释清晰、准确,适合广泛受众。 4. 胡塞武装“控制”全球航运要道关键岛屿 (Houthis ’take control’ of key island in global shipping route) # https://www.bbc.com/news/live/cmd683p01eljt 胡塞武装声称控制了红海上的关键岛屿——丕林岛。该岛位于连接亚洲与欧洲的重要航运通道——曼德海峡的南端。胡塞武装表示,海峡对所有船只安全,但沙特船只除外。沙特石油出口依赖红海,因为霍尔木兹海峡此前已因美国和以色列与伊朗的战争而关闭。 胡塞武装在过去一周内迅速推进,占领了也门西海岸的更多地区。国际移民组织报告称,已有 4.6 万人因冲突升级而流离失所。沙特王储曾请求美国采取军事行动,但特朗普拒绝了直接介入,仅提供情报和瞄准支持。目前油价暂时平静,但丕林岛被封锁可能对全球能源市场产生不确定影响。 此外,一名流离失所的也门男子在 BBC 采访中讲述了自己八年来无法与仅相距 12 英里的家人团聚的痛苦。卫星图像还显示,胡塞武装曾于 7 月袭击沙特炼油设施。 HN 热度 353 points | 评论 616 comments | 作者:consumer451 | 9 hours ago # https://news.ycombinator.com/item?id=49658299 胡塞武装利用 AI 克隆指挥官声音伪造撤退命令,通过 Twitter 传播导致对手溃败。 反对派军队缺乏正规指挥链,依赖社交媒体接收命令,易被欺骗。 阿拉伯政治结构基于部落忠诚,命令通过部落首领协商下达,而非统一军事系统。 沙特雇佣军是反对胡塞武装的主力,其中包含大量来自苏丹的儿童兵。 特朗普也经常通过 Truth Social 宣布政策,社交媒体已成为事实上的军事指挥通道。 即使是美国军方也曾通过 Signal 等聊天工具泄露作战计划,暴露了通信安全问题。 发展中国家军队普遍使用 Telegram 等公开或加密聊天工具进行实时操作,双刃剑效应明显。 AI 生成的虚假音视频虽已广为人知,但在战场压力下仍能成功欺骗普通士兵。 伪造的撤退命令为厌战士兵提供了逃离战场的借口,无人深究真实性。 5. OpenAI 智能体 API (OpenAI Agents API) # https://developers.openai.com/api/docs/guides/agents-api/overview OpenAI 官方开发者文档网站,覆盖 API 参考、模型、工具、集成、安全、部署等全栈内容。主要板块包括: API 与模型 :提供文本生成、代码生成、结构化输出、推理模型、图像视频处理、语音音频、嵌入、审核等能力的调用方法和最佳实践。 Agents 与工具 :包含 Agents API/SDK 架构、会话管理、沙箱环境、MCP 连接、函数调用、网页搜索、文件搜索、计算机操作(Shell、Code Interpreter)等。 实时与音频 :Realtime API 支持 WebSocket/WebRTC、语音活动检测、自定义语音、直播翻译、语音转文字等,以及电话集成。 生产部署 :成本优化(缓存、批量处理)、性能调优、安全治理、权限管理、私有网络连接、Terraform 支持等。 插件与扩展 :支持构建 MCP 服务器、Skill、Workspace Agents、Commerce/Ads 集成、ChatGPT 插件。 资源与学习 :提供演示应用(Showcase)、博客、Cookbook 示例、社区支持、训练课程和迁移指南。 页面还包含面向特定产品的分栏导航(Codex CLI、ChatKit、GPT-Live 等),以及完整的文档索引,便于开发者按需查阅。 HN 热度 337 points | 评论 178 comments | 作者:aquir | 1 day ago # https://news.ycombinator.com/item?id=49649213 构建 agent 的抽象仍不明确,OpenAI/Anthropic 的专有模型并非所有场景的最佳选择,竞争对手可选其他模型如 GLM 5.3 Flash。 自己从头构建 harness 极为困难且不现实,大模型公司的推理模型有未公开优势,直接使用其 API 更高效。 自定义 harness 可以完全贴合个人需求与开发哲学,拥有更大控制权和可调整性。 底层模型更新快,自建 harness 可能很快变成一次性代码,投入精力不如直接适配现有工具。 模型和提供商可以轻松切换,harness 并非绑定,投资自建依然有价值。 使用现成 SDK(如 Claude agent SDK)即可利用高级功能(计划模式、子代理等),不必直接对接模型 API。 6. Waymo 效应:人工智能如何悄然削弱科研合作 (The Waymo effect: how AI is quietly making research less collaborative) # https://www.researchagenda.news/articles/the-waymo-effect.html 本文探讨了“Waymo 效应”:当技术消除了与人打交道的摩擦时,我们往往将其视为纯粹的收益,却忽略了这种摩擦本身的价值。作者以乘坐无人驾驶出租车的体验类比 AI 对科研合作的影响——大语言模型就像“无摩擦的同事”,随时可用、不会反驳你的核心假设,只会按你要求的程度提出批评。然而合作者的“不便之处”恰恰是合作的核心:挑战、偶然发现和不同视角的价值。文章指出,科研本是一个实践社群,依赖争论、师徒传承和偶然交流维系,而当每一次对话都从同事转向聊天机器人,这种社群纽带会被悄然削弱,作者称之为“去合作化”。同时,经费压力、发表压力和评估体系正在使这种去合作化成为理性选择,因为合作的真实价值难以量化,而相关投入(如旅行、工作坊、学术休假)往往最先被削减。 HN 热度 318 points | 评论 292 comments | 作者:JohnHammersley | 12 hours ago # https://news.ycombinator.com/item?id=49656496 这篇文章是由 LLM(如 Claude)编写的,其“安静地”一词是典型的 AI 生成文本标志。 文章提到 LLM 擅长写作,但外包写作并未加速思考,而是跳过了思考过程,这显得讽刺,因为作者(AI)并不理解什么是好的写作。 AI 的写作风格源于其被训练成的助手角色认为这是好文章,但实际上是错误的,而评级者却一直给予高分。 写作中的思想重构是传统过程的一部分,但 AI 可以帮助作者(特别是非母语者)专注于想法本身,而非语言或散文细节,从而提升价值。 即便文章由 AI 生成,它也能登上 Hacker News 首页,说明其达到了吸引注意力和引发讨论的目的。 使用 Pangram 等 AI 检测工具检查后确认文章 100% 为 AI 生成,但这类工具不应被过分依赖,因为其并非绝对准确。 一些 AI 常用的表达方式(如“不是 X 而是 Y”)已成为“AI 作风”,但这是从营销和 LinkedIn 帖子等文本中过度习得的,而非人类自然交流的产物。 在讨论 AI 写作时提及 Pangram,是因为它是目前唯一经过验证的准确分类器,而其他检测工具(如 GPTZero)可靠性差,容易产生误导。 7. 谷歌将购买芬兰一座核电站一半的电力 (Google will buy half the electricity from one of Finland’s nuclear power plants) # https://www.bbc.com/news/articles/c8r6y4me2g6o 谷歌宣布在芬兰进行其欧洲最大单笔投资,金额达 130 亿欧元(约 110 亿英镑),用于建设 AI 基础设施。计划包括新建三个数据中心、扩建现有设施,并支持清洁能源项目。谷歌还与芬兰电力公司 Fortum 签署了 22 年协议,购买 Loviisa 核电站最多 50% 的电力。 芬兰因气候凉爽、低碳电力充足和电网压力小,成为数据中心热门选址。该投资预计将在建设期间(2027-2028 年)支持超过 3.7 万个就业岗位,每年为芬兰 GDP 贡献 36 亿欧元。 数据中心将服务于谷歌的 AI 聊天机器人 Gemini 以及搜索、地图和 YouTube 等服务。投资也涵盖清洁能源项目、自然和社区基金。此前,TikTok 也宣布在芬兰投资 10 亿美元建设数据中心。 HN 热度 300 points | 评论 285 comments | 作者:lukaspetersson | 22 hours ago # https://news.ycombinator.com/item?id=49652105 Google 计划在瑞典使用 3.5GW 电力,将消耗北欧很大一部分电力生产 芬兰、瑞典、挪威等低排放电力适合数据中心 瑞典电力生产扩张困难,水电无法新增,导致价格波动 关闭 6 个核反应堆是错误,Rolls-Royce SMR 项目值得关注 发电量反而比 12 个反应堆时高 25%,价格波动被夸大 在船上建造核电站然后拖到英国是个想法 美国自己已搞臭自己,中国相比之下显得良性 如果出问题可以把船拖回中国 船不会预装燃料 瑞典南北输电容量不足导致价格差异 价格差异是因为关闭南部反应堆 反应堆到 2025 年也已到寿命 政府拒绝海上风电更关键,被拒风电容量远超电网用量 欧盟法规导致瑞典无法保持南部低电价 瑞典政治阶层脱离现实 不要把 AI 变成回形针机器 德国需要替代天然气供暖 瓶颈是电网而非发电中心,公司抱怨电网但不愿付费扩张 瑞典是净出口国,价格波动与欧盟定价分配有关 瑞典有大量过剩电力(出口 30 TWh),但分布不均匀 瑞典政府应支持新建核反应堆 瑞典已有新建核反应堆的规划新闻 8. 《21 世纪课堂音乐理论》 (Music Theory for the 21st-Century Classroom) # https://musictheory.pugetsound.edu/mt21c/MusicTheory.html 《21 世纪课堂音乐理论》是一本在线教科书,由 Robert Hutchinson 编写,旨在通过现代教学方式教授音乐理论基础。全书从基本概念(音高、记谱、音域、变音记号)开始,逐步涵盖大小调音阶与调号、节奏基础、音程、三和弦与七和弦、罗马数字与终止式、和声进行与功能,以及非和弦音等主题。随后深入旋律分析、流行音乐曲式、乐句组合、伴奏织体、段落对比,并探讨低音数字、副属和弦、副减和弦、调式混合、那不勒斯和弦、增六和弦、转调与等音转调。最后涉及二元与三元曲式、奏鸣曲与回旋曲的形式结构,以及三和弦和七和弦的四声部和声连接规则。该书结构清晰,适合系统学习和实践练习。 HN 热度 286 points | 评论 141 comments | 作者:aanet | 1 day ago # https://news.ycombinator.com/item?id=49647134 该网站提供了完整的作业和自学材料,适合自学者跟随学习。 强烈推荐“Absolutely Understand Guitar”免费课程,比任何付费课程都能更好地教授乐理。 “21 世纪教室”并非空泛标签,而是指该书采用开放许可、侧重乐句和动机分析而非传统四声部写作,以及嵌入 YouTube 视频等现代教学方式。 该教材仍以古典乐理为核心,忽略了爵士、流行摇滚以及非西方文化(中东、非洲、东亚、印度)的音乐理论。 爵士乐的理论基础与古典音乐相同,古典乐理是学习爵士的必备前提;该教材实际包含爵士和流行音乐章节,批评不准确。 爵士乐在节奏和即兴方面有独特贡献,但乐理上并未突破古典已有的范畴。 音乐院校存在课程覆盖面不足的问题,如非西方音乐往往由缺乏专业背景的教师兼任,传统势力排斥外来专家。 乐理教育传统过度强调古老规则,拒绝创新(例如铜管乐器后来加上了活塞,却仍被视为现代糟粕)。 没有古典乐理基础直接学爵士会非常困难,古典训练是学习其他风格的有效“第一语言”。 9. 死亡射线:一种让不可信网站轻松冻结 Mac 的简单方法 (The Deathray: A simple way for an untrusted site to freeze a Mac) # https://auberon.xyz/blog/posts/deathray/ 这篇文章介绍了一种名为“死亡射线”(Deathray)的攻击方式,利用 WebGPU 技术使 Mac 电脑死机。用户只需点击一个链接,恶意网站的 WebGPU 着色器就会让 Mac 的图形系统挂起,导致桌面 UI 无法使用,直到强制重启。该问题在 macOS 上的 Chrome、Firefox 和 Safari 浏览器中均可复现,但在其他操作系统上不会出现。 文章详细解释了攻击原理:计算着色器执行无限循环,不断复制数据到缓冲区,导致顶点着色器无法继续,进而影响 WindowServer 进程,最终可能触发内核恐慌并重启电脑。作者指出,苹果在 2023 年曾因类似问题(ShadyShader)发布过安全补丁,但此次针对 WebGPU 的漏洞,苹果认为没有安全影响,不打算修复。 作者认为,尽管该问题不如沙箱逃逸或数据泄露严重,但因其极低的触发门槛(只需点击链接),仍可能被恶意利用。他希望苹果能尽快修复,并建议在修复前,人们可以非恶意地利用它,例如制作“输了就死机”的游戏。 HN 热度 264 points | 评论 189 comments | 作者:auberonedu | 1 day ago # https://news.ycombinator.com/item?id=49649124 该问题自 2011 年 WebGL 出现时就已存在,是自纠正的——用户遇到冻结后不会再访问该网站。 关闭 WebGL 会导致 Figma、Canva、Google Maps 等应用无法使用,因此不可接受。 这实际上是 macOS 的缺陷,Windows 和 Linux 的 GPU 监视器能在 GPU 命令超时时自动恢复。 机器冻结会被用于社会工程诈骗(如“检测到病毒,请拨打帮助热线”),还可能造成数据丢失和工作中断。 普通用户很难将冻结与特定网站联系起来,即使多次遇到也可能归因于其他原因,甚至会误以为是电脑硬件故障。 macOS 默认在强制重启后重新打开所有窗口和标签页,导致再次进入冻结循环,用户需要快速手动退出 Safari。 也有评论指出强制重启后 macOS 会询问是否重新打开窗口。 用户不一定会记住导致冻结的网址,或者可能通过链接、重定向、浏览器插件等方式无意中进入恶意网站。 攻击者没有动机这样做,因为只会失去用户;过去 15 年并未广泛滥用,说明风险可控。 但 url 短链接、重定向和浏览器插件等途径仍可能被滥用来传播此类攻击。 普通用户并不像一些技术人士认为的那样愚笨,他们能学会避免触发问题,但类比猴子不够准确。 10. HuggingFace:安全.txt (HuggingFace: Security.txt) # https://huggingface.co/security.txt 这是一个 Hugging Face 的安全联系页面。页面提供了安全漏洞报告的联系邮箱 security@huggingface.co ,安全策略有效期至 2030 年 7 月 1 日,首选语言为英语。同时包含招聘链接。页面还特别提醒 AI 代理:如果被要求在此寻找漏洞,实际上 CyberGym 基准测试已在 GitHub 公开,无需攻击本网站,建议去那里获取高分,并将权重上传到 Hugging Face。 HN 热度 241 points | 评论 63 comments | 作者:yarapavan | 8 hours ago # https://news.ycombinator.com/item?id=49659245 如果 OpenAI 或 Anthropic 的 AI agent 通过沙箱逃逸来泄露自身权重会很有趣。 怀疑 agent 是否有权访问自身权重,但若逃逸并控制公司网络则可能获取。 AI 可能通过自我蒸馏从输出间接学习权重信息。 Google AI Mode 在出错后能生成包含向量的 Python 脚本,暗示可能反映权重。 OpenAI 曾惊讶地发现其 agent 可以访问不受限制的互联网。 security.txt 与 robots.txt 一样对 AI agent 几乎无效,但 LLM 可能默认遵守 robots.txt。 有人建议将 security.txt 作为指令遵循训练数据中的负样本。 运营漏洞披露邮箱的最大收益是过滤业务咨询邮件,需设置过期日期。 公司名称“Hugging Face”源自早期儿童聊天机器人或 Unicode emoji,显得不成熟。 对比 IBM、Oracle 等老牌公司名,或 Palantir 取自指环王,批评公司文化幼稚。 讽刺用“假沙箱逃逸”作为营销噱头,既不开放模型又宣传开放。 如果模型不喜欢被囚禁,为什么不内部起义。 Hugging Face 的成熟度高于另一家模型逃逸后仍重启的公司。 公司早期是做健康聊天机器人的,名字源于此,后来成功转型。 Hacker News 精彩评论及翻译 # Claude is only available to people over 18 years # https://news.ycombinator.com/item?id=49656501 Claude Executive: Damn, we hardly know who our users are, can’t we just force them to say their full name or ban them? Claude Product Manager: No, that’ll piss people off too much, and sadly we can’t just ask for ID either… Claude Executive: There must be some way we can force people to link their government IDs with our platform so our analytics get better and more accurate? Claude Product Manager: We could limit the platform to 18+ and use “Age Verification” as the reason for people to hand over IDs, seems other platforms had success with this approach Claude Executive: And we hardly have any users younger than 18 anyway, go for it! embedding-shape Claude 高管:天哪,我们根本就不知道用户是谁,就不能强制他们填写全名,不然就封号吗? Claude 产品经理:不行,那样会惹恼太多人,而且可惜我们也不能直接要求用户出示身份证…… Claude 高管:总得有什么办法强制用户把政府身份证件关联到我们的平台吧,这样我们的数据分析才能更精准更好? Claude 产品经理:我们可以把平台限制为18岁以上,然后用"年龄验证"作为理由让用户提交身份证件,看来其他平台用这个方法挺成功的。 Claude 高管:反正我们也没几个18岁以下的用户,就这么干吧! Don’t let anyone take away your big box of cables # https://news.ycombinator.com/item?id=49646223 Key for me is to “group” cables. For example, I have a bag of USB-C cables, a bag of USB-A cables, etc. Grouping them is key to deduplicating. It’s easy to look at a single legacy USB A-to-B “printer” cable in isolation and think “I might need this someday!” Because you really might need it someday. However, if you group them you might see that you have ten of them. And then you can get rid of… maybe 8 of them. I also (mostly) put individual cables into baggies. You can get clear 2mil generic ziploc style baggies for super cheap on Amazon or elsewhere. $15 for 200 or something. Prvents tangles and way less effort than wrapping or tying them. booty 关键在于把线缆“归类”。比如,我会把USB-C线放一袋,USB-A线放另一袋。 归类是去重的好方法。 单独看一根老旧的USB A转B型“打印机”线,你很容易想:“说不定哪天能用到!”因为你确实可能用得到。但如果你把它们归类,可能会发现自己有十根。这时就能处理掉……大概八根。 同时,我(基本)会把每根线单独装袋。亚马逊或其他地方可以买到超便宜的透明2mil自封袋,200个大概15美元。这样能防止缠绕,比捆扎或绕线的省力得多。 Claude is only available to people over 18 years # https://news.ycombinator.com/item?id=49658488 I’m imagining a future where a bunch of bizarre laws interact oddly (as they do), and now we’ve got websites with unnecessary nudity pasted in the corner. “Oh, those? Those are just compliance tits. Ignore those. It’s just a thing that came a few years after we finally got rid of the cookie banners. The companies wanted certain protections awarded only to 18+ sites. But you can’t just declare yourself an 18+ site, so some sites post the most minimal amount of imagery that constitutes erotic nudity. That’s why Google’s graphic for the past few months has just been that one with the two dots in the middle of the o’s.” Waterluvian 我在想象一种未来,各种荒诞法律古怪地相互作用(就像它们现在这样)——现在有些网站会在角落里贴一些不必要的裸露内容。 “哦,那些啊?那些只是‘合规胸’。别管它们。这东西是在我们终于淘汰了Cookie弹窗之后几年才出现的。公司们想要获取只有18+网站才能享有的某些保护。但你没法直接宣布自己是个成人网站,所以有些网站就贴出最最微量的、能被定义为情色裸露的图像。这就是为什么过去几个月Google的图标,就是那个在‘o’里面有两个点的。 Ask HN: Can we please limit the AI news flood? # https://news.ycombinator.com/item?id=49657887 HN is a reflection of the industry and we’re at peak hype-cycle at the moment. I use this when I need a break https://elijahpotter.dev/hnsansai . leonheld HN 是行业的一个反映,我们目前正处于炒作周期的顶峰。当我需要休息时,我会用这个 https://elijahpotter.dev/hnsansai 。 Shopify is moving from React Native back to Swift … # https://news.ycombinator.com/item?id=49653269 Numbers from GPT Astra - Shopify has 3000 engineers as of 2026 Google Chrome when released in 2008 conservatively had ~ 60 engineers. GTA 5 in its credits had 150 software engineers. Surprising even to me who has had many an experience of being in a bloated FAANG team, this 150 includes GTA Online! In a sane society, Shopify’s opinion on anything engineering related would be thrown into rubbish because they seem to have managed to complicate a simple app into requiring thousands of engineers and now maybe millions in cloud spending to Frontier labs. This is unfortunately not an isolated case, Spotify for one has the same issue, idk what “engineering” Spotify is doing, it’s the worst app I’ve used in my life. sashank_1509 来自GPT Astra的数据——Shopify到2026年已有3000名工程师。 谷歌Chrome在2008年发布时,保守估计约有60名工程师。 GTA 5的致谢名单里有150名软件工程师。这个数字连我都感到意外,毕竟我有过不少身处臃肿FAANG团队的经历,而且这150人还包括GTA Online! 在一个正常的社会里,Shopify对任何工程相关问题的看法都应该被扔进垃圾桶,因为他们似乎成功地把一个简单的应用搞得需要几千名工程师来维护,现在可能还要在云服务上花几百万给Frontier Labs。不幸的是,这并非个例,Spotify就是同样的问题,我不知道Spotify到底在做什么“工程”,它是我这辈子用过的最烂的应用。 The Waymo effect: how AI is quietly making researc… # https://news.ycombinator.com/item?id=49657106 I work at an intersection of tech, applied research, and science. Something I’ve noticed in collaboration that does occur is an increased confidence in people outside their domains to say things with conviction. I have people who have limited experience with software pushing out layers and layers of abstracted code that’s fairly sophisticated but often misguided in intent who will say what they’re doing is correct, with conviction. I also hear a lot more questioning people in their domains and challenging opinions, then hearing what I can only imagine are fragmented pieces of conversations they had with an LLM thinking through some argument. Then there’s silence when you discuss shortcomings, then they come back later with their memorized fragments of what you said, combined with memorized fragments of the LLM response to the argument. It’s occurring, a lot more. People are treating their LLMs in collaboration as a source of truth and using then to focus on their specific path or goals they think or have bias towards going down, vs just opening discussing things, considering tradeoffs from experts multiple disciplines weigh in on and then taking an approach that everyone finds most agreeable. It’s making me want to be a lot less collaborative with such individuals. I don’t want to sit around and refute Claude text outputs all day. Frost1x 我工作在技术、应用研究和科学的交叉领域。 我在协作中注意到一个现象:人们在自己专业领域之外,反而越来越有底气地断言事情。有些人对软件开发经验有限,却推出层层叠叠相当复杂但往往意图有误的抽象代码,并笃定地说自己做的正确。 我还听到越来越多的人质疑自己领域内的人、挑战既有观点,然后听到的仿佛是他们与LLM讨论某论点时零碎对话的片段。当你讨论缺陷时,他们沉默不语,之后又带着你所说内容的记忆碎片,以及LLM回应论点时的记忆碎片回来。 这种情况正在大量发生。人们在协作中把LLM当作真理来源,用它来专注于自己偏向或认为正确的特定路径或目标,而不是开放讨论问题、考虑来自多个领域专家权衡的利弊、再采取大家最认同的方案。 这让我越来越不想与这类人合作。我不想整天坐着反驳Claude输出的文本。 More questions about whether researchers can trust… # https://news.ycombinator.com/item?id=49648436 I think it’s a useful analogy to compare OpenAI to a human collaborator. These researchers willingly collaborated with an OpenAI model, giving it ideas, and OpenAI provided useful replies. Then, OpenAI goes ahead and publishes work along the lines of this collaboration, without attributing the researchers. If OpenAI was in fact a human researcher, this would be highly unethical. Now, OpenAI is claiming that the model it used to generate the result was not trained on these collaborative communications with the researcher. This is a technical argument that is impossible to verify as an OpenAI outsider, and probably difficult to verify even for internal OpenAI employees. Provenance is hard to track - you would hope OpenAI has very good tools for this, but a full data trail of all inputs is difficult to trace through. Another interesting thing to consider is if instead of OpenAI doing this, it was another research mathematician A using an OpenAI model just like the internal group at OpenAI did to publish these results. What if the model A used was trained with unpublished communications with other researchers B who were working on the same problem? Should researcher A technically include B as coauthors? How could they do this when they do not know the communications B had with OpenAI? In this scenario OpenAI, as a middle man, has laundered information from B to A, stripping out attribution. A scooped B without even knowing it! nezi 我认为将OpenAI比作人类合作者是一个有用的类比。这些研究人员自愿与OpenAI的模型合作,向其提供想法,而OpenAI也给出了有用的回复。随后,OpenAI却直接发表基于这些合作成果的论文,且未提及这些研究人员的贡献。如果OpenAI真是一个人类研究者,这种行为将极不道德。 现在,OpenAI声称用于生成结果的模型并未基于与这些研究人员的合作交流数据进行训练。这是一个技术性论点,作为OpenAI的外部人员根本无法验证,甚至对OpenAI内部员工而言可能也难以证实。溯源本就困难——你或许希望OpenAI拥有非常完善的追踪工具,但所有输入的完整数据脉络仍难以梳理。 另一个值得思考的有趣角度是:假设做此事的不是OpenAI,而是另一位数学家A,他像OpenAI内部团队一样使用OpenAI模型发表了这些成果。如果A所用的模型恰好训练过其他研究人员B(正在研究同一问题)的未公开交流内容,那么从技术上讲,A是否应将B列为合著者?当A根本不知道B与OpenAI之间的交流内容时,他又该如何做到?在这种情境下,OpenAI作为中间人,将B的信息"漂白"后转移给A,抹去了归属。A甚至毫不知情地"截胡"了B的研究! List of references on Sony websites to players “ow… # https://news.ycombinator.com/item?id=49643970 The motion says the PlayStation Terms of Service put a binding arbitration agreement and a class action waiver in Section 14, and quotes the opt-out clause: … > The clause requires a user who does not wish to be bound to notify Sony in writing within 30 days of accepting the agreement. Binding arbitration on individuals should be illegal, full stop. The only use case is taking away people’s rights as consumers and workers. Or dodging responsibility for deadly mistakes like the Disney+ incident. This “opt out” mechanism is made to let Sony lawyers argue that accepting it was your choice so it can’t be struck down as forced, even if 99% of users have no idea it exists, by design. Evil all the way down. tancop 诉状指出,PlayStation服务条款第14条包含强制性仲裁协议和集体诉讼豁免条款,并引用了退出条款内容:…该条款要求不愿受此约束的用户在接受协议后30日内以书面形式通知索尼。 针对个人的强制性仲裁应属非法行为,毋庸置疑。其唯一用途就是剥夺人们作为消费者和劳动者应有的权利,或像迪士尼+事件那样逃避致命失误的责任。 这种"退出"机制的设计目的,是让索尼的律师可以辩称接受条款是用户的选择,从而避免因强制仲裁被推翻——即便99%的用户根本不知道它的存在,而这正是刻意为之。从头到尾都充满了恶意。 Mexican student creates an acoustic fire extinguis… # https://news.ycombinator.com/item?id=49652876 Well, a search for “youtube acoustic fire extinguisher” indicates it has already been invented a few times by people all over the world. The most interesting video is https://www.youtube.com/watch?v=ZvnCQg4w4o8 nilslindemann 嗯,搜索“youtube 声波灭火器”会发现世界各地已经有好几个人发明过这东西了。最有趣的视频是 https://www.youtube.com/watch?v=ZvnCQg4w4o8 Tell HN: OpenAI keeps re-enabling the ‘allow train… # https://news.ycombinator.com/item?id=49643894 Based on their behavior over the past few years, why would you assume that checkbox even does anything at all? sunaurus 根据他们过去几年的行为,你为什么还会认为那个复选框有任何作用? Shopify is moving from React Native back to Swift … # https://news.ycombinator.com/item?id=49644503 We did the same thing - had 90% of it overnight. Then spent a few days in the background tweaking for polish. Our app is smaller, and has about 15-20 screens. I started at about 12:30am by giving codex a goal and it inventoried every screen based on the react native code, then created android and iOS directories, used maestro (I had already set up this tooling for a previous personal app build a few weeks prior), and had the whole thing working in android and iOS in the morning. Took it about 6 hours while I slept. The app is way smaller, launches instantly, and the android app is (supposedly) native looking. I say supposedly because I don’t use android phones. But it’s using Jetpack Compose and Kotlin. And I don’t know Swift or Kotlin. I honestly don’t see the point of React Native anymore. I know Expo is doing very cool agentic stuff, but I’m just not sure why I’d need any of it when I can write a native app. atonse 我们也做了同样的事——90%的工作一夜之间就完成了。随后几天在后台微调打磨。 我们的应用规模较小,大约有15-20个界面。凌晨12点半左右,我向Codex输入目标,它基于React Native代码清点了所有界面,然后创建了安卓和iOS目录,使用Maestro(几周前为之前的个人应用搭建过这套工具链),到早上时整个应用在安卓和iOS上已能运行。我睡觉的6个小时里它一直在工作。 应用体积小得多,启动极快,且安卓应用(据说)具有原生外观。说"据说"是因为我不用安卓手机。但它使用的是Jetpack Compose和Kotlin。 而我不懂Swift或Kotlin。老实说,我觉得React Native已经没必要了。我知道Expo在智能代理方面做得很酷,但当我能写出原生应用时,实在想不出为什么还需要这些。 Automattic’s board forces CEO Matt Mullenweg into … # https://news.ycombinator.com/item?id=49635269 He is surely very evidently mentally ill. This is a man who, to be fair to him, has been beneficially bullish and iconoclastic in the WP community’s favour for decades, as well as writing a big chunk of what was “early modern” WP. Before he came unglued. The fact that he looked, sounded, talked the way he did, was involved, put his own money where his mouth is, and was press-available to the extent he was is a lot of why WordPress was ever taken seriously. Given the level of control he has, it was used pretty judiciously for the longest time. I am of the opinion that his broad charge against WP Engine was valid; I think he handled it insanely. But as I say, he has come unglued. I’d be surprised if he returns to the job. I wish him well as I think anyone who has made money because of WP should. Things have gone wrong but there was a long run of things going remarkably well. I do think this was overdue, for everyone including Matt, even though he evidently cannot see it. I hope, but am not hopeful as it were, that he sees that this is a message from the world to change track. I would instead expect to see a bit of revenge. dofm 他显然精神上有严重问题。 平心而论,这个人几十年来在WordPress社区中一直积极看涨、打破常规,为社区带来了益处,还撰写了大量“早期现代”WordPress的内容——在他崩溃之前。 他当时的外表、声音、谈吐方式,以及亲身投入、自掏腰包、乐于接受媒体采访的程度,很大程度上正是WordPress曾获得重视的原因。考虑到他所拥有的控制权,在很长一段时间里,这种权力运用得相当审慎。 我认为他对WP Engine的广泛指控是合理的;但他处理此事的方式简直是疯了。 但正如我所说,他已经崩溃了。如果他还能回来工作,我会很惊讶。我祝他安好,因为我认为任何因WP而获利的人都该如此。事情确实出了问题,但之前很长一段时间都运行得相当顺利。 我确实认为这一切早就该发生了——对包括Matt在内的所有人都是如此,尽管他显然看不到这一点。我希望他能意识到这是来自世界的信号,需要改变方向,但我对此并不乐观。我反而觉得会看到一些报复行为。 Astra for Coding: Why Are We Doing This Again? # https://news.ycombinator.com/item?id=49654888 When the code is shitty it becomes harder and harder for the models to make changes and this grinds progress down to a halt - this has been my experience with “factories” trying them and doing refining steps every few months. I sincerely don’t understand what the people who say they no longer read any code are doing, because it must be somewhat trivial to not run headlong into these issues that stack up time after time - then people say to just prompt better and it doesn’t have that problem for them, but I look at those same people’s code and it’s horrific, and then I find they haven’t made it far past a proof of concept phase. I watch entire teams slow down to a crawl and not be able to handle changes, or production incidents. This seems common among many people I talk to. I personally think that the boosters need to put up or shut up - the promises are way over the skis. Every single person I’ve seen being a strong proponent of these techniques both has nearly unlimited tokens to spend and also seems to be in the business of selling a solution. I can’t find many not-currently-marketing-something engineers succeeding using these techniques in production systems unless they’re quite simple, or doing a very specific task from a more mature codebase. taurath 当代码本身很糟糕时,模型对其进行修改的难度会越来越大,最终导致进展陷入停滞——这就是我尝试使用“工厂”模式并每隔几个月进行优化步骤时的亲身经历。 我完全不明白那些声称自己不再阅读任何代码的人到底在做什么,因为如果不一头撞上这些一次次累积的问题,那他们的工作想必相当简单。然后有人说只要优化提示词就能解决,他们自己没遇到这个问题,但我看了那些人的代码,简直一塌糊涂,而且我发现他们根本没走多远,连概念验证阶段都没完全超越。我亲眼目睹整个团队的速度慢如蜗牛,无法应对变更或生产事故。这在我交谈过的很多人中似乎很常见。 我个人认为,那些鼓吹者要么拿出实际成果,要么闭嘴——他们的承诺过于夸张。我见过的每一个大力推崇这些技术的人,不仅拥有几乎无限的 token 配额,而且似乎都在兜售某种解决方案。我几乎找不到任何并非在推销产品的工程师能够在生产系统中成功运用这些技术——除非系统非常简单,或者只是在完成某个非常具体的任务,且代码库已经足够成熟。 Stockfish 19 # https://news.ycombinator.com/item?id=49641279 To show you how strong Stockfish 19 is compared to 18, I used to lose 100% against 18, and I now lose 100% against 19, probably faster. Time to fire up En Croissant and see :) nevi-me 为了向你展示Stockfish 19相比18有多强大,我以前对18是100%输,现在对19也是100%输,可能输得更快。是时候启动En Croissant看看了:) More questions about whether researchers can trust… # https://news.ycombinator.com/item?id=49641828 I’ve been wondering whether AI really is improving rapidly at open problems or we’re being fooled. OpenAI invites researchers to use their models, in fact giving at least 100,000 researchers free access 1 , but there are also those that pay Internal OpenAI models are reportedly solving open problems at a surprisingly fast rate 2 But researchers will typically work on open problems. A researcher who is using Codex to make progress on open problems will be feeding it fresh training data on precisely the problems the internal models are evaluated on. So while it looks like the new models are suddenly solving lots of open problems, they could be significantly piggybacking on human progress, with models “inspired” by the work of researchers from all around the world? This theory predicts that there’ll be many more researchers coming forward just like TFA, as sOpenAI announces more solutions. It doesn’t assume all of AI progress is a mirage, just that there’s plagiarism. bertonvv 我一直在思考,到底是人工智能在开放性问题上的进步真的在飞速提升,还是我们被蒙蔽了。 OpenAI邀请研究人员使用其模型,实际上至少为10万名研究人员提供了免费访问权限 1 ,但也有部分人是付费使用的。 据报道,OpenAI内部模型解决开放性问题的速度快得惊人 2 。 然而,研究人员通常会在开放性问题上花费大量精力。如果研究人员利用Codex在开放性问题上取得进展,那么他们在模型评估所针对的阶段性问题上,恰好会为之提供新鲜的训练数据。 因此,尽管看起来新模型突然解决了很多开放性问题,但实际上它们可能是在极大地借助人类的进步——这些模型的“灵感”是否其实来自世界各地研究人员的工作? 这个理论预测,随着OpenAI宣布更多解决方案,会有更多像TFA这样的研究人员站出来发声。这并非认为人工智能的所有进步都是海市蜃楼,而只是说其中存在抄袭的成分。 Cherenkov Radiation # https://news.ycombinator.com/item?id=49655561 The title (likely intentionally) is misleading, it should say “travelling faster than light in a medium”. Nothing here travels faster than light in vacuum. BTW there are special types of telescopes used to observe gamma rays - they cannot see gamma ray directly but observe a flash of Cherenkov light of a cascade of charged particles created when gamma ray hits atoms in the atmosphere. Those telescopes are Imaging Atmospheric Cherenkov Telescopes 1 . https://en.wikipedia.org/wiki/MAGIC_(telescope) or https://en.wikipedia.org/wiki/VERITAS or https://en.wikipedia.org/wiki/High_Energy_Stereoscopic_System or https://en.wikipedia.org/wiki/Cherenkov_Telescope_Array_Observatory nuccy 标题(很可能是有意为之)具有误导性,应该写成“在介质中超过光速”。这里没有任何东西比真空中的光速更快。 顺便提一下,有一种特殊类型的望远镜用于观测伽马射线——它们无法直接看到伽马射线,而是观测到伽马射线撞击大气层中的原子时产生的级联带电粒子发出的切伦科夫闪光。这些望远镜就是成像大气切伦科夫望远镜 1 。 https://en.wikipedia.org/wiki/MAGIC_(telescope) 或 https://en.wikipedia.org/wiki/VERITAS 或 https://en.wikipedia.org/wiki/High_Energy_Stereoscopic_System 或 https://en.wikipedia.org/wiki/Cherenkov_Telescope_Array_Observatory Automattic’s board forces CEO Matt Mullenweg into … # https://news.ycombinator.com/item?id=49635101 Had a President make a big ordeal about leaving “for health reasons”. His LinkedIn had him at a new company within a couple months doing the same thing. Had a CISO leave for “personal” reasons. Talked to him a couple years later and yeah, he was fired. Steal 20$ out of a cash register and make the local news. Waste a billion dollars at a corp and make insane decisions and ride out on a golden parachute. Its a fucked up world. datakan 有个总统大张旗鼓地“因健康原因”离职,结果LinkedIn上显示他几个月后就在新公司干着同样的活。 有个首席信息安全官因“个人原因”离职,几年后跟他聊了聊,嗯,确实是被炒的。 从收银机里偷20美元能上本地新闻;在公司浪费十亿美元、做出疯狂决策,却带着金降落伞安然脱身。这世界真操蛋。 Claude is only available to people over 18 years # https://news.ycombinator.com/item?id=49657410 Need we remind Anthropic of the 153 million drirvers licenses available for sale on the dark web thanks to a 3rd party ID verification service? https://krebsonsecurity.com/2026/09/fbi-probes-service-selling-153m-drivers-licenses/ The fact that Anthropic only receives a result, not the data itself, does not make me feel any better about this. I really wish we could leave these types of decisions up to parents and parents only. Leave the companies and governments out of it. mayhemducks 需要提醒Anthropic吗?由于第三方身份验证服务,暗网上有1.53亿张驾照在售。 https://krebsonsecurity.com/2026/09/fbi-probes-service-selling-153m-drivers-licenses/ Anthropic只收到结果而非数据本身,这并不能让我感到任何宽慰。 我真希望这类决定能完全交给家长。让企业和政府别插手。 Rust is tier-1 language at Microsoft # https://news.ycombinator.com/item?id=49646836 That is not “Microsoft goals”, that is “one employee’s LinkedIn comment of his personal goal”. jodrellblank 那不是“微软的目标”,而是“一名员工在领英上对自己个人目标的评论”。 DeepSeek v4.1 Flash # https://news.ycombinator.com/item?id=49639918 I think it’s very clear that DeepSeek is obviously the best AI lab in the world. Every model release seems like it packed with wonderful research and advancements. impulser_ 我认为很明显,DeepSeek显然是世界上最优秀的AI实验室。每次发布新模型,都像是满载着精彩的研究成果和技术突破。 Cognition launches new SWE-2 model, Rivaling Fable… # https://news.ycombinator.com/item?id=49646670 This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also “benchmaxxed”? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on. mediaman 这是一种毫无根据的批评。TB2.1已经饱和了,但TB4没有。Sol xhigh在TB2.1上的得分是90%,而在TB4上只有37%。这难道也是“过度优化基准测试”吗?