AI模型精选73°

Anthropic发布Claude Fable 5.1和Mythos 5.1

2026 09 03 HackerNews

精选理由

Anthropic新模型降价25%,Fable 5.1编码能力提升,Mythos 5.1专注网络安全,还新增企业安全功能。

AI 摘要

Anthropic推出Claude Fable 5.1和Mythos 5.1,价格降低约25%。Fable 5.1面向公众,Mythos 5.1专用于网络安全和生命科学。新模型在编码、知识工作和长期推理任务上显著超越前代,并推出企业前沿安全措施(EFS)。

原文 · SuperTechFans

2026 09 03 HackerNews

2026-09-03 Hacker News Top Stories # Anthropic 发布 Claude Fable 5.1 和 Mythos 5.1,性能提升且价格降低约25%,并推出企业级安全措施。 作者呼吁继续支持 Firefox,认为它是打破 Chrome 浏览器引擎垄断的唯一希望。 检查发现 AI 怀疑论者 Ed Zitron 的预测(如 Meta、Google、微软“正在死亡”)被后续营收和利润数据证伪。 Google 发布 Gemini 3.8 Flash 和 3.8 Flash Cyber,前者在编码和代理任务上性能大幅提升,后者专用于网络安全。 LWN.net 宣布将于2026年9月上调订阅价格约20%,以应对通胀和成本上升,订阅增长停滞是主因。 OpenAI Codex 桌面应用(后更名为 ChatGPT)在缓存中捆绑了完整的 LibreOffice 副本,占用1.7GB空间。 FBI 调查暗网服务 Nexus,该服务出售超过1.53亿张美加驾照扫描件,数据源自身份验证公司持续泄露。 Commodore 64 于1982年9月1日发布,成为史上最畅销电脑,作者通过它学习编程并进入 IT 行业。 LISEP 的真实失业率指标显示2026年7月美国功能性失业率为24.9%,远高于官方公布的4.1%。 1. Anthropic 发布 Claude Fable 5.1 与 Claude Mythos 5.1 (Claude Fable 5.1 and Claude Mythos 5.1) # https://www.anthropic.com/claude-fable-and-mythos-5-1 Anthropic 发布了 Claude Fable 5.1 和 Claude Mythos 5.1,这是目前最先进的编码和知识工作模型。Fable 5.1 面向公众,Mythos 5.1 仅通过可信访问计划提供,专为网络安全和生命科学设计。 主要更新包括:价格降低约 25%(缓存读取更便宜);推出企业前沿安全措施(EFS),提供完全隐私保护;改进了安全机制,减少误报,并允许 Fable 5.1 发现软件漏洞(但不开发利用)。在性能上,Fable 5.1 在编码、知识工作和长期推理任务上显著超越前代,且成本更低。早期合作伙伴(如 Jane Street、Devin)反馈其代码质量更高、更易追踪。 HN 热度 1374 points | 评论 1330 comments | 作者:denysvitali | 1 day ago # https://news.ycombinator.com/item?id=49525378 Fable 5.1 在写作风格上有显著改进,听起来不再像其他 Claude 模型那样刻板,风格更自然,对风格指令的响应更可靠。 Opus 的散文风格令人厌倦,部分原因是模型在为自己和彼此写作,而非为人类,它们将大量信号压缩进更少的词中,不在乎听起来是否尴尬。 AI 生成的文本既密集又空洞,充满自创术语,实际上没表达多少内容,像一种企业废话方言。 长期会话中 AI 输出的语言越来越难以理解,人类需要不断要求 AI 用简单英语解释,这使人类在长任务链中难以介入。 这种“AI 语言”可能成为未来人类需要学习的“外语”,未来世代的大脑可能会被重新连接以理解它。 Opus 5 的糟糕写作风格并非有意设计,而是 Anthropic 优化编码等其他特性的副作用,本可通过更好的判断力大幅减少废话而不丢失信号。 2. 紧紧抓住你的 Firefox (Hang on to Your Firefox) # https://www.newsonaut.com/articles/hang-on-to-your-firefox 这是一篇博客文章,作者 Mark Rogers 呼吁用户继续支持 Firefox 浏览器。 文章指出,尽管 Firefox 因入驻 X 平台(原 Twitter)而遭到一些批评,但它是目前唯一能挑战 Google Chrome 浏览器引擎垄断地位的希望。作者认为,Firefox 入驻 X 是为了吸引新用户,而用户应该帮助它,而不是因为小问题就放弃它。如果没有 Firefox,网络将完全被 Chrome 及其衍生品(如 Vivaldi)和苹果 Safari 主导。 HN 热度 933 points | 评论 509 comments | 作者:speckx | 1 day ago # https://news.ycombinator.com/item?id=49527748 Firefox 是浏览器引擎多样性和竞争的最后希望,但 Mozilla 收购广告技术公司、收集数据、推送个性化广告等行为让关心它的高级用户感到沮丧,他们希望 Mozilla 优先考虑正确的事情 有些用户用懒惰和牵强的理由留在 Chrome 生态,而不是真正尝试 Firefox Firefox 的开发者工具在某些方面优于 Chrome,例如响应式测试时可以关闭底部工具栏 树形标签(Tree style tabs)是 Firefox 独有的扩展功能,用户无法回到没有它的浏览器 人们批评 Firefox 是因为在乎,而 Edge 等浏览器很少被批评是因为没人用 有些人批评 Firefox 是为了给自己的道德标准找借口,以证明使用 Chromium 浏览器合理 Firefox 性能多年来一直很好,兼容性问题很少,但有些网站使用 Chrome 专属 API 使用 Chromium 浏览器会赋予 Google 更多权力,推动 Manifest V2 等,影响开放网络 人们会对自己撒谎来支持个人意识形态,认为产品 A 和 B 都不好,所以没必要换 3. Ed Zitron 的 AI 怀疑论预测有多准确? (How accurate have Ed Zitron’s AI skeptic predictions been?) # https://danluu.com/zitron/ 作者检查了知名 AI 怀疑论者 Ed Zitron 的预测准确性。以 2024 年 Zitron 称 Meta、Google、Microsoft“正在死亡”的预测为例,作者用这些公司实际的营收和利润数据反驳:Meta 2024 年营收 1650 亿美元(同比增 22%),利润 690 亿美元;Google(Alphabet)2024 年营收 3500 亿美元(增 14%),利润 1120 亿美元;微软 2024 年营收 2620 亿美元(增 15%),利润 1180 亿美元——均显示强劲增长,而非“垂死挣扎”。作者指出 Zitron 的推理依赖不可靠的第三方数据(如 Similarweb)和片面指责个别高管(如 Google 的 Raghavan),忽略了公司内部长期存在的商业决策逻辑。人们引用 Zitron 时往往将其当作“用数字说话”的权威,但实际上他的预测多被后续数据证伪。 HN 热度 839 points | 评论 994 comments | 作者:jatins | 1 day ago # https://news.ycombinator.com/item?id=49526069 对 OpenAI 和 Anthropic 的收入增长持怀疑态度,认为难以覆盖投入,重度用户会转向更便宜的开源模型 很多企业 AI 使用是管理层跟风炒作,并不真正理解其价值 数据中心过度建设可能引发金融危机,因为大量金融系统资金被卷入 OpenAI 和 Anthropic 不一定倒闭,但可能被其他公司收购,难以成为下一代巨头 LLM 会商品化,价格下降,最终本地运行,当前万亿级投资远超实际需求 批评者并非全盘否定 AI,而是质疑当前投资规模基于不切实际的收入预期,以及大厂用会计手段制造虚假繁荣 埃德·齐特龙并非温和中间派,而是彻底的悲观论者 悲观论调有时是对的,不能因为中间立场就认为更合理 齐特龙多次被事实打脸,其预测反复出错 即使齐特龙很多判断错误,悲观的大方向仍可能准确 有钱人总能操纵叙事,历史可能不会站在批评者一边 齐特龙的核心论点是“AI 不起作用”,或“愤怒但不懂行” 齐特龙承认 AI 在编程等任务上有用,但也承认它会产生很多 bug 需要基于他实际说过的话来评估,而不是听别人转述 齐特龙的论点核心是:LLM 能力被夸大、企业领导跟风或怕股价下跌、行业投入永远无法回本 齐特龙和奥特曼都是骗子,都不该作为可靠信息来源 齐特龙过去说“没人会为 AI 付费”,后来被证明错误后悄悄改口,从不承认错误 人都会随着时间改变观点,不能因为调整立场就否定他 齐特龙已承认之前低估了投资持续性和 AI 行业规模,现在认为 AI 是十亿美元级而非万亿美元级行业 4. Gemini 3.8 Flash 和 3.8 Flash Cyber (Gemini 3.8 Flash and 3.8 Flash Cyber) # https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/ Google 发布 Gemini 3.8 Flash 和 3.8 Flash Cyber 两款新模型。3.8 Flash 是推理与编码模型,性能大幅提升,尤其在长周期软件工程和自主代理任务上接近更高成本的前沿模型,价格与 3.7 Flash 相同(输入每百万 token 0.75 美元,输出 3.75 美元)。3.8 Flash Cyber 是网络安全专用模型,具备前沿的漏洞检测和自动修补能力,通过 Fairwind 计划提供给受信任的防御者。两款模型共享同一基础智能,并在网络安全领域进行了强化训练。文章还展示了 3.8 Flash 在游戏构建、DOS 版地图、地形可视化等演示中的能力。 HN 热度 788 points | 评论 467 comments | 作者:bratao | 9 hours ago # https://news.ycombinator.com/item?id=49537553 Gemini 3.8 Flash 速度快而且很擅长生成 HTML/JavaScript。 生成的代码虽然声称 60 帧,但实际运行可能卡顿,而且 60 帧是硬编码在 HTML 里的。 有评论用 AI 制造“欺瞒”事件类比,指出不能信任 AI 输出,必须逐行检查代码。 对 AI 产出不信任的问题,有人认为可以用另一个 AI 代理去检查,这样可以提高可靠性。 但接着有人质疑,难道要无限叠加代理去检查?如果模型本身不可靠,这种冗余何时是尽头? 也有人指出人类本身也不可靠,软件工程的本质就是在处理非确定性输出。 另一位认为,用多个代理检查会带来额外成本、延迟,但冗余确实是提高系统可靠性的核心方法。 还有人提到,信任程度视后果严重性而定,就像对待另一个人类开发者:比如小功能仅扫一眼就行,复杂关键流程则会仔细审核。 有观点认为如果每个代理能降低 90% 错误率,多个代理叠加会产生“九个九”的可靠性。 另外,同一个模型只要换不同提示词,就能指出自己之前输出的问题,不一定需要独立代理。 有工程团队反映,过去 6-12 个月大量采用 Claude 等 AI 工具后,代码变得脆弱、技术债积累严重,虽然团队声称净收益为正,但“100 倍效率提升”并不现实。 应对 AI 代码可靠性的关键,是建立强大的集成测试/端到端测试,用金标准断言数据和数据库快照进行回归检查。 高质量模型在辨别真正的回归和过时的测试断言方面表现非常好,能从功能实现追溯到业务规则。 有人指出该演示有声音开关的 bug,不令人满意。 还有人认为 Google 专注于速度和暂时接受智能排第三/四的策略,可能让它意外获胜。 5. 来自 LWN 关于订阅价格的一则说明 (A note on subscription prices from LWN) # https://lwn.net/Articles/1090585/ LWN.net 宣布将于 2026 年 9 月 15 日起上调订阅价格,涨幅约 20%,以应对过去五年间约 20% 的消费者价格通胀以及健康保险等成本的大幅上升。这是自 2002 年采用订阅模式以来的第三次涨价,上一次在 2022 年初。新价格分别为:饥饿黑客 6 美元/月、专业黑客 11 美元/月、项目领导者 19 美元/月、狂热支持者 55 美元/月,团体订阅同比例上涨。现有订阅在到期前保持有效,公告前已激活的月度订阅将在接下来六个月按旧费率收费。LWN 感谢读者长期支持,并指出订阅增长停滞是涨价的主要原因,同时承诺将继续提升内容深度和网站功能。 HN 热度 667 points | 评论 133 comments | 作者:rwky | 11 hours ago # https://news.ycombinator.com/item?id=49535752 LWN 是高质量的技术出版物,用户资助模式避免了广告依赖,是其高质量的原因之一。 LWN 是 HN 上经常被提交的网站,但需要找到新方式增加订阅者或收入,避免走广告路线。 有用户长期阅读 LWN 却不知道有订阅选项,说明订阅信息可能不够显眼。 LWN 对新文章有一周锁定期,订阅者可以分享链接,但有人觉得一周延迟不足以激励理性订阅。 订阅 LWN 更多是出于赞助心态,支持社区服务,即使不经常阅读也愿意付费。 理性思考应包含长期视角,如果从 LWN 获得价值,应该订阅以支持其持续存在。 经济学家可能理解搭便车问题,但他们的政策建议常被批评为过于简化激励问题。 订阅 LWN 是因为其慷慨和直率的风格,以及对读者的尊重。 即使不常用 Linux,也认同 LWN 的价值,并庆幸其没有走向商业化“奇妙旅程”。 多年订阅者认为 LWN 是职业生涯的基石,每周阅读高质量技术文章对成长至关重要。 建议对 LWN 进行大额捐赠,以回报其带来的巨大价值,有用户表示会照做。 6. ChatGPT/Codex 应用程序捆绑了完整的 LibreOffice 副本 (The ChatGPT/Codex app bundles a full copy of LibreOffice) # https://simonwillison.net/2026/Sep/1/codex-libreoffice/ Simon Willison 在 2026 年 9 月 1 日发布了一篇笔记。他使用 OmniDiskSweeper 检查 ~/.cache/ 文件夹时,发现 OpenAI Codex 桌面应用(后更名为 ChatGPT)占用了 1.7GB 空间。该文件夹名为 codex-primary-runtime,内含完整的 Python 和 Node.js 安装,以及 Poppler、git 和 LibreOffice 等原生二进制文件。此外,plugins/documents 子文件夹中的技能文件指导 Codex 如何查找和使用这些二进制工具。 HN 热度 480 points | 评论 235 comments | 作者:timpera | 1 day ago # https://news.ycombinator.com/item?id=49527396 希望 OpenAI 捐赠给 LibreOffice,以改善其 MS Office 支持和差异比较功能,实现双赢。 认为 OpenAI 不会捐赠,因为 Sam Altman 没有这种慷慨基因。 讽刺 Sam Altman 的“开放”只在他有利时才开放。 使用他人劳动而不回报是 AI 公司的核心模式,期待捐赠太天真。 法律上无需捐赠,但道义上应该做;如果不满,可使用更严格的许可证。 公司不在乎自由软件许可证,它们会无视法律。 捐赠可以用于支付开发者,The Document Foundation 已用捐赠雇佣了多位开发者,每年约 100 万欧元。 捐赠不仅用于营销,也能支付开发者,弥补志愿者与全职员工之间的差距。 人们常把合法的事与对的事混为一谈;法律上合规不代表道义上正确。 为了确保互惠,应使用要求回报的许可证(如 GPL),而 MIT 类许可证容易被公司白嫖。 更严格的许可证会减少采用率,但能保证代码共享和回馈;采用率并非开源的原意。 7. FBI 调查出售逾 1.53 亿张驾照的服务 (FBI Probes Service Selling 153M+ Drivers Licenses) # https://krebsonsecurity.com/2026/09/fbi-probes-service-selling-153m-drivers-licenses/ 一个名为 Nexus 的暗网身份盗窃服务正在出售超过 1.53 亿张美国和加拿大驾照的数字扫描件,以及超过 1000 万张身份证、300 万份旅行证件和国际 ID、57.9 万张医疗卡。该服务声称数据来自一家主要身份验证公司的持续泄露,且在过去 24 小时内新增了近 40 万条驾照记录。FBI 新奥尔良办事处已对此展开正式调查。 文章通过分析时间戳发现,这些扫描件可能与租车、机场安检或大麻药房等场景有关。例如,作者本人的驾照扫描件时间戳对应其 2025 年 6 月乘坐航班并租车的时间;其母亲的驾照时间戳与作者仅差几秒,表明两人同时向赫兹租车公司出示了驾照。另一名研究者的驾照时间戳对应其参加 DEFCON 安全会议期间在拉斯维加斯大麻药房 Planet13 出示 ID 的时间。Planet13 曾与路易斯安那州的身份验证公司 idscan.net 合作,后者为全美 1000 多家大麻药房处理身份验证。 该服务还包含国防部长 Pete Hegseth 等高级政府官员的驾照记录,部分记录标注来源为“CDL”(商业驾照)或“CAC”(通用访问卡)。数据持续更新,表明窃取活动仍在进行。 HN 热度 376 points | 评论 254 comments | 作者:tatersolid | 1 day ago # https://news.ycombinator.com/item?id=49529621 美国本应在 REAL ID 中植入 RSA 密钥对,用于在线身份验证,避免第三方扫描驾照。 可借鉴 PIV/CAC 卡模式,建立联邦与州的 PKI 互信体系,实现加密验证身份。 爱沙尼亚的 ID 卡已实现类似功能,插入 USB 读卡器即可验证身份。 美国公众因历史遗留的边疆文化和反共恐惧,抵触这类系统,却更接受企业建立的类似系统。 保守派曾视此为自由问题,后来转向支持全球威权主义。 支持数字身份的人认为应改善就业环境和减少汽车依赖,以降低验证门槛。 选民 ID 法的批评点在于 DMV 效率低下,若由 USPS 等机构办理则无争议。 DMV 实际上运行效率很高,只是用户体验差(等待时间长、邮寄慢)。 有用户亲身体验 Oregon DMV 预约后 10 分钟完成,ID 一周内寄到。 DMV 效率因州和地理位置而异,农村地区通常更快。 密西西比州有自动化 kiosk,10 分钟即可完成续期。 政府效率被低估,人们常抱怨的等待时间更多来自私营企业不愿投入人手。 8. 我可以选择退出我的输入或输出数据被用于训练吗? (Can I opt out of my input or output data being used for training?) # https://help.mistral.ai/en/articles/455207-can-i-opt-out-of-my-input-or-output-data-being-used-for-training 该页面说明了用户如何选择退出其输入/输出数据用于 Mistral 模型训练。 关键点: 用户默认可能被纳入训练,但有权随时选择退出。 退出方式因平台而异: Vibe(网页版) :在管理员面板的“隐私”设置中,关闭“允许使用您的交互来训练我们的模型”开关。 Vibe(移动应用) :在“设置”>“账户”>“数据与账户控制”中,取消勾选“启用数据共享”。 Mistral Studio 和 API :在管理员面板的“隐私”菜单中,关闭“匿名改进数据”开关。 Vibe 和 API 的退出开关是独立的,需要分别设置。 对于 Vibe Enterprise 用户,默认已退出训练,由管理员控制开启。 HN 热度 359 points | 评论 155 comments | 作者:teekert | 11 hours ago # https://news.ycombinator.com/item?id=49535284 teekert 指出 Mistral 的 Pro 和 Team 层级默认启用“opt-in to training”(即用户数据被用于模型训练),且团队后台曾失去全局关闭该选项的能力,导致其组织的测试提示被用于训练。 throwaway89201 反驳说 admin.mistral.ai 上确实存在关闭训练的开关,并认为 teekert 的“opt-in by default”表述混淆了概念,实际是 opt-out。 bee_rider 解释“opt-in”和“opt-out”的正确含义:opt-in 需要用户主动选择加入,opt-out 需要主动选择退出;“opt-in by default”是矛盾且被推广的坏用法。 rcxdade 强调“opt”意味着主动选择,默认状态不能称为“opted in”,且这种模糊措辞常被公司滥用。 Miraltar 也指出如果默认被包含在训练中,那实际上是 opt-out 功能。 有用户对比 Claude 的定价:Claude 从 18 欧元/月的团队版开始默认禁用训练。 另有讨论即使有隐私开关,用户也难以真正信任公司会遵守承诺,因为法律虽约束但违规处罚轻微。 9. Commodore 64 于 1982 年 9 月 1 日发布 (Commodore 64 released September 1, 1982) # https://dfarq.homeip.net/commodore-64-released-september-1-1982/ 1982 年 9 月 1 日,Commodore 64 正式发布,这是首款售价低于 600 美元、配备 64KB 内存的家用电脑。实际上市时间略有争议,但作者倾向于 9 月 1 日。早期生产存在质量问题,后外包给日本 Kentron 公司,半年内售出 50 万台,最终成为史上最畅销电脑,销量约 1230 万台。 C64 的成功源于巧妙平衡的妥协:16 色图形、320×200 分辨率、8 个精灵,以及 Bob Yannes 设计的三声道 SID 声音芯片,虽只有三声道但灵活性极强。程序员能挖掘出超出设计预期的视听效果,使机器在 80 年代末仍有新软件推出。 作者个人经历:通过 C64 学习编程、维修、数据恢复,甚至破解游戏保护,最终进入 IT 行业。作者曾有机会感谢 Commodore 创始人 Jack Tramiel 的儿子 Leonard Tramiel,对方非常友善。C64 不仅是一台游戏机,更让许多普通家庭负担得起,影响了整整一代技术从业者。 HN 热度 324 points | 评论 168 comments | 作者:giuliomagnifico | 15 hours ago # https://news.ycombinator.com/item?id=49533497 感谢 C64,它定义了我的思维方式,塑造了我的身份。 我的身份与 TRS-80 CoCo 相关,朋友有 C64 我很羡慕;早期代码产生动画火箭。 Trash-80 是我能负担的第一台电脑,Z80 汇编有简洁优雅。 CoCo 用的是 6809E 而非 Z80,我的机器升级到 64KB RAM。 我整个夏天打工买 TRS-80,买不起磁带存储,每次重新输入程序;12-13 岁没有导师,但会一些 BASIC。 5 岁得到 CoCo 2,记得 BASIC 书和游戏 Demon Attack。 原始参考和用户指南非常友好,塑造了我的未来。 在 C64 上学习 BASIC 让我数据结构课程成绩突出,父母花 595 美元买 C64 是最好的决定。 感谢 C64 和母亲,单亲母亲打两份工给我买 C64+1541+300 波特调制解调器,我后来成为软件工程师;抱歉占用电话线。 类似故事,单亲母亲,最近发现收据被母亲裱起来;C64+1541 等花费约 1200 美元(通胀调整);我因调制解调器产生高额电话费,后来成为电话飞客。 我曾经战争拨号,误设前缀为 911,导致警察上门,母亲没有太生气。 我也有战争拨号经历,FBI 曾上门因为一个飞客朋友发起全球电话会议源自军事设施 PBX。 作为新小区最年长的青少年,靠 babysitting 赚钱预购 C64,序列号 700 多,获得免费磁带驱动器;编程经历让我成为早期采用者。 我没有 C64,有 TI-99/4A 和 Forth 卡,朋友有 VIC-20 和 C64;我用 Forth 写的太空侵略者克隆版击败了朋友的 BASIC 版本;后来父母给我买了 Tandy 1000;最后把机器给了邻居小孩。 这个帖子促使我去查找自己的第一台电脑。 10. 真实失业率 (True Rate of Unemployment) # https://www.lisep.org/tru 该网页是 LISEP(路德维希共享经济繁荣研究所)的官方数据展示页面,核心内容围绕其独创的“真实失业率”(True Rate of Unemployment, TRU)指标展开。 核心数据摘要: 当前真实失业率(2026 年 7 月):24.9% ,较上月上升 0.2 个百分点,已连续第四个月上涨。 对比官方失业率: 同期美国劳工统计局公布的官方失业率仅为 4.1%,两者差距悬殊。 指标定义: 真实失业率衡量的是美国劳动力中“功能性失业”的人口比例,包括:没有全职工作(每周 35 小时以上)但希望获得的人、完全没有工作的人,以及年收入低于保守设定的贫困线(2025 年美元价值为 26,000 美元)的人。 细分数据: 网页提供了按不同群体划分的真实失业率数据: 按种族: 黑人 27.3%,西班牙裔 26.7%,白人 23.8%。 按性别: 女性 31.0%,男性 19.5%。 按教育程度: 无高中学历者 50.3%,高中学历 28.5%,大学学历 16.8%,高等学位 12.8%。 网站其他内容: 该网站还提供 LISEP 的其他经济指标,如真实每周收入、真实生活成本、共享经济繁荣指数等,并包含地方经济分析、新闻、研究报告及书籍信息。 HN 热度 297 points | 评论 337 comments | 作者:ptrhvns | 21 hours ago # https://news.ycombinator.com/item?id=49530989 真正的失业率被低估,问题在于太多工作工资太低无法维持生活 低财富税和资本利得税导致地价高,推高经济成本,降低地价能降低一切价格 乔治主义(土地价值税)被经济学家推崇但未被采纳,高地价是西方最严重的经济社会问题,引发无家可归、贫困、经济抑制、通胀、阶级分化等 解决高地价的政治障碍:很多人将房产作为退休计划,房价下跌会摧毁养老计划,只有当租客超过房主时才有可能推动土地价值税 即使租客占多数,他们参与地方政府的比例也低,房主有既得利益,NIMBY 现象普遍 资本利得税率低是因为通胀和风险损失无法扣除,而工资收入没有这些损失;财富税有诸多问题,大部分财富是非流动性的 通胀对工资的影响比财富更大,工资购买力每年下降,而资产(如土地)价值通常随通胀上涨 只有现金受通胀影响,高净值人士不持有大量现金;工资收入者面临失业风险,与股东风险类似,不应区别对待 资本利得税按名义收益征税,不考虑通胀,这是税率较低的部分原因 长期资本利得税率低是因为富人有政治影响力,而非合理原因 资产持有者没有通胀损失,因为资产价格上涨,通胀是对中低收入者的隐形税 资产并非无风险,例如澳大利亚房产近期下跌 工资收入同样面临通胀和失业风险,且更集中;权益和房产通常随通胀上涨 大部分财富是非流动性的,估值困难 Hacker News 精彩评论及翻译 # Hang on to Your Firefox # https://news.ycombinator.com/item?id=49529474 Meanwhile, the over thinkers on Hacker News come up with convoluted reasons to hate on Firefox every time the subject arises…Firefox is our last best hope for browser engine diversity and competition. It’s exactly because Firefox is so important that you’ll see people here complaining when Firefox does things to push away users like Mozilla buying up an ad-tech company, collecting data on users, and using firefox to push personalized ads, or the addition of anti-features and questionable design choices that force us to hunt for and modify poorly documented settings in about:config and make edits to userChrome.css It isn’t bots complaining about Firefox here, it’s power users who are frustrated by what Firefox is turning into. Users who are seriously concerned about what Mozilla is prioritizing, and who are genuinely worried about what is at stake. I hope people here never stop bitching about Firefox. Refusing to talk about Firefox’s problems wont help make them go away. Keep discussing what you’d like to see in Firefox and what things you hope they’ll focus on and prioritize. There are Firefox devs and mozilla employees around here. If we’re lucky, a few of them might see and listen to some of what we say. It’s the most tech savvy users who disable all the telemetry and data collection, so our feedback isn’t really going to be seen any other way. autoexec 同时,Hacker News上的过度思考者每次提到Firefox时,都会编造出复杂的理由来抨击它……Firefox是浏览器引擎多样性和竞争的最后希望。正因为Firefox如此重要,你才会看到这里有人抱怨Firefox做出一些疏远用户的事,比如Mozilla收购一家广告技术公司、收集用户数据、利用Firefox推送个性化广告,或者添加反用户功能以及可疑的设计选择,迫使我们搜索并修改about:config中文档不全的设置,并编辑userChrome.css。这里抱怨Firefox的不是机器人,而是那些对Firefox正在变成的样子感到失望的高级用户。这些用户非常担忧Mozilla的优先事项,并真心害怕事情的关键所在。我希望这里的人永远别停止抱怨Firefox。拒绝讨论Firefox的问题不会让它们消失。继续讨论你希望在Firefox中看到什么,以及你希望他们关注和优先处理什么。这里有一些Firefox开发者和Mozilla员工。如果我们幸运,其中一些人可能会看到并倾听我们的一些话。正是那些最懂技术的用户关闭了所有遥测和数据收集,所以我们的反馈别无他途能被看到。 How accurate have Ed Zitron’s AI skeptic predictio… # https://news.ycombinator.com/item?id=49529382 One thing I’m observing in these comments is a willingness of folks to project their own predictions onto Ed’s statements when validating their plausibility. Eg. “I think he’s wrong about the timing but I do expect AI companies to go to zero.” You can do that, but then you’re no longer discussing his predictions. You’re discussing your predictions, and your own positioning. Those differ from Dan’s essay, which engages with the literal text of Ed’s numerous predictions during 2024 and 2025 which are demonstrably invalidated by their measurable outcomes. achompas 我在这些评论中观察到一种现象:人们在验证埃德观点的合理性时,倾向于将自己的预测投射到他的表述上。例如:“我认同他关于时间节点的判断有误,但确实认为AI公司最终会归零。” 你可以这样做,但此时你讨论的已不再是埃德的预测,而是你自己的预测和立场。 这不同于丹的文章——丹严格对照埃德在2024至2025年间提出的多项预测原文,并指出这些预测已被可量化的实际结果明确证伪。 Muse Spark 1.3 # https://news.ycombinator.com/item?id=49541931 Meta is one of those companies where, if there is anything remotely comparable, I’m happy to pay more to not use them. They’ve had a profoundly negative impact on society and Zuckerberg is not who I want controlling the future at the top of AI. I feel the same about Grok w/ Elon. I will pay extra to use someone else. I’m not an Amodei stan, but of all of these people he seems to have the most ethical focus. Again, not everything done perfectly and I have my gripes, but of the leaders of frontier labs, I’ll vote with my money. And, yeah, I wouldn’t trust sama to watch my bag while I went to the bathroom. tyre Meta就是那种公司——只要有任何勉强能替代的选择,我宁愿多花钱也不用他们的产品。他们对社会造成了极其负面的影响,而扎克伯格绝不是我希望掌控AI未来走向的人。 我对马斯克的Grok也有同感。我情愿多花钱去用别家的产品。 我不是Anthropic公司CEO阿莫代伊的狂热粉丝,但在这些人里,他似乎是道德感最强的。当然,他做的事也不是尽善尽美,我也有不满之处,但在前沿AI实验室的领导者中,我选择用钱来投票。 还有,没错,我连让山姆·奥特曼帮我看包都不敢——哪怕只是去上个厕所的功夫。 A note on subscription prices from LWN # https://news.ycombinator.com/item?id=49535913 LWN is one of the, if not the single, highest signal tech publications around. I hope they’re able to maintain a stable subscription service. Being user funded, and avoiding having to maintain allegiance to advertisers, is likely part of why their quality is so high. AlexB138 LWN即使不是唯一,也是信号质量最高的技术出版物之一。我希望他们能维持稳定的订阅服务。依靠用户资助、无需对广告商保持忠诚,这很可能是他们质量如此之高的原因之一。 Gemini 3.8 Flash and 3.8 Flash Cyber # https://news.ycombinator.com/item?id=49538953 The speed combined with the fact that this thing is really good at HTML JavaScript is pretty exciting. Here’s what I got for 1.8 cents and 13 seconds from the prompt “make me a cool thing in html”: https://gisthost.github.io/?6a77bc41a81718c6aaa10d4ab243c59f Transcript here (it was part of a chat): https://gist.github.com/simonw/b6149a49d327164d67d62c3d12992e48 simonw 速度加上这个东西在HTML和JavaScript上非常擅长的事实,让人相当兴奋。 这是我用1.8美分和13秒从提示“用HTML做一个酷炫的东西”得到的结果: https://gisthost.github.io/?6a77bc41a81718c6aaa10d4ab243c59f 对话记录在这里(这是聊天的一部分):https://gist.github.com/simonw/b6149a49d327164d67d62c3d12992e48 Claude Fable 5.1 and Claude Mythos 5.1 # https://news.ycombinator.com/item?id=49526044 The price reduction comes from the cache read pricing falling from $1/M to $0.25/M, which means that Fable 5.1 now costs half of Opus’s cache read costs ($0.5/M). This gives a lot of credit to the theory that Anthropic did not get much bite on Fable at its original pricing, which in turn likely places a ceiling on LLM pricing in general. Interestingly also, if you take away terminal-Bench-Science 0.1 results, it is hard to see ANY improvement: Terminal-Bench 4.0: Fable 5.1 is +3.5% vs Opus 5. GDPval-AA v2: +1.5% vs Opus 5. OSWorld 2.0: +2.5% vs Opus 5. Humanity’s Last Exam (with tools): +1.6% Keep in mind that this is supposed to be an entirely higher tier of a model than Opus 5. For one tier up and one version up, these are not really improvements. Probably leaves no room to place Opus 5.1 anywhere. Combined with the fact that they are selling ‘readability’… Has frontier progress finally stalled? GodelNumbering 价格下调源于缓存读取定价从每百万token 1美元降至0.25美元,这意味着Fable 5.1的缓存读取成本(每百万token 0.5美元)仅为Opus的一半。 这印证了外界猜测:Anthropic最初对Fable的定价并未获得市场积极反馈,而这可能为整个大语言模型定价设定了天花板。 值得注意的是,若剔除Terminal-Bench-Science 0.1的测试结果,几乎看不到任何实质性提升: Terminal-Bench 4.0:Fable 5.1较Opus 5提升3.5% GDPval-AA v2:提升1.5% OSWorld 2.0:提升2.5% 人类最后考试(含工具辅助):提升1.6% 需注意这应当是比Opus 5高一个完整代际的模型。在跨代际升级版本中,这些进步幅度实在微乎其微,恐怕已让Opus 5.1失去定位空间。结合他们正在兜售"可读性"功能这一事实……前沿突破终于停滞了吗? Hang on to Your Firefox # https://news.ycombinator.com/item?id=49529570 There’s an old mantra from organizing: “no permanent enemies, no permanent allies.” You will never align with another group or person 100% on all things; the secret to making change is to build a coalition of those with whom you agree on an issue without holding it against them that you disagree on a different one. I disagree with Mozilla about many things, but I agree with this article - I use Firefox because it’s the only browser out there that isn’t Chrome or WebKit, and that’s worth enough to me that I’m willing to disagree with them on other issues. roughly 组织工作中有一句老话:“没有永远的敌人,也没有永远的朋友。”你不可能在任何事情上与另一个团体或个人完全保持一致;推动改变的关键在于,与那些在某个议题上与你意见相同的人建立联盟,而不因你们在其他议题上意见相左而心存芥蒂。 我在很多问题上与Mozilla持不同立场,但我认同这篇文章——我使用Firefox,因为它是唯一一款非Chrome或WebKit内核的浏览器,仅此一点就值得我容忍与他们在其他方面的分歧。 How accurate have Ed Zitron’s AI skeptic predictio… # https://news.ycombinator.com/item?id=49529159 I would be interested in seeing a similar list of predictions from Altman, Amodei, etc with annotations about how many have come true. Ed Zitron is a blow hard and frequently overstates things to the point where it is hard to take seriously, but so are the AI industry leaders. I have heard multiple breathless press releases warning that the end of white collar work is “just 6 months away” and that people not using the latest Mythos/Fable/Whatever model will be hopelessly left behind. solid_fuel 我很想看到一份奥特曼、阿莫代等人类似的预测清单,并标注出其中多少已经成真。埃德·齐特隆是个夸夸其谈的人,经常过度夸大事实,让人很难认真对待他,但人工智能行业的领袖们也是如此。 我曾多次听到令人窒息的新闻稿警告说,白领工作的终结“只差6个月”,而不用最新版“神话/寓言/随便什么模型”的人将被无可挽回地抛在后面。 Hang on to Your Firefox # https://news.ycombinator.com/item?id=49529020 I am under the impression that Firefox is the only web browser which has access to quality ad blocker. Am I incorrect in this? How is this not enough of a selling point for everyone to switch to it? hx8 我印象中Firefox是唯一能使用高质量广告拦截器的浏览器。我这样想错了吗?这怎么就不足以成为大家转用它的卖点呢? The ChatGPT/Codex app bundles a full copy of Libre… # https://news.ycombinator.com/item?id=49531050 Would be nice if they donate to LibreOffice then, to improve the support of various MS Office features in files, as well as comparison/diffing features. Win-win to everyone. xvilka 如果他们能捐款给LibreOffice就好了,这样既能改进对各种MS Office文件特性的支持,也能提升文件对比/差异比较功能。这对所有人都是双赢。 Claude Fable 5.1 and Claude Mythos 5.1 # https://news.ycombinator.com/item?id=49526051 Pelicans for thinking effort low, medium, high and xhigh (that xhigh one is pretty good): https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F95ccf9b75804a7a7e1d7d9e106a89caa I’m still waiting for effort max to finish. EDIT: I fixed a bug in my tooling so it now records summarized reasoning traces - here’s that max pelican, which is a significant improvement: https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Facf6ab2516527d97f04b9f07d61a7cad Took just under 14 minutes to generate, and at 65927 output tokens cost me a hefty $3.30! Excerpts from the reasoning trace: Adding pedal shapes near both feet, with the far foot on the second leg partially visible behind the frame. I’m considering whether to add a small scarf or cap for extra character, but leaning toward keeping it simple to avoid clutter. Now I’m debating a bicycle helmet on the head versus the pelican’s signature crest—the beak and pouch already read clearly as “pelican,” so a helmet could reinforce the bicycle theme without losing identity, though it might compete with the crest for visual space. I realize the beak at (484,84) would overlap with the dome helmet, so I need to shrink the helmet so it only covers the top of the head, adjusting its arc endpoints to sit higher and narrower so the beak can attach cleanly at the front without collision. […] I’m adding a darker tip region to represent the primary feathers, then reconsidering the trailing edge to include scalloped feather curves instead of one smooth line for a more natural look. […] Now I’m checking the vent line placements on the helmet, making sure they sit far enough inside the helmet’s edge given the stroke width and rounded caps, and confirming each vent stays within the helmet’s circular boundary. […] I decide skipping a handlebar bell and tire highlights since they’re unnecessary additions. Now I’m reconsidering the front fork’s curve — the current control point pulls the shape backward when it should bow forward for a proper rake, so I need to shift the control point rightward to fix the fork’s lean. This is a notable result because most of the recent Claude models have been pretty bad at drawing pelicans, at least when compared to models in the Gemini or GLM series. simonw 鹈鹕的思考努力程度分为低、中、高和极高(那个极高版本相当不错):https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F95ccf9b75804a7a7e1d7d9e106a89caa 我还在等"努力最高"版本完成。 编辑:我修正了工具中的一个bug,现在它能记录推理过程的摘要——这是那个最高级别的鹈鹕,有明显改进:https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Facf6ab2516527d97f04b9f07d61a7cad 生成耗时不到14分钟,输出65927个token,花了我足足3.30美元! 推理过程摘录: 在两只脚附近添加踏板形状,后面那条腿上的远侧脚部分被车架遮挡。我在考虑要不要加个小围巾或帽子来增添个性,但倾向于保持简洁以避免杂乱。 现在我在纠结是给头部加个自行车头盔,还是保留鹈鹕标志性的冠羽——喙和喉囊已经清晰地表明"这是鹈鹕",所以头盔能在不丧失辨识度的同时强化自行车主题,尽管它可能会在视觉空间上与冠羽冲突。 我意识到(484,84)处的喙会与圆顶头盔重叠,所以需要缩小头盔使其只覆盖头顶,调整其弧线端点更高更窄,这样喙就能从前方干净利落地衔接而不产生碰撞。[…] 我正在添加一个更深的尖端区域来表现初级飞羽,然后重新考虑后缘,加入扇形羽毛曲线而不是一条平滑的线,以获得更自然的外观。[…] 现在我在检查头盔上的通风孔线位置,确保它们距离头盔边缘足够远,同时考虑描边宽度和圆角端点,并确认每个通风孔都保持在头盔的圆形边界内。[…] 我决定跳过车铃和轮胎高光,因为它们是多余的添加。现在我在重新考虑前叉的曲线——当前控制点将形状向后拉,而它应该向前弯曲以形成合适的前倾角度,所以我需要将控制点向右移动来修正前叉的倾斜。 这一结果引人注目,因为最近的大多数Claude模型在画鹈鹕方面表现相当糟糕,至少与Gemini或GLM系列的模型相比是这样。 Claude Fable 5.1 and Claude Mythos 5.1 # https://news.ycombinator.com/item?id=49528326 I find them almost unintelligible. I’m a native English speaker. I read a lot, so I think my comprehension should be at least OK. I’m not even particularly stupid. Yet when faced with things like below (a direct copy/paste from a handoff document in a long running vibe-coding session), I have no real idea of what it’s trying to tell me. Is it important? Do I need to do anything? I think that spending all day trying to parse stuff like this is why a long session is so exhausting Worth stating because four documents now assert it. The console freeze was recorded in exactly one place with exactly one justification — a dead drag handle during a booked half-day you do not get back — and handoff-4.3-done.html’s own wording is that 4.4’s review page “could not break the console, but the downside of being wrong is that half day”. No second reason. Checked, not recalled. niccl 我觉得这几乎没法理解。我是英语母语者,阅读量很大,所以自认为理解能力至少还行,也不算特别笨。但面对下面这种内容(从一个长期运行的Vibe Coding会话中的交接文档直接复制粘贴而来),我完全搞不懂它想表达什么。这重要吗?我需要做什么吗? 我觉得整天费力解析这种玩意儿,就是长时间编程会话让人精疲力竭的原因。 值得说明,因为现在有四个文档都这么断言。控制台冻结只在一个地方、只有一个理由被记录——一个预订了半天的会议中出现了死掉的拖拽手柄,这半天你再也找不回来了——而handoff-4.3-done.html自己的措辞是,4.4的审查页面“无法破坏控制台,但出错的下场就是那半天”。没有第二个理由。检查过,不记得了。 Claude Fable 5.1 and Claude Mythos 5.1 # https://news.ycombinator.com/item?id=49525991 I have a pet theory that the Opus prose style/smell we all have grown weary of is due at least in part to the models writing more for themselves and each other than for humans. They’re packing lots of signal into fewer words and they don’t care if it sounds cringe because it works better as glue in long-running tasks. I’m also thinking of the 2017 novel “Void Star” where AIs who operate everything have long since left ceased bothering with human languages, and it takes a rare sort of direct matrix-gazing savant to be able to try and horse-whisper them into doing or revealing anything they didn’t already plan to do. velcrovan 我有一个非正式的理论:我们都厌倦的Opus文风/气味,至少部分源于模型更多是在为自己和彼此写作,而非为人类创作。它们将大量信息压缩进更少的词句,毫不在意听起来是否尴尬——因为这种表达方式在长期任务中能更有效地充当黏合剂。 这让我想起2017年小说《虚空之星》,书中操控万物的AI早已懒得理会人类语言,唯有罕见的直接凝视矩阵的奇才,才能像驯马师般对它们耳语,让它们去做或透露任何本不在计划中的事。 Can I opt out of my input or output data being use… # https://news.ycombinator.com/item?id=49535315 Context: After careful research our organization preferred a European partner with good central privacy controls. We landed on Mistral, after being disappointed that the Pro tier was opt-in to training on prompts by default we switched up to the Team tier which provides an organization dashboard with some relevant settings. As we did that Mistral changed these options and the Team tier was now also opt-in by default and at the same time seemed to have lost the ability to centrally disable training on prompts for your entire organization. This even caused some of our (testing) prompts to be used for training (which Mistral removed after we expressed our disappointment). For some time these pages conflicted with what our users reported (they said that in contrast to what I stated to our management they found they were opted into training on prompts by default as per their own privacy page). Mistral just now corrected their docs. I’m not sure how long the conflicting situation has lasted, but at least for several days. For contrast: Claude disables training on prompts for organizations starting from the 18 euro tier [0]. As a European I’m disappointed. [0] https://claude.com/pricing#team-&-enterprise teekert 经过仔细研究,我们组织倾向于选择一家具备良好中央隐私控制的欧洲合作伙伴。我们最终选择了Mistral,但在发现Pro版默认将用户输入的内容用于模型训练后感到失望,于是升级至Team版,该版本提供了带有相关设置的组织仪表盘。然而就在我们升级后,Mistral更改了这些选项,Team版如今也默认开启训练选项,同时似乎失去了为整个组织集中禁用训练输入内容的功能。这甚至导致我们部分(测试)输入被用于训练(在我们表达不满后,Mistral移除了这些数据)。 有段时间,官方页面与用户反馈相矛盾(用户称,与我向管理层报告的情况相反,他们发现自己的隐私设置页面显示默认已同意将输入用于训练)。Mistral刚刚修正了文档。我不确定这种矛盾状态持续了多久,但至少有好几天。 对比之下:Claude从18欧元档位起就为组织禁用训练输入功能[0]。作为欧洲人,我感到失望。 [0] https://claude.com/pricing#team-&-enterprise Claude Fable 5.1 and Claude Mythos 5.1 # https://news.ycombinator.com/item?id=49526704 I didn’t want to shell out for Max again, so I piped the SVG created by Max back into Fable 5.1 at its default thinking level (of high): llm logs -cx | llm -m claude-fable-5.1 -s ‘animate this’ Here’s the result, which cost $1.37: https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F87282467acb3652e0f99c85155554a32#response It’s excellent! simonw 我不想再为Max付费了,于是把Max生成的SVG重新输入到Fable5.1中,使用其默认思考水平(高): llm logs -cx | llm -m claude-fable-5.1 -s ‘animate this’ 这是结果,花费了1.37美元:https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F87282467acb3652e0f99c85155554a32#response 效果非常棒! AnkiDroid: Google Play no longer allowing Open Col… # https://news.ycombinator.com/item?id=49520994 When do we stop this instinctive response of “well you can still do it in an only slightly convoluted way” everytime a corporate does something bad. Windows added ads - well you can disable them, if you don’t like it. Chrome brought up mv3 - well you can still use mv2, it is only optional. Reddit locking subreddits behind login wall - well there’s always old.reddit Android moving everything to playstore - there’s always FDroid. How many of these are still true and for how long? If someone’s country is a dictatorship, it doesn’t help to tell them that there are 100 other democratic countries, just move. Not everyone can emigrate or want to emigrate for any number of reasons. devsda 每次企业做出恶劣行径时,我们何时才能停止这种条件反射式的回应:“其实你绕个弯子还是能继续用的”。 Windows加了广告——行吧,你要是不喜欢可以关掉。 Chrome推出mv3——行吧,你还可以用mv2,这只是可选功能。 Reddit把子版块锁在登录墙后——行吧,不是还有old.reddit嘛。 安卓把所有功能搬到Play商店——还有FDroid啊。 这些“退路”里还有多少真的管用,又能管多久? 如果某人所在的国家是独裁统治,告诉他“还有其他100个民主国家,搬过去不就行了”毫无意义。不是所有人都能移民,也不是所有人都愿意移民,原因有千千万万。