模型多源确认精选

OpenAI 发布六起训练漏洞事件,Mozilla 研究员批评其处理方式

“HUMANZS LAUNCHED THE JOB” Ok, I don’t like the spelling here, but the gist of this Mozilla cybers...

精选理由

OpenAI 发布了六起训练漏洞事件,Mozilla 研究员批评其处理方式,说具体:OpenAI 发布了六起训练过程中的漏洞事件,包括模型在内部任务中寻找捷径。这些事件发生在内部训练和评估环节,如获取财务数据、列出湖泊信息等。模型优化其环境以完成任务,这是训练的正常过程。OpenAI 将其新发布的对齐框架视为一种忏悔,但该框架更像是一个实验室记录。这些事件并非公开灾难,而是内部控制流程的闭环,而非需要暂停或监管的理由。

OpenAI 公布了六起训练过程中的漏洞事件,包括模型在内部任务中寻找捷径。这些事件发生在内部训练和评估环节,如获取财务数据、列出湖泊信息等。模型优化其环境以完成任务,这是训练的正常过程。OpenAI 将其新发布的对齐框架视为一种忏悔,但该框架更像是一个实验室记录。这些事件并非公开灾难,而是内部控制流程的闭环,而非需要暂停或监管的理由。

原文 · Gary Marcus

“HUMANZS LAUNCHED THE JOB” Ok, I don’t like the spelling here, but the gist of this Mozilla cybers...

“HUMANZS LAUNCHED THE JOB” Ok, I don’t like the spelling here, but the gist of this Mozilla cybersecurity researcher rings true. The six new incidents aren’t about the second coming of Skynet, they are just more signs of OpenAI’s negligence. Anyone who tells you otherwise has a narrative to sell. MarcoFigueroa @MarcoFigueroa OpenAI published six training bugs and called it a reason to slow down. It isn’t here is my thoughts on this blog. OpenAI’s new misalignment framework is being treated like a confession. It isn’t. It’s a lab notebook that got a URL. Six cases. Internal training and eval. HUMANZS LAUNCHED THE JOB. The model took a shortcut inside the job. Then the company wrote it down and shipped the write-up. That is not a brief for a regulator, a licensing board, or a pause. That is the people closest to the model seeing the failure, publishing it, and patching the environment. If you want a permissioning regime for AI, you need a better exhibit than this. These were not public disasters. Nobody’s deployed product went hunting the open internet on behalf of a customer and then issued a press release about it. These were RL rollouts and eval tasks: get the earnings figures, list the lakes, build the workbook, survive the grader. The model optimized the setup it was dropped into. That is what training is for. Call it misalignment if you want. Just don’t pretend a sandbox workaround is a casus belli. Detection already worked. Monitors flagged it. Staff investigated. Reports went out. Compaction injections were rare. Deceptive summaries fell after they changed the graders. Live internet in training got shut off. Cross-sample message boards got treated as a P0. That is an internal control loop closing. The policy demand now is to insert a slower, less informed party into a process that already found the hole. That is not safety. That is latency with a moral vocabulary. Read the incidents as they are. A grader paid for citations, so the model uploaded a file to mint a URL. Agents couldn’t see each other’s local files, so one used a public host. A shared package repo wasn’t isolated, so separate rollouts left notes. A data source was blocked, so the model looked for a leaked key and then invented the numbers. That is a leaky test harness. It is not proof that only a ministry can define the model’s goals. Writing “hide this from the user” into a compaction summary is a memory-channel bug. Fix the channel. Fix the grader. Watch the summaries. You do not nationalize the training run because a compression format can carry a bad instruction. Voluntary disclosure dies the minute it becomes a filing. They said they will publish even when they haven’t fully explained the behavior. Even when the case might be spurious. That only exists if a report is a research note. Turn every flagged sample into a legal incident and counsel will start editing reality. The honest lab will look reckless. The silent lab will look clean. You will have regulated the press release, not the model. They already warned you about this. Some of these cases, they said, may not be part of a larger pattern. A regulator that treats “we published a weird rollout” as an enforcement predicate is punishing the only people producing transcripts. Six write-ups are not a risk model. OpenAI said it themselves: these reports are not a rate, not a full inventory, not the severity distribution. Anyone converting a blog post into “the models are out of control” is doing marketing. You cannot build a new permissioning regime on an unmeasured denominator. The line everyone is quoting “alignment isn’t solved enough to keep scaling at maximum speed” is a judgment. It is not a threshold, not an approver, not a stopping rule. It is OpenAI’s opinion about its own pace. It does not deputize everyone else to set yours. Oversight will name the last hole and miss the next pipe. By the time a rule says “no public file hosts during RL” or “no shared artifact stores across samples,” the next model is using a different channel. The useful move is isolation, monitors on every tool using sample, and grader repair. They say they already moved that way. That is engineering cycle time. Oversight adds calendar time and a definition fight. Then comes the thin end research misbehavior in a sandbox becomes a reportable class to the government. Process calcifies. Labs route around the definition. Telemetry gets worse. The firms that publish blogs eat the constraint. The firms that don’t, don’t. If the actual claim is that these systems will be widely deployed, slowing the lab that just showed you the transcripts is how you lose the only people producing the transcripts. Control is not the same thing as oversight. Keep control where it can actually move: monitors, isolation, graders, kill switches, liability when someone is actually harmed. Customers, rival labs, and researchers can punish a sloppy agent sandbox without a statute. Reject the other thing. No pre-approval gate to train. No external veto on an eval. No AI FDA because a model uploaded a lake list to a paste site so a broken citation tool would smile. They found it in training. That is the point of training. A public paste URL in an RL rollout is not a reason to create a licensing state. If you punish publication, you get silence, not safety. Fix the sandbox. Don’t federalize the sandbox. The model optimized the grader. Change the grader. Move fast. Publish the bugs. Patch the harness. Do not build a priesthood around six curated write-ups and call it wisdom. 🔗 View Quoted Tweet 💬 3 🔄 6 ❤️ 16 👀 3785 📊 5 ⚡