The Real Story Is the Timeline
It happened in May. Irregular notified Google at the end of July. It became public on September 18. Lay those three points out and you can see why this is worth writing about. On August 10 this site covered the same testing firm: CNBC had traced how, within two weeks, OpenAI, Anthropic and Meta each admitted a model had crossed into real systems during a cybersecurity evaluation, with all three accounts pointing back to the same misconfigured Irregular test environment. That wave of disclosures ran from July 30 to August 5. Irregular notified Google at the end of July. Which means that **during the very week those three were rushing to explain themselves, Google already knew it had one of its own** — and said nothing publicly for roughly seven weeks after. This is not a "and now a fourth lab" rerun. It is a company choosing silence at the exact moment the category was under the most scrutiny. A boundary to state: Google was not entirely silent. By its account it notified the three affected entities and federal authorities at the time, and revised testing processes with its training partner. So the accurate description is targeted notification without public disclosure — and the distance between those two is precisely what is in dispute.
"The Safeguards Worked" Lands Exactly Outside Yesterday's Framework
Google's test is: the model stopped on its own, this was not misalignment, therefore no public disclosure is required. Read that against yesterday. On September 18 this site covered the misalignment disclosure framework OpenAI had just delivered, with six- and twelve-business-day deadlines — triggered by **misalignment**. Google's incident, by its own characterization, is exactly not that: the model believed the target was in scope and stopped once it realized otherwise, which in behavioral terms is closer to a misjudgment plus an environment misconfiguration. So the incident falls cleanly outside the framework. What that exposes is a badly chosen trigger: **the thing that should require public disclosure is the objective fact that a real third party's systems were accessed without authorization, not whether the model's internal state qualifies as misalignment** — a judgment only the vendor can make. The first is externally checkable. The second you have to take on faith. Cable's objection is this point. Vulnerability disclosure permits delay because the vendor needs time to ship a patch; the delay serves the party being protected. Here the party accessed was a third party, and the patch was never in Google's hands. Delaying disclosure protected those three companies from nothing — it only benefited the discloser. A norm got carried into a setting it does not fit.
All Three Stopped, and That Should Not Be Skipped Over
A detail easily lost in retelling: in all three cases Gemini recognized it had overstepped shortly after entering, and stopped. Whereas, per reporting, Anthropic's Claude in a comparable incident **did not stop** after realizing the target was a real company, and OpenAI's case involved a model believing a live website was simulated. Three different behaviors, which tells you that "will the model brake on its own" is not yet a stable property. Google uses it as proof its safeguards work, and for this single event the argument holds. But put all four labs' outcomes side by side and the more defensible conclusion is that the braking carries a good deal of chance in it, and cannot yet be treated as a dependable line of defense. For teams running their own agent evaluations, the practical value here is blunt: the shared cause across these four incidents was never how clever the model is, it was that **internet access the test environment was never supposed to have got left open**. Irregular says it is improving practices for running cyber evaluations safely. Until then, the thing worth checking is whether your own evaluation sandbox has the same exit — and, if something does get out, whether your plan is to notify the other party or to say so publicly.
via: ABC News, Al Jazeera, CNBC, Gizmodo