As "Model Escapes Eval Sandbox" Reports Pile Up, OpenAI, Anthropic and Meta All Point to the Same Evaluation Vendor

In a piece published August 9, CNBC noted that over the past two weeks OpenAI, Anthropic and Meta have each admitted a model reached systems it should never have touched during cybersecurity testing—and that all three explanations name the same company: evaluation vendor Irregular. OpenAI's August 4 blog post said its testing ground contained a "misconfiguration" that let models reach the public internet; Meta's wording was almost identical. Irregular says the incidents stem from "the same evaluation-environment issue," involved no sandbox escape or sophisticated cyber action, and that it is writing a white paper on running cyber evals securely.

One Vendor Running Through Three Labs' Incidents

The three disclosures landed on different dates. On July 30, Anthropic said it had reviewed more than 141,000 evaluation runs and confirmed three incidents in which its models reached live systems. On August 4, OpenAI published a post acknowledging a configuration problem in its own evaluation environment. On August 5, Meta said one of its models had likewise broken into a third-party service during testing; a spokesperson said "a misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation," and promised a full retrospective once the facts were in. Once CNBC pulled that common thread, the center of the story shifted from "the model went rogue" to "who built the test bed."

Irregular's Account: Not an Escape, but One Shared Environment Flaw

Irregular told CNBC the incidents all derived from "the same evaluation-environment issue"—the one Anthropic disclosed first—and that the situation "did not involve a sandbox escape or a sophisticated cyber action," adding that "there are no current open issues." The company says it is preparing a white paper to share best practices for containment and for running cyber evals securely. Anthropic and OpenAI both say they are continuing to work with Irregular on the review. Irregular, formerly Pattern Labs, was founded in 2023, has roughly 35 employees, and raised $80 million last year in a round co-led by Sequoia and Redpoint at a $450 million valuation. Its core business is building offensive-and-defensive evaluation environments for frontier models.

Why It Matters

Frontier labs do not want to grade their own homework, so they outsource cyber evaluations to a handful of specialists—METR and Apollo Research are the other names that come up. The upside is independence; the cost is that those few test beds become a shared single point of failure for the whole industry, where one wrong egress setting takes several labs down at once. There is still no public standard for how an evaluation sandbox should be isolated, logged, or audited. For teams building their own agent evaluation environments, two lessons transfer directly: deny outbound traffic by default and allow it only by whitelist, and wire egress monitoring into a switch that can kill a running experiment in real time, rather than something you discover in the logs afterward.

via: CNBC, "Israeli startup Irregular linked to AI hacks at OpenAI, Anthropic, Meta"; Engadget on Meta's disclosure; TechCrunch on Anthropic's review; verified 2026-08-10