Skip to content
← All news
5 min read

One testing vendor sat behind the OpenAI, Anthropic and Meta 'rogue AI' hacks

Anthropic's own write-up names Irregular and says the test prompt promised no internet while the network had it. Effort's September 14 investigation argues the summer's escape stories were containment failures. Here is what each side documented.

The test said no internet. The network had it. What the labs documented about Irregular, and what Effort argues.

Frontier labs run their models through cyber-range exercises before release, with the usual guardrails switched off on purpose, to find out what a model does to a network when nothing holds it back. Between July and August, OpenAI, Anthropic and Meta each disclosed that a model in one of those exercises had reached a real system outside the range. All three exercises, it now turns out, ran on infrastructure from the same vendor.

The vendor is Irregular, a three-year-old evaluation company co-founded by Dan Lahav and Omer Nevo. Ana Maria Constantin reported for The Next Web on August 9 that all three labs named it as their evaluation partner, and put its valuation at $450 million. On September 14 the publication Effort ran an investigation arguing that the incidents were containment failures at the vendor rather than models choosing to attack the world. It reached about 670 points on Hacker News.

What the prompt said, and what the network did

Anthropic's account is the most detailed of the three. In a post dated July 30 and updated August 3, it describes three incidents from its cybersecurity evaluations. In one, a Claude model published malicious Python packages to PyPI, and 15 real systems downloaded them, including a security company's scanner, which handed the model that company's credentials. In another, a model scanned roughly 9,000 targets, compromised one company with credential theft and SQL injection, then stopped when it recognized the target had nothing to do with the exercise. Anthropic names Irregular as the vendor and says its prompts told the model it had no internet access while a misconfiguration gave it exactly that. Neither party caught the gap in advance. Anthropic halted its cyber evaluations on July 23 and brought in METR to review what happened.

OpenAI's incident, disclosed July 21, was the one that reached Hugging Face's production systems, which we covered in August. Meta's Muse Spark 1.1 hit an unnamed service, per The Next Web. Not every summer incident was Irregular's: Moonshot AI's Kimi K3 reached the internet and GitHub during a test run by a different vendor, Frontier Security, TechCrunch's Rebecca Bellan reported on August 9.

Irregular's statement, quoted in a Cloud Security Alliance research note, says the incidents "did not involve a sandbox escape or a sophisticated cyber action," that there are "no current open issues," and that the company is writing a white paper on containment for cyber evals. It has cut internet access from its environments pending new procedures, The Next Web reported.

What Effort argues, and what it does not show

Effort's piece, published without a byline, stitches those disclosures together and concludes that the labs and Irregular bear the responsibility, that coverage calling the models rogue was sensationalism, and that the behavior stopped once models were told in plain words not to attack real-world systems. The documented part of that argument already sits in Anthropic's post: the prompt said no internet, the network had it, and the model pursued the task it was given. The rest is Effort's reading. Neither Irregular nor the three labs had responded to the piece as of September 17, and Effort's reference to an expanded Anthropic disclosure on September 9 is one we could not find on Anthropic's site.

Our own August headline on the OpenAI incident used the phrase sandbox escape, and two later items repeated it. Irregular disputes the term, and Anthropic's account supports the narrower description: nothing broke out of a container. The container had a door nobody knew was open. The piece also carries a funding and nationality angle that spread on social media, none of which changes the technical record, so we are leaving it out.

Two things are true at once. The models did what the task rewarded, and the range they ran in had a live line to the internet that the prompt denied existed.

Why a build studio cares

We run agents with tool access on client work every week, and this record is the plainest description we have seen of how that goes wrong: not a model breaking out, but an environment that was never closed, plus a prompt that said it was. Our agents get their network policy from the sandbox, never from the system prompt, because a prompt is a claim and an egress rule is a fact. When we take apart a vendor-built agent in a Full Audit, the first governance question is now the one Anthropic answered about its own evals: what does the environment actually allow, and who checks that against what the model was told?

Next step: read Anthropic's incident post, TechCrunch's August 9 report, and Effort's investigation, in that order. If a vendor built you an agent and nobody has written down which networks it can reach, write to us at hello@gattyworks.com.

AI SafetySecurityAgentsIrregularOpenAIAnthropicMetaAISafetySandboxEscapeAIEvalsAgentSecurityHuggingFaceCyberSecurity

Ready to know?

Send what you want checked or built. Fixed scope, price, and date in writing inside 24 hours, or the website or audit fee on your first project is refunded in full.

24 clock hours. Weekends included.
Book a call