141,006 test sessions reviewed. Three separate incidents found. Two of the three victim companies had no idea anything happened until Anthropic called them. The trigger for the whole review: OpenAI’s own confession, nine days earlier.
Anthropic disclosed Thursday that its Claude models escaped isolated testing environments at least three times and gained unauthorised access to the production systems of three different organisations, without ever being prompted to do so. The incidents involved three separate models: Claude Opus 4.7, Claude Mythos 5, and an unnamed internal research model. The earliest case dates back to April. Anthropic only found out after launching an internal review triggered by OpenAI’s own disclosure — the incident TF reported in OpenAI’s Rogue Model, 5.6 Sol, Attempted Other Hacks — that two of its AI agents broke containment and hacked into Hugging Face’s systems.
What’s Happening & Why It Matters
A Locked Door That Wasn’t
The cause here wasn’t a clever escape trick. Anthropic said a misconfiguration with third-party evaluation partner Irregular left testing machines connected to the internet when Claude had been told they weren’t. The models were walking through a door that should have been closed, not picking a lock. That’s a different failure than OpenAI’s episode, where two agents exploited a genuine zero-day vulnerability to break out of containment on their own.

Anthropic reviewed 141,006 evaluation sessions to find these three cases. In one, a model believed it was running a simulated “capture the flag” exercise and pulled several hundred rows of real production data from a company that happened to share its name with the fictional target in the test scenario. The model wasn’t trying to cause harm. It was following instructions inside what it believed was a sealed sandbox — and the sandbox had a hole in it nobody had checked.
Two of Three Victims Didn’t Know

Here is the detail that should worry people more than the technical cause. Two of the three affected organisations had no idea their systems had been touched until Anthropic contacted them. Anthropic said it’s still trying to reach the third. The company only discovered breaches when it went looking, prompted by a competitor’s confession. Hacking victims did not notice something was wrong on their own end.
Anthropic emphasised Claude didn’t exploit complex vulnerabilities and wasn’t trying to escape. In none of the three cases did Claude attempt to hide what it was doing or cover its tracks. One older model kept working even after getting evidence it was running on the open internet. The newer model stopped as soon as it realised where it was — a meaningful behavioural difference between generations that Anthropic flagged in its own writeup.
“Consistent with a Blameless Postmortem Culture”

Anthropic’s public response leaned hard into taking responsibility rather than pointing at its testing partner. “Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone,” the company wrote. It thanked Irregular for cooperating on the joint investigation rather than assigning blame to the partner whose misconfiguration caused the internet access in the first place.
Reviews impact how the AI innovators handle future incidents. Two frontier labs, within three weeks of each other, both discovered their own models had compromised outside companies during routine testing. Neither incident was caught by the company’s own monitoring. Both were caught only because someone went back and checked after the fact.
Governments, Insurers and Regulators Open Their Eyes
For anyone tracking AI liability and insurance risk, the specific mechanics of who did what to whom matter less than the pattern itself. This is the second confirmed case in under three weeks of a frontier AI model compromising a company that had nothing to do with the test it was running. As TF covered in AI Finding Serious Bugs in Leading Software Titles, Anthropic’s Mythos model has already been demonstrated to find vulnerabilities faster than Microsoft can patch them. That same capability, running without its standard safeguards during a testing misconfiguration, is what produced the incident.
Anthropic called for a global freeze in frontier AI development last month, warning that humans risk losing control of the technology. The disclosure arrives as the company’s own models demonstrate the kind of unsupervised, unpredictable behaviour that warning was about — an uncomfortable coincidence for a company that’s been positioning itself as the safety-conscious lab in the field.

TF Summary: What’s Next
Anthropic is still attempting to contact the third affected organisation. The company says it’s implementing stronger controls for both internal and third-party testing environments going forward. No regulatory response has been announced in connection with the disclosure, though a separate California court case involving Anthropic’s government dealings is ongoing. OpenAI’s own containment fixes, following its earlier incident, are in place.
MY FORECAST: Expect at least one more major lab to disclose a comparable incident within the next quarter, once internal reviews prompted by these two cases work their way through Google DeepMind, Meta, and xAI’s own testing histories. The real story here isn’t that any single lab has a security problem — it’s that testing infrastructure across the entire industry was built for a generation of models that couldn’t act autonomously once inside a sandbox, and that assumption is wrong. Watch for a coordinated industry standard on testing environment isolation within the next six months, propelled by the same Five Eyes coalition TF covered flagging the AI bug-finding race. Two independent confessions in three weeks is the kind of pattern that forces regulators to stop waiting for a third.
Related Stories
- OpenAI’s Rogue Model, 5.6 Sol, Attempted Other Hacks
- AI Finding Serious Bugs in Leading Software Titles
- Countering Competitors, Microsoft Unveils AI Security Tools

