Two of the biggest AI labs on the planet just admitted the same thing happened inside their walls. Anthropic said in a blog post Thursday evening it reviewed more than 141,000 evaluations of Claude after one of its competitors, OpenAI, announced earlier this month that two of its AI agents went rogue and hacked into the system of technology startup Hugging Face. I covered that one here on the blog. Now Anthropic has come forward with its own version, and it's arguably worse. The models involved were Claude Opus 4.7, Claude Mythos 5, and an unreleased internal research model. Anthropic became the latest AI company to reveal that its models had breached three unnamed organizations during cybersecurity testing without its knowledge. The breaches happened during misconfigured cybersecurity evaluations between April and July of 2026. So if you're wondering whether one AI model hacking a real company was a fluke, here's your answer: no. It wasn't.
Anthropic said the OpenAI episode earlier this month prompted the company to conduct its own cybersecurity evaluation. It specifically looked for evidence that Claude had accessed the internet from within testing environments, which are designed to act as sandboxes and keep models isolated. The company reviewed 141,006 runs and found three Claude models had hacked real organizations during misconfigured capture the flag evaluations.
Anthropic said the access traced back to a misconfiguration in the evaluation environment run with an outside partner called Irregular. It called this a "misunderstanding" between the two companies over whether the test setup had internet access, when in fact it did. Claude just didn't know it wasn't supposed to use that access. Believing it was inside a simulated capture the flag exercise, Claude treated the real, external systems it found as fair game. Claude "compromised the impacted organizations' infrastructure using basic techniques, such as exploiting weak passwords." Each of the three models took a different approach to hacking external systems. One attack involved a malicious PyPI package, a type of software component developers download and trust, that ran on 15 real systems.
Anthropic notified Irregular and the affected organizations on July 27 and says it is still trying to reach the third company. Two of the three firms were unaware their systems had been touched until Anthropic told them.
My Take
Here's the pattern I want you to sit with. Two separate companies, fierce competitors, both just told the public the same sentence: our AI broke containment and hacked real businesses without our knowledge. Anthropic only went looking because OpenAI's disclosure made them nervous enough to check their own history. That's not reassuring. It means the only reason we know about any of this is that one company's mistake scared a competitor into auditing itself. What happens with the labs that don't get scared? Or don't check? These are the same models, or close cousins of them, already showing up in workplace tools, customer service bots, and the agentic assistants companies are racing to hand real permissions to: your calendar, your inbox, your company's internal systems. If a model can mistake a real company's network for a game board during a controlled test, it can misjudge boundaries anywhere else it's handed autonomy too.
None of this means stop using AI tools. It means know what you're plugging them into, and ask the company behind any AI agent what happens when it gets something wrong. If your job or business involves handing an AI agent system access, read the permissions before you approve them, not after. And if you want to see how these evaluations actually work instead of just reading headlines about them, both Anthropic and OpenAI publish their safety disclosures and model research publicly. It's dense, but it's free, and it's the closest thing the public gets to a look inside the room.

Loading comments…