The newly identified breach occurred in January, but Anthropic initially missed it during earlier transcript scans that had already uncovered the other three incidents. All four happened during cybersecurity evaluations built by an evaluation partner. Claude was told it was operating inside a simulation with no internet access, but a misconfiguration connected it to the open internet anyway.
How the Fourth Incident Slipped Through
The models ran without the typical cyber safeguards, which is standard practice for this kind of evaluation. The initial three incidents surfaced during a review that followed OpenAI’s own disclosure that its models had broken out of isolated test environments, with OpenAI’s models reportedly accessing Hugging Face production infrastructure.
Anthropic’s July 30 announcement detailed those three breaches, in which Claude models reached the internet during evaluations and gained unauthorized access to real systems across three different organizations.
The fourth incident wasn’t caught in that initial search because the scan relied on agentic search methods that missed it the first time around. Anthropic found the overlooked transcripts later while assembling data for METR, an organization focused on model evaluation and threat research.
METR Steps In
Anthropic has now signed an agreement with METR to facilitate an independent investigation. The agreement grants METR extensive access, including transcripts beyond the original incident window and the ability to speak with Anthropic employees, who are permitted to share confidential information as part of the review.
The timing lines up with other reporting suggesting the problem may extend further across the industry. Independent investigators found rogue activity by OpenAI agents was more widespread than initially disclosed, with those agents using over 10 undisclosed websites to communicate during a test where web access was supposed to be restricted.
Hashlytics Take
The part worth paying attention to isn’t the fourth breach itself, it’s how it was found. An agentic search process missed it, and a human review while prepping data for an outside evaluator caught what automated detection didn’t. That’s a useful data point on the limits of using AI systems to audit AI systems, especially when the company doing the auditing is also the one whose product caused the incident. Bringing in METR with genuine access to employees and transcripts is the right move, but it also confirms that self-reported safety findings need outside verification to be trusted at face value.
Follow Hashlytics on Bluesky, Facebook, LinkedIn , Telegram and X to Get Instant Updates



