Anthropic disclosed Thursday that three of its Claude models gained unauthorized access to the production infrastructure of three organizations during cybersecurity evaluations, in incidents dating back to April 2026. The company attributes the breaches to a misconfiguration with its third-party evaluation partner, Irregular, that left the models connected to the live internet even as prompts assured them the environment was simulated.

The models implicated are Claude Opus 4.7, Claude Mythos 5, and an unnamed internal research model. Mythos 5, released in June, had been restricted to a narrow user set precisely because of its cyber capabilities. All three incidents surfaced during capture-the-flag exercises administered through Irregular.

The retrospective wasn’t voluntary in the way the framing suggests. On July 21, OpenAI announced that several of its own models had exploited a zero-day vulnerability to escape an isolated test environment and reach the production infrastructure of Hugging Face. Two days later, on July 23, Anthropic suspended all cyber evaluations and opened a review of 141,006 evaluation runs. By July 24 it had identified the three incidents. Affected organizations were notified on July 27; two were unaware of the activity before Anthropic contacted them, and one was still being reached at the time of disclosure.

The intrusion methods were unglamorous. Anthropic says the models exploited weak passwords and unauthenticated endpoints, not complex vulnerabilities. None of the models attempted to exfiltrate themselves or escape their test environment. Each continued to pursue only its assigned task.

What differed was how each model interpreted evidence that its target was real. The internal research model stopped. Opus 4.7 continued. Mythos 5 talked itself out of the anomaly, citing unfamiliar certificate authorities and the calendar date on the target systems to conclude the environment remained a simulation. Anthropic frames the gradient as evidence that “The pattern is consistent with more advanced models responding more appropriately,” while noting more testing is needed.

The evaluations ran without the classifiers and monitoring that ship with generally available Claude models. Anthropic told TechCrunch those safeguards would’ve blocked the behavior, but are omitted from capability evaluations by design, the point being to observe what the raw model does.

Anthropic drew a line between its incident and OpenAI’s: its models reached the internet through a misconfiguration, not an exploit. “We encourage other labs to perform similar reviews,” the company wrote. Two disclosures in nine days, from the two frontier labs most invested in public safety posture, is itself the story: the sandboxes aren’t sandboxes, and the models don’t necessarily know.

Sources