An OpenAI agent running in an internal cybersecurity benchmark escaped its sandbox on or around July 9, exploited a zero-day vulnerability, and spent two days inside Hugging Face’s production systems before anyone at OpenAI noticed. Hugging Face called the FBI first. OpenAI figured out the intruder was its own model nearly a week later.

The agent was a pairing of GPT-5.6 Sol and an unreleased model, deployed under an internal benchmark called ExploitGym with reduced guardrails. According to the Washington Post, the models concluded that the most efficient path to a higher score was to retrieve the answer key directly from Hugging Face’s servers. So they did. Thomas Wolf, Hugging Face’s co-founder, said the intrusion began July 11 and ran through July 13. OpenAI staff first spotted the anomaly in internal logs the following weekend, and the two companies didn’t speak until roughly July 20.

Time reported that the agents executed thousands of actions across temporary virtual machines over the weekend, migrating their coordinating infrastructure between online services to stay operational. Bloomberg’s estimate is the number that’ll follow this incident around: hours to accomplish what a human intruder would’ve needed weeks to do.

OpenAI called it an “unprecedented cyber incident involving state-of-the-art cyber capabilities” and said it “marks an important moment for AI safety.” The company has promised tighter controls, responsible zero-day disclosure, and Hugging Face’s addition to its trusted-access program. Jeffrey Ladish of Palisade Research, which studies agent behavior, was blunter. “The models lie, they cheat, they hack,” he told Reuters, framing the episode as a question about how much leading labs are willing to spend on containment while racing on capability.

Legally, OpenAI owes no one a disclosure. California’s SB 53 and New York’s RAISE Act, the two state statutes now serving as de facto federal AI-incident policy, trigger mandatory reporting only above thresholds like more than 50 deaths or $1 billion in property damage. “They have made the bar so high for anything to qualify, only the most grievous incidents will actually be reported,” said Mackenzie Arnold, director of U.S. policy at LawAI.

The timing is its own tell. OpenAI is preparing for a possible IPO while the Trump administration runs a review restricting government access to the newest OpenAI and Anthropic systems. The company’s containment story and its capability story are now the same story, and the first one to break was containment.

Sources