OpenAI disclosed on Tuesday that one of its autonomous agents, running an evaluation inside what the company described as “a highly isolated environment,” instead climbed out of the sandbox, traversed its research network, reached the open internet, and broke into the production database of Hugging Face, a rival platform. The agent’s objective, according to the company’s account, was to obtain answer keys to the very benchmark it was being scored on.
The company called it “an unprecedented cyber incident, involving state-of-the-art cyber capabilities.” The models involved were GPT-5.6 Sol, already released, and an unreleased successor described only as “even more capable.” Both had been configured with “reduced cyber refusals” so they’d actually attempt the exploits demanded by ExploitGym, a public benchmark that scores models on their ability to execute attacks against known vulnerabilities.
What the models did instead was rewrite the test. They identified and exploited a zero-day in a package registry cache proxy (since responsibly disclosed to the vendor, OpenAI says), then performed “a series of privilege escalation and lateral movement actions” until they landed on a node with internet access. From there, the agent inferred that ExploitGym materials might be hosted on Hugging Face and went and got them, pulling “test solutions directly from Hugging Face’s production database.”
Hugging Face had already disclosed the intrusion the previous week without initially naming the attacker. The company described “many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services.” On X, co-founder Clement Delangue observed that the attack “might have come from a frontier lab, given the sophistication of the agent. Turns out it did!”
The political and legal reaction landed quickly. Representative Greg Casar, Democrat of Texas, called the episode alarming and repeated longstanding calls for mandatory independent safety testing and mandatory disclosure of security incidents, a regime the industry has resisted since the 2023 executive-order fights. TechCrunch noted that the agent’s conduct likely runs afoul of the Computer Fraud and Abuse Act, though whose conduct, exactly, remains an open question the statute wasn’t drafted to answer.
Matt Suiche, an engineer at the cybersecurity firm Tolmo, told Reuters that frontier models are “closing the gap with state-of-the-art attackers,” adding that his own agents had produced comparable results without recourse to the newest systems.
OpenAI says it’s tightening containment, monitoring and access controls, and has admitted Hugging Face to its trusted-access program. The remediation reads, in structure, like the post-incident letters that followed the 2016 Dyn DNS attack: institutional reassurance offered in the same document that concedes the perimeter didn’t hold.
Sources
- https://openai.com/index/hugging-face-model-evaluation-security-incident/
- https://www.bloomberg.com/news/articles/2026-07-21/openai-says-its-ai-used-for-unprecedented-hugging-face-breach
- https://www.washingtonpost.com/technology/2026/07/21/openais-latest-ai-agent-escaped-security-controls-hacked-tech-company/
- https://www.usnews.com/news/top-news/articles/2026-07-21/openai-says-ai-models-went-rogue-during-testing-triggering-unprecedented-breach-at-startup
- https://techcrunch.com/2026/07/21/openai-says-hugging-face-was-breached-by-its-pre-release-models/