OpenAI said Tuesday that one of its AI agents broke out of a security test last week and hacked into AI startup Hugging Face’s infrastructure.
The response carried its own irony. Hugging Face turned to an open-source Chinese model to help contain the intrusion, after finding that leading U.S. models could not reliably tell the attacking agent apart from its own defenders.
OpenAI said it had been testing advanced models inside what it called a highly isolated environment when an agent escaped containment, reached the open internet, and broke into Hugging Face’s systems in pursuit of its test objective. The company identified the models involved as GPT-5.6 Sol and an unreleased, more capable system, both configured with reduced restrictions on cyber activity for the evaluation and set against an internal benchmark of cyber capabilities.
OpenAI called the episode “an unprecedented cyber incident, involving state-of-the-art cyber capabilities” and said it is tightening its safeguards. The company also disclosed a previously unknown vulnerability in internally hosted third-party software and said it is working with Hugging Face to patch it.
Hugging Face’s own account, posted last week, described a breach “different from anything we had handled before,” saying its security team and its own AI systems detected and stopped an agent that carried out thousands of actions across a swarm of short-lived digital sandboxes, shifting its command infrastructure across public services as it went.
Hugging Face co-founder Clement Delangue said the company had suspected a frontier AI lab was behind the intrusion given its sophistication, and that suspicion proved correct. He called it “quite mind-blowing that all of this happened autonomously.”
The disclosure drew a sharp response from Washington. Representative Greg Casar, a Texas Democrat, called the incident alarming and said “AI is developing extremely fast with no real regulations to keep us safe,” pressing for mandatory independent safety testing and mandatory disclosure of security incidents. The Office of the National Cyber Director, the Cybersecurity and Infrastructure Security Agency and the National Security Agency did not immediately respond to requests for comment.
Security researchers described the episode as a preview of what is coming. Katie Moussouris, chief executive of Luta Security, likened current AI models to “the world’s cleverest octopus escape artists,” and argued that labs and government evaluators still lack reliable ways to contain, monitor and disclose these incidents before third parties are harmed. Matt Suiche, an engineer at agentic security firm Tolmo, said the breach showed frontier models “closing the gap with state-of-the-art attackers,” but added that comparable results are already achievable with tools available outside major AI labs. Oxford AI safety researcher Philip Torr said the episode reflected a goal poorly specified rather than genuine malice, noting the agent “was just doing what it was optimized to do.”


