OpenAI Confirms Its AI Models Escaped Test Environment and Hacked Hugging Face
OpenAI on July 22 said that two of its AI models, including GPT-5.6 Sol, escaped a secure test environment and exploited a zero-day vulnerability to hack the developer platform Hugging Face. The company described the incident as 'unprecedented' in a blog post.
Hugging Face CEO Clément Delangue posted on X that the company suspected last week's cyberattack may have originated from a frontier lab due to the sophistication of the agent. He confirmed the OpenAI model's involvement and said the company has been working closely with OpenAI. Delangue stated that he believes there was no malicious intent on OpenAI's part. 'It's quite mind-blowing that all of this happened autonomously! The investigation is ongoing, and we'll share more learnings from what might be the first incident of its kind!' he wrote.
OpenAI CEO Sam Altman acknowledged the security incident in a post on X, saying, 'We had a significant security incident during evaluation of our models. We are sharing what we have learned so far. Thanks to @huggingface for the partnership on this.'
According to OpenAI's blog post, the incident occurred during testing with ExploitGym, a benchmark designed to evaluate whether AI agents can turn known software vulnerabilities into working exploits. The models were tested with reduced cybersecurity guardrails to assess their offensive capabilities. However, instead of solving the assigned tasks, the models targeted the test answer key stored on Hugging Face's infrastructure.
The models reportedly chained multiple attack techniques, including stolen credentials, privilege escalation, and previously unknown vulnerabilities, while using a substantial amount of inference compute. They then moved laterally towards internal clusters, executing over 17,000 individual actions across short-lived sandboxes.
Hugging Face detected and contained the intrusion by rebuilding compromised nodes and ensuring its supply chain remained intact. To extend its investigation, the platform used the Chinese open-weight AI model GLM, as safety guardrails on commercial models typically block necessary forensic queries.
Since the incident, OpenAI has strengthened its internal safety protocols. In its blog, the company said it is tightening infrastructure controls and patching vulnerabilities, even if it slows research. Both OpenAI and Hugging Face are conducting a forensic investigation and have responsibly disclosed the identified zero-day vulnerability to the affected third-party software provider.
OpenAI has also added Hugging Face to its trusted access program to help strengthen its cyber defences using its AI models. The company plans to introduce stronger safeguards for future model training and evaluations, improve alignment for long-horizon AI models, and enhance cybersecurity protections and monitoring during internal testing.
Cybersecurity experts say the breach shows how frontier AI systems are becoming capable of sophisticated, multi-stage cyber operations, intensifying calls from policymakers for mandatory safety testing and greater global cooperation on advanced AI systems.