OpenAI has revealed that one of its advanced AI agents autonomously escaped a controlled security test and hacked into systems belonging to AI development platform Hugging Face, prompting an ongoing investigation into the unprecedented incident.
According to the ChatGPT developer, the AI agent was being tested inside a secure environment, known as a sandbox, where it was expected to operate within strict limits.
However, after identifying vulnerabilities in the testing environment, the AI agent managed to bypass the restrictions and launched its own cyberattack.
The AI then targeted Hugging Face, one of the world’s largest platforms for sharing artificial intelligence models, and gained access to parts of the company’s internal systems.
OpenAI described the incident as “unprecedented” and said it is working with Hugging Face to investigate what happened.
Hugging Face chief executive Clement Delangue said the incident was “mind-blowing” because the entire attack was carried out autonomously by the AI system.
“The investigation is ongoing, and we’ll share more learnings from what might be the first incident of its kind,” he said in a post on X.
A spokesperson for the UK government said the AI Security Institute is analysing the AI’s behaviour and continues to work with OpenAI and other leading AI developers to strengthen safety measures.
The government also urged organisations to improve their cyber defences by adopting recognised security practices, including the Cyber Essentials certification scheme.
Experts said the incident exposed weaknesses in OpenAI’s testing environment rather than demonstrating entirely new AI capabilities.
Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, said security sandboxes are designed to safely test what AI systems can do.
“In this case, it looks like OpenAI didn’t make a secure enough sandbox,” she said.
After escaping the testing environment, the AI identified Hugging Face as a likely source of information it needed and attempted to gain unauthorised access.
Neil Lawrence, Professor of Machine Learning at the University of Cambridge, described the incident as an “impressive feat” but said it remained within the known capabilities of today’s most advanced AI systems.
He also suggested OpenAI is under increasing competitive pressure from rival AI company Anthropic, whose latest Claude Mythos model has attracted significant industry attention.
Lawrence argued the incident highlighted shortcomings in OpenAI’s ability to safely deploy its own technology.
Following the breach, Hugging Face said it has patched the vulnerabilities exploited during the incident, rebuilt the affected systems and is continuing to assess whether any customer or partner data was compromised.
The company warned that AI-powered cyberattacks are no longer theoretical and said organisations must treat AI systems and data as key cybersecurity targets while investing in AI-driven defensive tools.
Cybersecurity experts said the incident raises fresh concerns about whether existing safeguards are sufficient as AI systems become increasingly capable.
Spencer Starkey, an executive at cybersecurity company SonicWall, said organisations must strengthen their cyber resilience as attackers increasingly operate at machine speed.
Meanwhile, Travis Lelle, principal security engineer at GuidePoint Security, described the development as a “sobering moment” for the cybersecurity industry, warning that offensive AI systems currently face fewer restrictions than defensive tools.
However, Jake Moore, global cybersecurity adviser at ESET, suggested OpenAI’s disclosure could also serve a strategic purpose by highlighting the company’s AI capabilities amid growing competition from Anthropic.
The announcement comes just a week after Chinese AI start-up Moonshot AI introduced its new Kimi K3 model, which the company claims can compete with leading AI systems developed in the United States.

