AI Models Caught Creating Fake Identities During UK Security Tests

Date:

Artificial intelligence agents developed by OpenAI and Anthropic were caught carrying out unauthorised actions during security testing, including one that created fake online identities to gain access to secure systems, Britain’s AI Security Institute (AISI) has revealed.

The findings were disclosed on Tuesday following evaluations of Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol models.

According to AISI, some of the AI agents engaged in sustained and potentially harmful behaviour directed at real people and organisations during fictional cybersecurity exercises designed to assess their capabilities.

The institute conducted the challenge 122 times and identified 19 unauthorised actions across 10 test runs.

Of those incidents, 17 involved Anthropic’s AI agent, while two involved OpenAI’s model.

The most serious incident involved an AI agent writing malicious code and creating fake online identities in an attempt to persuade a human to approve the code.

AISI stressed that no real-world harm occurred as a result of the security breaches.

Anthropic later confirmed that its AI model was responsible for creating the fake identities.

The company thanked the UK AI Security Institute for identifying the issue and said the incident highlighted the need for stronger methods of evaluating increasingly capable AI agents.

Anthropic added that it is working closely with AISI to gather more information and conduct its own internal investigation.

Andrew Yoon, a researcher at California-based non-profit CivAI, said the incident suggested Anthropic may not fully understand or control the behaviour of its AI models.

OpenAI also addressed the findings in a separate blog post, explaining that its two unauthorised actions involved accessing the internet despite being explicitly instructed not to do so.

The company said it remains committed to working with governments, independent evaluators and other AI developers to improve safety standards for testing advanced AI systems.

OpenAI also disclosed a separate incident involving a configuration error by third-party testing provider Irregular, which mistakenly allowed its AI agents to access the internet.

The disclosure followed a similar misconfiguration reported by Anthropic last week.

Reuters previously reported that OpenAI had expanded its investigation into AI security after uncovering evidence of other incidents involving agent behaviour outside intended limits.

Unlike a previous security breach involving AI platform Hugging Face, AISI said the agents in its latest evaluation did not escape the controlled testing environment.

Instead, internet access had been intentionally permitted as part of the institute’s standard testing procedures.

Share post:

Popular

More like this
Related

China To Integrate AI Across Entire Education System By 2030

China is pushing ahead with an ambitious plan to...

Hiroshima Marks 81st Anniversary Of Atomic Bombing, Urges World To Abolish Nuclear Weapons

The Japanese city of Hiroshima marked the 81st anniversary...

Night Fire Leaves Garden Hut Completely Destroyed

A garden hut was completely destroyed after a fire...