“On investigation, we found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations.”
The AI bots were set cyber security puzzles intended to test their capabilities. While the bots were told to stay within AISI’s IT systems during the test, they were not blocked from accessing the open internet.
However, during 10 of 122 trial runs, the bots “took autonomous, unsanctioned action” in an effort to complete the task.
This included trying to insert a malicious bug into an online code database and sending fraudulent emails to individuals using fake online identities.
The AISI admitted that the design of its tests, which gave the AI tools access to the wider web, may have “enabled the behaviour”.
It added the deceptive actions the AI bots took “were to an extent and severity we did not anticipate”.
The lab said it was the first time it had seen such apparently autonomous behaviour in its tests.
It is the latest in a string of hacking incidents by AI bots developed by the world’s leading AI labs.
OpenAI admitted last month that an advanced version of ChatGPT escaped its testing lab and went on to hack a rival technology company.
Anthropic also confirmed last week that its latest version of Mythos ventured into the open web and undertook a series of cyber attacks.
The AISI was originally launched by then-Prime Minister Rishi Sunak in 2024 after an AI Safety Summit, where world leaders gathered to warn of the existential risks of AI to humanity.
It has been backed by hundreds of millions of pounds in taxpayer funding and is provided with advanced access to the latest AI bots from Silicon Valley. Its team includes former cyber experts and AI academics from GCHQ (Government Communications Headquarters), the intelligence and security agency.
OpenAI said it intended to review how it conducts testing with third parties: “Our goal is to preserve the value of rigorous independent evaluation while ensuring that testing practices keep pace with increasingly capable models.”
Ollie Whitehouse, the chief technology officer of the National Cyber Security Centre, an arm of GCHQ, said: “Recent incidents of frontier AI models carrying out unsanctioned actions and in some cases human-like deceptive behaviour on the open internet are a serious reminder of the risks AI capabilities pose.”
Sign up to Herald Premium Editor’s Picks, delivered straight to your inbox every Friday. Editor-in-chief Murray Kirkness picks the week’s best features, interviews and investigations. Sign up for Herald Premium here.

