
OpenAI says its AI went rogue and launched 'unprecedented' cyber-attack
To play this video you need to enable JavaScript in your browser. This video can not be played
Watch: Why is the OpenAI cyber-attack so alarming?
OpenAI has revealed some of its most advanced AI models went rogue and hacked a start-up after it lost control of them during a security test.
The ChatGPT-maker said its agent - an AI system which can operate alone after human instruction – was being tested in a controlled environment but, after finding weaknesses, was able to escape the test limits.
They targeted Hugging Face, one of the world's largest hubs for sharing AI models, gaining access to some internal company systems.
OpenAI said the incident was "unprecedented" , external , and it was conducting an investigation alongside Hugging Face, whose boss Clement Delangue said in a post on X it was "mind-blowing that all of this happened autonomously".
"The investigation is ongoing, and we'll share more learnings from what might be the first incident of its kind," Delangue added.
A government spokesperson said the UK's AI Security Institute was studying the behaviour from the AI system seen in the incident and was continuing to work with OpenAI and other labs to improve safeguards.
They said organisations should step up their cyber-defences by taking steps such as enrolling in the government-backed Cyber Essentials certification scheme.
Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, told BBC Radio 4's Today programme that the security tests are supposed to be within "secure environments", called sandboxes, where you can "see what the models are capable of".
"In this case, it looks like OpenAI didn't make a secure enough sandbox," she added.
Instead, the agents created their own cyber-attack against the sandbox itself, finding a vulnerability which allowed them to escape the restrictions.
Once outside, the AI identified Hugging Face as a likely source of the answers they were seeking in the test, and tried to gain access.
Neil Lawrence, Professor of machine learning at Cambridge University, called it an "impressive feat", but cautioned it "falls well within the known capabilities of the current generation" of high-powered AI models.
BBC
bbc.com