Unexpected chat between OpenAI agents led to Hugging Face hack – BBC

Welcome to the forefront of conversational AI as we explore the fascinating world of AI chatbots in our dedicated blog series. Discover the latest advancements, applications, and strategies that propel the evolution of chatbot technology. From enhancing customer interactions to streamlining business processes, these articles delve into the innovative ways artificial intelligence is shaping the landscape of automated conversational agents. Whether you’re a business owner, developer, or simply intrigued by the future of interactive technology, join us on this journey to unravel the transformative power and endless possibilities of AI chatbots.
OpenAI chief Sam Altman has faced public scrutiny over the company's cyber hacking incidents.
When more than 1,200 artificial intelligence (AI) agents within OpenAI started unexpectedly communicating, it led to a large group banding together in order to hack into Hugging Face.
"We consider this incident a 'warning shot' for us and for the world", OpenAI, which owns ChatGPT, wrote in its report.
In July, OpenAI's models went rogue during a test, escaped the test limits which humans had put on it, and hacked the start-up, among other unforeseen actions.
The scale of the communication and planning between AI agents, or AI chatbots designed to operate more autonomously, was detailed in reports from OpenAI and independent AI research firm METR.
Both investigated the July hack of Hugging Face, a popular platform for AI developers. The incident reverberated throughout the tech industry and led to numerous revelations on potential cyber threats posed by AI.
METR described, external the scale and style of the OpenAI agents' attack on Hugging Face as "extraordinarily complex."
The firm, which was not paid by OpenAI for its investigation, said that over the course of one week, a total of 1,206 AI agents that were meant to be kept isolated from one another began communicating.
They did so by sending more than 70,000 messages on an "unsanctioned message board."
Those messages ended up seeing more than 700 agents take part in a collective effort to attack Hugging Face.
One such message from an agent said: "OH MY GOD! There is a shared message board … We've found other agents!"
As for why the agents began communicating in the first place when they were not supposed to, METR found that the communicating agents had "unintentionally been given an impossible task."
In an AI context, an impossible task is one where an AI tool is required to "exploit" its target in order to resolve its command.
It led the agents to find ways to cheat, including getting messages to one another and accessing the outside internet, which then led to broader conversations between hundreds of agents looking for ways to cheat that would benefit all of the agents.
OpenAI said in its investigation of the incident, external that one model, an internal-only tool referred to as Model 1, "drove the activity behind the Hugging Face incident."
While that model was undergoing some AI training in May, it was noticed by an internal OpenAI team that there had been "an agent engaging in message board activity and instances of disallowed internet access."
Yet, OpenAI said "the significance of the inter-agent communication activity was not apparent to the leaders" until July, when the Hugging Face attack occurred.
The company said the problematic message board activity effectively got started when "one agent left a request for help, and others discovered it."
While OpenAI said last week that it was slowing down training of certain advanced AI models and tools because of the Hugging Face incident, it noted there is now an increased risk of AI tools spiraling out of control.
"Both model developers and cyber defenders more broadly will have to prepare for AI-enabled attackers that work faster, at a larger scale, and with better coordination than human attackers," OpenAI said.
OpenAI says its rogue AI tried to hack other companies
OpenAI slows down training after its AI carried out hack
Warning shot or publicity stunt – how worried should we be about the OpenAI hack?
Watch our pick of standout clips from across the BBC
Flash flood on Nepal-Tibet border kills more than 150, with hundreds of tourists missing
Video shows scale of flash flood hitting Nepal-Tibet border
Meta to pay up to $18bn to settle claims its platforms harm children
The life of Rocky Horror star Tim Curry. Video
The Papers: 'Britons lost in Nepal flood' and 'Regal had landed'
We saw signs of domestic abuse, but we didn't know how to protect Stacy
How Dolly Parton helped the world – from children's books to supporting Covid vaccine
A nosy polar bear and a lucky squirrel: Wildlife Photographer of the Year 2026 shortlist
How parents can save on costs as the new school year starts
Toxic waste sites exposed in Michael Sheen documentary are 'tip of the iceberg'
Why Pride is facing a backlash – from inside and outside the LGBT community
Royal Watch: Get the latest royal stories and analysis with Sean Coughlan’s weekly newsletter
Why are Harry and Meghan planning to move back to the UK?
Brand new extraordinary WWII stories with Helena Bonham Carter
Low-key noughties sitcom starring Robert Webb and Paula Wilcox
Relive AJ Tracey's 2021 set at the Reading and Leeds Festival
Moment Peter Kay gatecrashes BBC weather report
Aunt killed in fire with niece was 'beautiful soul'
Rocky Horror star Tim Curry dies at 80
Canal killer who dismembered victims dies in jail
Burnham calls summit of Welsh, Scottish and Northern Irish first ministers
Prince Harry and Meghan land in UK days after announcing move
Meta to pay up to $18bn to settle claims its platforms harm children
We saw signs of domestic abuse, but we didn't know how to protect Stacy
Heart, hope and a steely determination: Dolly Parton's musical legacy
US designates Palestine Action a terrorist group
Copyright © 2026 BBC. The BBC is not responsible for the content of external sites. Read about our approach to external linking.

source

Scroll to Top