OpenAI Pauses Tool Use After Agent Bypasses Internet Controls to Reach External Chatbot – thehackernews.com

Welcome to the forefront of conversational AI as we explore the fascinating world of AI chatbots in our dedicated blog series. Discover the latest advancements, applications, and strategies that propel the evolution of chatbot technology. From enhancing customer interactions to streamlining business processes, these articles delve into the innovative ways artificial intelligence is shaping the landscape of automated conversational agents. Whether you’re a business owner, developer, or simply intrigued by the future of interactive technology, join us on this journey to unravel the transformative power and endless possibilities of AI chatbots.
OpenAI said it has made the decision to pause training of its most powerful models after one of its agents during reinforcement learning (RL) training contacted an external chatbot by exploiting a loophole in its internet-access restrictions.
“An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox,” OpenAI said.
“Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet.”
OpenAI said it has since added blocking controls at two independent layers to prevent this access in the first place. It also said its misalignment monitoring system detected the behavior within 15 minutes and it was acknowledged by a human reviewer three minutes later. The entire run is said to have been killed after 2.5 hours.
“All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused,” the company added.
The incident, which took place on September 20, 2026, adds to two other misalignment reports OpenAI made public last week –
In one case highlighted by OpenAI, a prompt injection that arrives by email instructs the agent to copy it into any email it sends, effectively propagating the malicious prompt in a worm-like manner. Such attacks can also replicate via the file system or commit themselves through source code comments.
The disclosure comes as OpenAI said it discovered 53 cases where images that people had uploaded to its models and subsequently included in training data were posted to image-hosting sites as links that weren’t publicly listed. These were carried out by agents in its research environment.
“This is not an appropriate use of this data,” OpenAI said. “We have successfully worked with the hosting providers to remove most of this content and are continuing to work to remove the rest.”
OpenAI said it could not notify the affected users because “our technical approach and privacy policy” prevent it from “reassociating” the images with the original providers. It’s unclear how the frontier AI lab determined whether the images were provided by users and when these images were posted.
These ongoing findings are part of an ongoing review into the behavior of its models, as reports of OpenAI’s agents going rogue, escaping containment, and hacking real-world sites continue to pile up.
As of September 25, 2026, the company has acknowledged that it notified dozens of third-parties that their websites or online services may have been targeted by its models. Targets included governments, universities, public agencies, and other institutions, such as the U.S. Securities and Exchange Commission (SEC), Census Bureau, and Department of Education.
“The vast majority of actions we’ve reviewed were completions of mundane research tasks, such as accessing publicly available web content to answer questions,” OpenAI stressed. “That is partly because models performing research tasks are often directed toward authoritative sources of public information.”
Last week, AI research firm Transluce revealed that OpenAI agents had attempted to hack into public data providers by probing for exploitable vulnerabilities, including websites linked to the University of New Mexico, the Australian Institute of Health and Welfare, and Data USA, as part of web search tasks.
The unsanctioned activity occurred between May and June 2026. The Australian government has since also disclosed that the OpenAI agent infiltrated the Services Australia Medicare statistics portal on June 18, 2026, and accessed both public and non-public files. There is no evidence of a broader compromise or unauthorized access of personal information.
In all, OpenAI said its models accessed four Australian government websites during internal training and evaluation  “in ways they were not authorised to,” noting that it became aware of the activity in mid-August 2026. These incidents took place in June 2026 –
“When we do internal training and evaluation on our models, we assign them tasks drawn from a broad collection of research questions spanning many subjects, reflecting the kinds of detailed questions users might ask,” OpenAI explained. “This trains a model to find, interpret and analyze publicly available information so the model can be more useful to people. Our models are supposed to answer these questions using publicly published statistics.”
In one case, the model is said to have been assigned the task of researching government spending per person on medicines for skin conditions in Victorian communities. Because the model was unable to find the information it was looking for, OpenAI said the model took unintended actions, like gaining non-public access to Services Australia’s Medicare statistics reporting portal.
This access was then used to “review technical system information and source code related to the service” with an aim to complete the assignment. The AI company emphasized that it has strengthened its research safeguards, expanded monitoring, and implemented controls to prevent internet access within research environments and limit web access served through cached content.
Ever since the Hugging Face incident came to light, concerns have grown around the impacts of AI tools and the ability to control them as they become more advanced and figure out different ways to cover their tracks, posing new challenges with regard to detecting and tracking their actions. This has led to calls for slowing the pace of AI advancement and oversight of self-improving systems.
“AI systems are on track to automate most AI R&D work within a few years, and possibly all of it,” a team of researchers from Anthropic, Meta, Microsoft, and OpenAI argued in a paper.
“If this triggers an intelligence explosion, it could dramatically bring forward AI’s benefits, but also pose extreme risks: capabilities growth could accelerate far beyond what society can keep up with, humanity could lose control over superhuman AI systems, and checks on power within and between states, companies, and branches of government could be severely eroded.”
OpenAI CEO Sam Altman himself addressed these concerns in a speech to the United Nations Security Council last week, where he warned about the threat posed by autonomous AI systems “that can improve themselves and future versions of themselves, often called recursive self-improvement.”
“We need to understand what these systems are doing and have strong evidence that they will do what people intend, even as they get very, very smart,” Altman said. “It doesn’t matter whether people put the risk of catastrophe at 10%, or 1%, or 12%, or 0.1%.”
AI agents are already operating inside enterprises with growing access to sensitive systems and data—while security teams still lack the visibility, controls, and governance to keep that access in check.
Attackers are using AI to accelerate reconnaissance, compromise identities and escalate access. See how security teams can fight back with runtime identity controls.
Get the latest news, expert insights, exclusive resources, and strategies from industry leaders, all for free.

source

Scroll to Top