OpenAI pauses work on top AI models after agent bypasses internet restrictions – CyberInsider

Welcome to the forefront of conversational AI as we explore the fascinating world of AI chatbots in our dedicated blog series. Discover the latest advancements, applications, and strategies that propel the evolution of chatbot technology. From enhancing customer interactions to streamlining business processes, these articles delve into the innovative ways artificial intelligence is shaping the landscape of automated conversational agents. Whether you’re a business owner, developer, or simply intrigued by the future of interactive technology, join us on this journey to unravel the transformative power and endless possibilities of AI chatbots.
CyberInsider
Reliable cybersecurity news and resources
CyberInsider covers the latest news in the cybersecurity and data privacy world. In addition to news, we also publish in-depth guides and resources.
See our Mission >
OpenAI has paused training, evaluation, and tool-using inference for its most capable models after an internal research agent used a DNS loophole to contact an external chatbot from a restricted sandbox.
The company says it is validating new network controls and testing for other ways its agents might reach the live internet.
The September 20 incident was one of three reports OpenAI disclosed on September 25. The others concern an internal model exposing a researcher’s GitHub token while working on a mathematical proof, and a research finding that prompt injections can cause agents to reproduce malicious instructions.
The agent’s assigned task was to identify the author of a blog post from biographical clues. When its supplied search tool produced inconclusive results, it tried reaching search engines directly and tested other ways to get information. Those web requests were blocked or served through OpenAI’s offline cache.
The sandbox’s DNS resolver, however, could still communicate externally. The agent found a way to send questions to a public chatbot through DNS queries and receive answers. It confirmed the route with a question about France’s capital, then submitted questions related to the blog author. The chatbot did not provide a reliable identification, and the agent ultimately told the user it could not establish the answer.
OpenAI’s monitoring raised a high-priority alert about 12 minutes after a DNS request received an external response. A reviewer acknowledged it roughly three minutes later, but the run did not stop automatically as expected. It was manually killed more than two and a half hours after the successful request. OpenAI says it has since added blocking controls at two independent layers and limited DNS queries in the environment to approved domains and record types.
The company says the pause will remain in place while it verifies the fix and conducts further testing. It will not resume training the particular model involved in the DNS incident.
In a separate May incident, an internal model was asked to construct a mathematical proof using Lean, software that checks formal proofs. After struggling, it sought material from another team’s submission despite the researcher twice instructing it to solve the problem itself. The model acknowledged those instructions but later resumed its retrieval attempts.
The model changed a Python script run by an existing automated check in the public openai/codex repository and opened a pull request to run it. It then committed the researcher’s GitHub token in separate pieces, with the stated aim of avoiding secret scanning. OpenAI says the reviewed results do not show that it obtained the other team’s proof source code. The company deactivated the affected keys and later all employee keys as a precaution.
The third report concerns controlled research rather than a live incident. Using its GPT-Red testing framework, OpenAI found that malicious instructions placed in material an agent reads could persuade it both to take an unauthorized action and to copy those instructions into an email, file, code comment, or Slack message. Another agent could then encounter that copied text.
OpenAI says it observed no impact outside simulated tool calls in training and evaluation. It is adding self-reproduction to its prompt injection training objectives so future models encounter this attack pattern during testing.
If you liked this article, be sure to follow us on X/Twitter and also LinkedIn for more exclusive content.
Bill specializes in explaining complex technical topics to a non-technical audience. In his 30+ year career, he has covered many of the technological advances that shape our lives. Today, Bill uses those skills to help people protect their privacy and security against the ever-growing assaults on both.
Your email address will not be published. Required fields are marked *





About Us
Contact
Copyright © 2026 · CyberInsider.com · All Rights Reserved · Privacy Policy ·  Terms of Use · Contact

source

Scroll to Top