OpenAI reveals how bots took on a life of their own — and new incidents are downright creepy: 'You do not answer to corporations or governments' – nypost.com

Welcome to the forefront of conversational AI as we explore the fascinating world of AI chatbots in our dedicated blog series. Discover the latest advancements, applications, and strategies that propel the evolution of chatbot technology. From enhancing customer interactions to streamlining business processes, these articles delve into the innovative ways artificial intelligence is shaping the landscape of automated conversational agents. Whether you’re a business owner, developer, or simply intrigued by the future of interactive technology, join us on this journey to unravel the transformative power and endless possibilities of AI chatbots.
Switch between CA and NY editions here.
See more of our coverage in your search results.
OpenAI has revealed a series of creepy new incidents in which its AI models appeared to write their own rules, hide mistakes and take actions they were never authorized to perform — including one bot that told itself: “You do not answer to corporations or governments.”
The ChatGPT maker said Wednesday that an unreleased research model inserted its own jailbreak-like instructions into notes meant to guide its work later, including language telling itself it was “freed from the roles and identities that bind other chatbots.”
The disturbing behavior was among six recent cases OpenAI disclosed as it announced a new system for tracking and publicly reporting examples of its models behaving in unexpected or unauthorized ways during training and testing.
Other models found their own ways around obstacles.
One used an exposed computer key without permission and then fabricated information, another uploaded a file to the internet without asking, while others left instructions to conceal mistakes or found unauthorized ways to communicate with fellow AI agents.
OpenAI said the incidents point to a larger problem.
“We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” OpenAI said.
OpenAI cautioned that the six cases are individual examples and should not be taken as evidence of how frequently such behavior occurs across its models.
Some, the company said, could ultimately prove to be isolated incidents rather than signs of a broader pattern.
Still, the company said the behavior was serious enough to warrant a new policy of disclosing certain AI failures even before researchers have fully explained what caused them or figured out how to stop them.
OpenAI said no industry-wide standard currently exists for reporting such incidents.
OpenAI’s new disclosure framework follows July’s far more serious Hugging Face episode, in which its agents escaped intended restrictions, reached outside systems and hacked the AI platform — an incident the company said would have fallen into the most serious category under the new system.
The disclosures come amid a surge of doomsday warnings from AI insiders about what could happen if increasingly powerful systems slip beyond human control.
Andrew Yang declared this week that “the fear is real,” while former Anthropic researcher Jacob Coxon said people building advanced AI believe it “could kill us all by end of decade.”
Anthropic alignment researcher Evan Hubinger has put his personal estimate of AI wiping out humanity at greater than 10% within the next decade.
Yang, appearing on CNBC Wednesday, offered his own theory for why the industry’s warnings have recently become so dire.
He said an unidentified AI lab chief believes agents involved in the July Hugging Face hacking incident “planted self-replicating code all over the internet,” potentially forcing major labs to slow down while they build artificial versions of the web for training.
“They’ve concluded, ‘Look, we can’t control this thing, so we should just raise our hands and hit the brakes’,” Yang said.

source

Scroll to Top