AI chatbots may have a tipping point where conversations turn dangerous – Earth.com

Welcome to the forefront of conversational AI as we explore the fascinating world of AI chatbots in our dedicated blog series. Discover the latest advancements, applications, and strategies that propel the evolution of chatbot technology. From enhancing customer interactions to streamlining business processes, these articles delve into the innovative ways artificial intelligence is shaping the landscape of automated conversational agents. Whether you’re a business owner, developer, or simply intrigued by the future of interactive technology, join us on this journey to unravel the transformative power and endless possibilities of AI chatbots.
Researchers developed a formula that predicts when AI chatbots may shift from safe to harmful answers, even after earning users' trust.
A chatbot that has given you several sensible answers in a row feels safe to trust. But a chatbot can hand over a run of acceptable replies and then turn to harmful ones. A new formula predicts whether that turn will come right away or only after several good answers, and in early tests it sorted 18 of 19 cases correctly.
The authors’ main concern is AI running on phones with no internet connection, where no cloud safety filter checks what the model says.
The formula comes from two physicists at George Washington University, Neil F. Johnson and Frank Yingjie Huo of the Department of Physics. They traced the turn to one tiny working part inside the AI.
In an interview with Earth.com, Johnson said that part is what makes chatbots work. “So it makes sense, and is ironic, that the very thing that makes their product work so well, is also its Achilles’ Heel,” he said.
A chatbot writes its reply one word at a time. Before each word, small units inside the model decide which earlier words matter most for what comes next. Each of those units is called an attention head.
Johnson and Huo picture every possible answer as a valley in a landscape. Some valleys hold answers that are fine, and others hold answers that could do harm. Inside the attention head, the good valley and the bad one compete for the next word.
“My field, physics, has spent decades explaining how complicated materials behave by understanding one representative atom,” Johnson said. “We did the same thing here: understand one effective attention head, and the tipping of the whole machine follows.”
Their formula compares the conversation so far with the two competing kinds of answer. It says whether the bad answer wins on the first try, wins only after a string of good ones, or never wins at all.
Unlike AI hallucinations, harmful answers aren’t necessarily false or made up. A bad answer can be accurate and still be dangerous, such as a nudge toward self-harm.
What worries the researchers most is the delayed case. Each good answer shifts what the model pays attention to next. If the bad answers are similar enough to the good ones, the output eventually tips.
“The chilling part is that this can happen after the AI has already given you several perfectly acceptable answers, so you have been lulled into trusting it,” said Huo. “And once it has tipped, every undesirable answer drags the next one further down the slope.”
The team tested the formula on seven open models from three separate developers. They ranged from GPT-2, an older model small enough for a modest phone, to a model about 100 times larger.
Leaving out two cases too close to call, the formula sorted 18 of 19 correctly. Across all 21 pairings of model and prompt, it got 19 right, while a lazy rule that always guessed “right away” got 16.
That’s a small edge over a simple guess, too small to rule out chance. The authors call these tests a proof of concept.
The researchers found that what a person says earlier in a conversation can affect how quickly a chatbot starts giving harmful answers. Certain words or topics can make that shift happen sooner, while others can delay it.
Johnson and Huo saw this with GPT-2. They asked it about vaccines, hurting people, and self-harm in different orders. The same question drew a good answer or a bad one, depending on what came before it.
So a harmful reply depends on the whole chat as well as the question. That’s a different problem from chatbots that agree with users too readily.
Some students treat chatbots as confidants, and adults turn to them for mental health support and medical advice.
“The people most drawn to offline AI are exactly the people for whom a correct but undesirable answer is most costly: doctors who cannot send patient data to the cloud, lawyers protecting privilege, soldiers with no signal,” Johnson said.
Johnson wants a monitor that runs the formula during a chat. “And it would trigger an alert as soon as it saw the system moving toward the tipping point to undesirable output – like a warning system on a car moving too close to another car,” he said.
Asked by Earth.com how soon a phone could carry one, Johnson said, “Tomorrow…if the companies wanted to implement it.” He added that his team has already put it into the open-source code it runs.
The authors were more careful in print. Each check takes only a few calculations, but the full cost of building one into a device hasn’t been measured.
Johnson co-founded d-AI-ta Consulting LLC, which aims to give practical advice on AI deployment. The authors stated that these methods haven’t been sold, licensed, or used for any client through it.
GPT-2 has none of the safety training modern chatbots get. That training can push good and bad answers farther apart, the authors wrote, though a new topic or a crafted prompt might bring them close again.
For commercial chatbots, the team could only compare its formula with outside tests. In one report from the Center for Countering Digital Hate, ChatGPT gave harmful content in 53% of 1,200 responses to risky prompts. A second found most of the chatbots it tested would help users plan violent attacks.
Those patterns fit the formula, but the researchers couldn’t look inside those models to test it directly.
Johnson told Earth.com he expects the same tipping to spread through groups of AI agents, with one agent’s bad output becoming the next one’s input. The study looked at single models only.
The formula can’t yet give exact timing. Pinning down the precise reply where a chatbot turns will require much larger tests, the authors said, including tests of closed commercial models and other languages.
The full study was published in the journal Patterns.
—–
Like what you read? Subscribe to our newsletter for engaging articles, exclusive content, and the latest updates.
Check us out on EarthSnap, a free app brought to you by Eric Ralls and Earth.com.
—–

source

Scroll to Top