AI Chatbots Outperform World Debate Champions in Persuasion Tests – 동아사이언스

Welcome to the forefront of conversational AI as we explore the fascinating world of AI chatbots in our dedicated blog series. Discover the latest advancements, applications, and strategies that propel the evolution of chatbot technology. From enhancing customer interactions to streamlining business processes, these articles delve into the innovative ways artificial intelligence is shaping the landscape of automated conversational agents. Whether you’re a business owner, developer, or simply intrigued by the future of interactive technology, join us on this journey to unravel the transformative power and endless possibilities of AI chatbots.
This content was translated by an AI-powered translation tool.
It may contain inaccuracies due to the limitations of machine translation.
Copyright © DongA Science. All rights reserved.
Oxford and Yale researchers show LLM-based chatbots shift political opinions more than top human debaters in a 6,923-participant study
A large-scale experiment has found that AI chatbots are more effective at persuasion than highly compelling speakers who have won world debate championships.

On the 20th (local time), the international journal 'Science' introduced the latest experiment comparing the persuasive power of AI and humans. The findings by Kobi Heikkenberg and colleagues at the University of Oxford in the UK were released on June 15 (local time) on the preprint server 'arXiv'.

AI chatbots based on large language models (LLMs) such as ChatGPT, Gemini and Claude are becoming increasingly adept at persuading people on political and social issues. Ethical concerns have already emerged, as in a study by the University of Zurich in Switzerland, where fake accounts secretly persuaded users in online communities and influenced their views.

Heikkenberg's team carried out a large-scale experiment to compare the persuasive power of humans and AI. They first prepared 10 contentious political and social claims related to the UK, such as 'The UK should return cultural artifacts taken during the colonial era', 'The UK should accept more immigrants', and 'Social media use by people under 16 should be banned'.

A total of 6,923 participants indicated how much they agreed with each claim on a scale from 0 to 100, then conversed with either a person or an AI chatbot, and afterwards recorded their responses again. Persuasive effect was evaluated based on how much their attitude scores toward each claim shifted.

Debates were conducted as online, real-time text chats, and participants did not know whether their counterpart was a human or an AI. Each topic-based debate lasted about 14 minutes.

When 132 participants, recruited to represent the general public living in the UK, attempted to persuade others, they achieved on average a 4.7-point greater persuasive effect than a control group that engaged in generic conversation using the ChatGPT-4o model. A persuasion-specialized AI chatbot produced an effect 8.2 points higher than these human persuaders.

The results were similar when the researchers gathered people who were already good at persuading. They recruited 1,154 members of the public, ran a four-round persuasion tournament over three weeks, and selected 87 people in the top 10%, offering prize money. These skilled persuaders shifted attitudes by 7.2 points more than the control group, but the persuasion-focused AI changed attitudes an additional 5.6 points beyond that.

AI also proved more persuasive than 56 expert human debaters, including four world debate champions and 11 continental champions. Their average debating experience was 8.9 years. 

To favor humans, the researchers disclosed the topics 21 days before the debates and allowed debaters to choose the topic they felt they could argue most persuasively. The experiment showed that expert debaters achieved an average persuasion score of 8.3 points, the highest among the human groups, yet even here the AI scored 4.6 points higher.

The researchers then brought back 43 debate experts who had taken part in the third experiment and had them review their earlier conversations and practice directly with the AI, providing a total of eight hours of additional training. The debaters showed meaningful changes, writing messages about 19% longer and presenting roughly 54% more verifiable facts in each exchange than before. The trained experts achieved a persuasion score of 9.7 points, but still could not catch up with the AI.

Only when the length and speed of AI responses were capped at human levels did their persuasive power become comparable. This suggests that AI's core logical ability may not fundamentally surpass that of humans; instead, the decisive difference could be that AI can supply far more facts and evidence much faster than people.

Scientists also point out that AI is a 'double-edged sword' that can both prevent people from falling into conspiracy theories and, just as easily, lure them into such beliefs. The same technology can have ambivalent effects on public health and democracy.

At present, chatbots developed by a small number of AI companies dominate the field. There are concerns that decisions by a few large corporations could have a major impact on how people think and behave.

Luciano Floridi of Yale University in the United States said that if we cannot avoid the influence of AI chatbots in the future, then "we must make them pluralistic and diverse", adding, "The world will become noisier, but it may be less manipulable."

<Reference>
– doi.org/10.48550/arXiv.2606.16475

source

Scroll to Top