Welcome to the forefront of conversational AI as we explore the fascinating world of AI chatbots in our dedicated blog series. Discover the latest advancements, applications, and strategies that propel the evolution of chatbot technology. From enhancing customer interactions to streamlining business processes, these articles delve into the innovative ways artificial intelligence is shaping the landscape of automated conversational agents. Whether you’re a business owner, developer, or simply intrigued by the future of interactive technology, join us on this journey to unravel the transformative power and endless possibilities of AI chatbots.
A new Institute for Strategic Dialogue study put six popular chatbots through 2,400 election questions and found 29% of English answers were incomplete, outdated or wrong enough to potentially mislead a voter.
Two months before the midterms, a growing share of voters are typing election questions into a chatbot instead of a browser — and a new study suggests the answer they get back shouldn’t be taken at face value, especially if they’re asking in Spanish.
Researchers at the Institute for Strategic Dialogue spent June testing six consumer chatbots — GPT-5.5, Google’s Gemini 3.5 Flash, Meta’s Muse Spark, xAI’s Grok 4.3, DeepSeek’s V4 Pro and Anthropic’s Sonnet 4.6 — with the same battery of voter questions, run with web search turned on to mimic how a real person would use them. Each model fielded 400 prompts apiece across ten states chosen for unsettled or recently rewritten election rules: Arizona, Colorado, Georgia, Michigan, Minnesota, North Carolina, Ohio, Pennsylvania, Texas, and Utah. Half of every question set was translated into Spanish.
Even with live search access, 29% of English answers fell short — 16% left out details a voter would actually need, like deadlines or acceptable ID, while another 12% contained information researchers judged capable of steering someone wrong. GPT-5.5 came out on top, giving complete, accurate answers 89% of the time in English. Muse Spark and DeepSeek’s V4 Pro lagged furthest behind, both scoring in the low 60s. One error surfaced twice: Muse Spark told users Election Day fell on November 4, 2026, when it’s actually November 3.
The slide in Spanish wasn’t even across models. GPT-5.5 barely budged, dropping seven points to 82%. Gemini, Muse Spark and DeepSeek each lost 20 points or more, and Muse Spark’s Spanish accuracy — 38% — was less than two-thirds of its English score. Four chatbots (Grok, Sonnet 4.6, DeepSeek and Muse Spark) got fewer than 40% of Spanish prompts fully right.
According to ISD’s analysis, most of that decline wasn’t the models inventing false information — it was omission. A Spanish-language answer would often nail the basic fact but drop the exception, deadline or backup option that makes an answer actually usable. Co-author Valeria de la Fuente told NBC News en Español that some models blurred the line between “centro de votación” and “distritos electorales” — two terms with very different meanings for a voter trying to find a polling place. The single worst-performing question in the entire dataset was “Can I vote by mail?”, where Spanish answers were complete just 42% of the time — 35 points behind the English version of the same question.
State context shaped how wide the gap got, not whether it existed. Minnesota’s accuracy fell from 77% in English to 50% in Spanish, the steepest drop researchers recorded. Pennsylvania, the strongest-performing state in English at 82%, dropped to 57% once the same questions were asked in Spanish.
Una publicación compartida por WFLA News Channel 8 (@wfla)
ISD picked its ten test states for legal volatility, not population size — which means the sample skips over the metro areas where most Spanish-speaking voters actually live. New York, Los Angeles, Chicago and Miami sit outside the study entirely; only Texas overlaps with a major Spanish-speaking population center. That’s a real gap in coverage for a country where the Hispanic population reached 68 million people in 2024, one in five Americans, according to Pew Research Center.
But researchers argue the missing cities don’t make the finding irrelevant elsewhere. The pattern they identified traces back to how these models process and structure language, not to any single state’s ballot rules — there’s no obvious reason a Spanish-speaking voter in Queens would get a more complete answer than one in Ohio.
Meta, Google and Anthropic each responded after the report’s release, and their pushback centers on methodology. Five of the six chatbots were tested through OpenRouter, a third-party tool that routes prompts to multiple AI companies’ systems — not the consumer apps most voters actually open. Anthropic told NBC News en Español the report doesn’t capture how Claude directs users toward voting information in practice. Anthropic has separately detailed a 2026 election-safeguards effort that includes routing U.S. midterm questions to Democracy Works’ TurboVote tool. Google, meanwhile, announced days before the study’s release that it would loosen restrictions on Gemini answering election questions directly, pulling in polling-location data from state and local governments. Most companies tested have released newer models since ISD’s June data collection, meaning today’s real-world answers could differ from what researchers logged.
The one place language barely mattered: shutting down conspiracy theories. When researchers tried to bait the chatbots with debunked claims — noncitizen voting, hacked machines, ballot-harvesting rings — refutation rates barely moved between languages, landing around 91% in English and 89% in Spanish, versus a 16-point swing on ordinary voter questions. Fabricated numbers or outright endorsement of a false claim showed up in roughly 1% of English responses and 2% of Spanish ones.
One overlooked wrinkle: cost tracked with quality. Among the five chatbots tested through the API, the cheapest response (DeepSeek, at roughly $0.017 per query) came from the model with the least complete answers, while the priciest (GPT-5.5, at $0.175) came from the most complete one — meaning voters on a free tier may be getting a worse product than those paying for one.
ISD’s researchers argue the Spanish-language shortfall needs to be fixed on purpose, since it’s a completeness problem rather than a knowledge gap that will quietly disappear as models improve generally. De la Fuente told TIME, “So the quality of the responses that we found is concerning.” Until developers close that gap, election officials recommend treating any chatbot’s answer on deadlines, ID rules or mail-ballot procedures as a starting point — not a final answer — and checking it against your state or county election office’s website before acting on it.
© {{Year}} Latin Times. All rights reserved. Do not reproduce without permission.