Why you can’t trust AI chatbots for science – Lab Muffin Beauty Science

Welcome to the forefront of conversational AI as we explore the fascinating world of AI chatbots in our dedicated blog series. Discover the latest advancements, applications, and strategies that propel the evolution of chatbot technology. From enhancing customer interactions to streamlining business processes, these articles delve into the innovative ways artificial intelligence is shaping the landscape of automated conversational agents. Whether you’re a business owner, developer, or simply intrigued by the future of interactive technology, join us on this journey to unravel the transformative power and endless possibilities of AI chatbots.
Lab Muffin Beauty Science
Lazy AI-generated content is everywhere now. Unfortunately, AI chatbots like ChatGPT, Claude and Gemini aren’t reliable for factual content, especially science content.
AI chatbots don’t simply report information from a database, and they haven’t “learned all of human knowledge”. This is a common misconception that’s perpetuated by AI companies – when people think AI chatbots are kind of magic, they tend to use and trust them more.
AI chatbots like ChatGPT, Claude and Gemini are based on large language models (LLMs). Fundamentally, these put words that are usually near each other, near each other, without really understanding what the words actually mean.
An analogy: they basically have a huge set of loaded (uneven) dice for outputting text. The dice are chosen based on the prompt, and the previous words (not just the previous word – this is partly why modern chatbots are much more powerful than older ones).
How the “dice” are made (this is very simplified):
The overall effect is that the LLM essentially mashes together text that’s selected as relevant. Other sources like web search results can be incorporated, but the LLM architecture is still used to interpret prompts and output text responses.
how LLMs work - dice analogy
This leads to particular accuracy issues…
LLM-based outputs appear sensible and human-like because language correlates with concepts, but they aren’t the same (this is essentially correlation isn’t the same as causation).
Essentially, AI chatbots are always BS-ing – they output text without reference to reality.
Chatbot outputs often correlate with reality, so they can be useful. Mashing examples of text together works surprisingly well for some tasks! But it also doesn’t work well for many tasks that are simple for humans to do.
Note: This is why you can’t “ask AI to do something” without verifying it. It’s more likely to mash together examples of text where that task was done, than to actually do it.
The accuracy of AI chatbots depends on the accuracy and quantity of inputs (training data text, search results).
There are many examples of AI summaries of studies that are woefully incorrect – for example, a study I coauthored was described as research from “a Harvard team”, even though only 1 out of 4 authors are from Harvard.
It’s currently unclear whether techniques that (in theory) should increase accuracy, actually do. This includes:
These can give more convincing illusions of thinking – seemingly logical reasoning for an answer that wasn’t determined by logic, but from standard LLM-based pattern-matching.
Bottom line: You need to know the “real answer” (and whether there even is a real answer) to know if AI-generated text is accurate or not.
If you know the topic, then in theory, you could just check the AI chatbot’s output. But that’s where our cognitive biases kick in…
Fluency heuristic: Fluent, easy-to-process text feels more truthful to us
Confidence effect: We tend to see sources that sound more confident as more credible 
Confirmation bias: AI chatbots tend to give us the answers we want
Automation bias: We tend to trust decisions made by machines
Authority bias: People who don’t understand how AI chatbots work believe they’re authoritative
These mental shortcuts require a lot of effort to override – you need to fight against thousands of years of human evolution. It’s extremely easy to just nod along and agree.
On top of being inaccurate and extra hard to factcheck, there are other reasons why AI-generated text is one of the worst starting points for factual content, largely as a result of the way LLMs work:
Anchoring effect: It’s hard for us to override how the starting point frames things
Verbose: LLMs tend to add a lot of unnecessary fluff and repeat the same information multiple times, so it takes longer to read and check
Bad at organising concepts: If you’ve read an AI summary of a meeting or an email, you’ll know it groups things weirdly
Jagged intelligence: Mashing text together only works for some tasks, so LLMs make weird inhuman mistakes we’re not used to looking for
Compounded biases: LLMs output common patterns from its sources, including racial, gender etc. biases – this is getting worse as LLMs are trained on more LLM outputs (model collapse)
Anecdotally, I’ve struggled to find creators posting science content with obvious AI slop signs who aren’t also making factual mistakes. Writing in your own style is usually far easier than getting the facts right, and factchecking AI text is very difficult with all these cognitive biases at play.
That’s why spotting AI slop is a red flag for inaccurate science content, and it’s more complex than just seeing an em dash! Stay tuned…
This article was adapted from my video on AI slop from science communicators. An infographic version can be found on Instagram or YouTube.
Tully SM, Longoni C, Appel G. Lower Artificial Intelligence Literacy Predicts Greater AI Receptivity. J  Marketing. 2025;89(5):1-20.
George D Montanez. LLMs, Model Collapse, and the Conservation of Information (lecture, on YouTube as “Model Collapse Ends AI Hype”).
On Bullshit, Hallucinations, and AI. Jordan Harrod, es machina (Substack). February 26, 2025.
No, “AI” is not a Stochastic Parrot 🦜. Margaret Mitchell. Medium. March 5, 2026.
The architecture behind web search in AI chatbots. Ida Silfverskiöld, Towards Data Science. December 4, 2025.
Baethge C, Jergas H. Systematic review and meta-analysis of quotation inaccuracy in medicine. Res Integr Peer Rev. 2025;10(1):13.
Smith N, Cumberledge A. Quotation errors in general science journals. Proc. R. Soc. A 2020;476(2242):20200538.
Chen Y et al. (Anthropic Alignment Science Team). Reasoning models don’t always say what they think. May 8, 2025. arXiv:2505.05410.
Sharma A, Chopra P. EsoLang-Bench: Evaluating Genuine Reasoning in Large Language Models via Esoteric Programming Languages. May 11, 2026. arXiv:2603.09678
Vishwanath K, Alyakin A, Ghosh M, et al. General-purpose large language models outperform specialized clinical AI tools on medical benchmarks. Nat Med. 2026;32(7):2405-2409. doi:10.1038/s41591-026-04431-5
Free Sunscreen Science E-Summit!
All About Bee Venom and Honey in Skincare
Fading Hair Dye With Low Damage
Powered By Chemicals Coffee and Tea Molecule Mugs




about michelle
Hi! I’m Michelle, chemistry PhD and science communicator, and I’m here to help you figure out which beauty products are and aren’t worth buying, using science.

source

Scroll to Top