Why benchmarking clinical LLMs from OpenEvidence, Doximity is complicated – STAT

Welcome to the forefront of conversational AI as we explore the fascinating world of AI chatbots in our dedicated blog series. Discover the latest advancements, applications, and strategies that propel the evolution of chatbot technology. From enhancing customer interactions to streamlining business processes, these articles delve into the innovative ways artificial intelligence is shaping the landscape of automated conversational agents. Whether you’re a business owner, developer, or simply intrigued by the future of interactive technology, join us on this journey to unravel the transformative power and endless possibilities of AI chatbots.
Home
What's the word?
Test your knowledge with our new weekday mini crossword
By Brittany Trang
July 29, 2026
Health Tech Reporter
Brittany Trang
Brittany Trang, Ph.D., covers AI in health and medicine: Does it actually work? Who benefits, or might be harmed? She writes the weekly AI Prognosis newsletter. Follow her on Threads, Mastodon, and Bluesky. You can reach Brittany on Signal at btrang.01.
You’re reading the web edition of STAT’s AI Prognosis newsletter, our subscriber-exclusive guide to artificial intelligence in health care and medicine. Sign up to get it delivered in your inbox every Wednesday. 
I saw “The Odyssey” during its opening weekend. Ever since then, I have been questioning whether I’m illiterate or whether Christopher Nolan is a poor storyteller. This London Review of Books evaluation of the film, written by the woman whose translation of “The Odyssey” Nolan apparently read, has freed me from my wondering. (h/t to my colleague Matthew Herper)
Advertisement
Hot takes on Homer’s epic, or hot tips about Epic Systems: [email protected]
You might recall that in mid-June, there was a Nature Medicine study that pitted clinical AI systems OpenEvidence and UpToDate Expert AI against general LLMs. It set off a reaction in the clinical AI world like no other paper has. “The results rang out like a gunshot,” as STAT health tech correspondent Katie Palmer describes it.
The controversy surrounding the study, and everything that came after, exemplifies the problems I have with benchmarks.
Advertisement
Katie summed it up well when I talked to her yesterday: “The way that benchmarks have been talked about generally, and specifically in clinical AI, tends to summarize them into the headlines,” she said. “Every study needs a headline and every story needs a headline, but as we both know, and as I think most people in the industry know, an individual benchmark doesn’t mean much.”
STAT+ Exclusive Story
Already have an account? Log in
Already have an account? Log in
Monthly
$39
Totals $468 per year
Totals $468 per year
Starter
$30
for 3 months, then $399/year
Then $399/year
Annual
$399
Save 15%
Save 15%
11+ Users
Custom
Savings start at 25%!
Savings start at 25%!
2-10 Users
$300
Annually per user
$300 Annually per user
To read the rest of this story subscribe to STAT+.
Brittany Trang
Health Tech Reporter
Brittany Trang, Ph.D., covers AI in health and medicine: Does it actually work? Who benefits, or might be harmed? She writes the weekly AI Prognosis newsletter. Follow her on Threads, Mastodon, and Bluesky. You can reach Brittany on Signal at btrang.01.
Tech is transforming health care and life sciences. Our original reporting is here to keep you ahead of the curve.
Your data will be processed in accordance with our Privacy Policy and Terms of Service. You may opt out of receiving STAT communications at any time.
By Brittany Trang
By Brittany Trang
Advertisement
By Brittany Trang
By Brittany Trang
By Brittany Trang

By Jason Mast

By Vishal Khetpal

By Susan Mayne

By Jason Karlawish

By Mario Aguilar
Share options
X
Bluesky
LinkedIn
Facebook
Doximity
Copy link
Reprints
A decade of reporting from the frontiers of health and medicine
Company
Account
More

source

Scroll to Top