Skip to content
AI Atlas
PaperActive

The widening evaluation gap in medical large language model research 2023 to 2026

arxiv.org/abs/2609.11770

quality89

Updated 6 h ago · first seen 11 Sept 2026

paper_01M294G4VKZ01G7CBT9EESRV9C

Published
11 Sept 2026
T1 · 6 h ago
arXiv
2609.11770
T1 · 6 h ago
Category
cs.CL
T1 · 6 h ago

Abstract

Large language models are superseded every few quarters; clinical evidence takes years. We asked whether medical research is keeping pace with the systems it evaluates. PubMed returned 11,628 records for January 2023 to June 2026 across fourteen clinical domains, growing 45-fold; 2.5% used a randomised, controlled or prospective design. Evaluation lag, from a study's newest named model release to its own publication, widened from 1.33 to 6.08 quarters. Because discontinued models age mechanically, we benchmarked this against a counterfactual holding model composition fixed: migration to newer systems offset only 56% of the drift (95% CI 50-65). Randomised trials evaluated models a median 4.6 quarters older than other designs (P = 3 x 10^-19), yet among studies naming a model still under development no design differed from any other; 62% of randomised trials evaluated a discontinued family. Rigour and currency are in tension, and that tension reflects model selection rather than research timelines.

Authors 3

Raad Bin Tareaf, Murad Al-Rajab, Samia Loucif

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

arXiv id
2609.11770

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

Categories
cs.CL

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

Primary category
cs.CL

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

6 h ago

Conflicts

None