The widening evaluation gap in medical large language model research 2023 to 2026
Updated 9 h ago · first seen 11 Sept 2026
paper_01M294G4VKZ01G7CBT9EESRV9C
- Published
- 11 Sept 2026
- T1 · 9 h ago
- arXiv
- 2609.11770
- T1 · 9 h ago
- Category
- cs.CL
- T1 · 9 h ago
Abstract
Large language models are superseded every few quarters; clinical evidence takes years. We asked whether medical research is keeping pace with the systems it evaluates. PubMed returned 11,628 records for January 2023 to June 2026 across fourteen clinical domains, growing 45-fold; 2.5% used a randomised, controlled or prospective design. Evaluation lag, from a study's newest named model release to its own publication, widened from 1.33 to 6.08 quarters. Because discontinued models age mechanically, we benchmarked this against a counterfactual holding model composition fixed: migration to newer systems offset only 56% of the drift (95% CI 50-65). Randomised trials evaluated models a median 4.6 quarters older than other designs (P = 3 x 10^-19), yet among studies naming a model still under development no design differed from any other; 62% of randomised trials evaluated a discontinued family. Rigour and currency are in tension, and that tension reflects model selection rather than research timelines.
Authors 3
Raad Bin Tareaf, Murad Al-Rajab, Samia Loucif
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 9 h agohigh
- Arxiv announce type
- new
Source:arXiv (Atom API + RSS)T1observed 9 h agohigh
- arXiv id
- 2609.11770
Source:arXiv (Atom API + RSS)T1observed 9 h agohigh
- Categories
- cs.CL
Source:arXiv (Atom API + RSS)T1observed 9 h agohigh
Source:arXiv (Atom API + RSS)T1observed 9 h agohigh
- Primary category
- cs.CL
Source:arXiv (Atom API + RSS)T1observed 9 h agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 9 h agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
9 h ago
Conflicts
None
No models linked to this paper yet.
- Authors
- Raad Bin Tareaf, Murad Al-Rajab, Samia Loucif
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · Published
Publishedpublished_at1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| 11 Sept 2026 | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
New paper: The widening evaluation gap in medical large language model research 2023 to 2026
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.CL | feed | T1· Official | 1 h ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.