Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs
Published 15 Sept 2026arXiv:2609.13582
Updated 29 h ago · first seen 15 Sept 2026
paper_01M2JK0CCXFREJWD9T2KXNC265
Abstract
A clinical agent benchmark can report the same verdict on identical inputs while the agent files a materially different order on each run. Such agents order tests, request medications and place referrals, yet benchmarks typically score one run per task and rarely ask whether identical inputs produce identical actions; MedAgentBench, the benchmark we use, scores a single attempt and says so. To measure this gap we introduce "same-input rerun", which replays a task with every input held fixed and compares the orders rather than the score, with six reliability metrics, and apply it to 1000 MedAgentBench runs across 50 tasks from its five write-capable families, two open-weight models below ten billion parameters quantised to four bits, and two temperatures. The study establishes that action-level divergence exists and can pass unrecorded by the score, not that any rate generalises. Under the 8B model at temperature 0.7, all 43 ordering groups emit a different set of orders across five identical runs, 26 emit the order on some runs and not others, and 28 record a different coded value, dose or analyte. In 22 of those 43 the benchmark reports the same failing verdict for materially different behaviour, as it does for all 10 divergent groups of the 4B model at 0.7. Orders also reach different endpoints across runs, one of which the record server rejects while the agent is told it succeeded. These findings motivate repeated-run evaluation, action-level stability reporting and execution-faithful environment feedback in clinical-agent benchmarks.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 3
- Property changedPaperSame Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs
Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs: arxiv announce type changed from new to cross
Arxiv announce typenew→crossarxiv - Property changedPaperSame Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs
Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs: arxiv announce type changed from cross to new
Arxiv announce typecross→newarxiv - New paperPaperSame Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs
New paper: Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs
arxiv
Sources
Sources 3
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.