Skip to content
AI Atlas
PaperActive

Can LLMs Follow Medical Expert Logic? A Benchmark for Hierarchical Logical Consistency in Risk-of-Bias Assessment

arxiv.org/abs/2609.11185

quality89

Updated 1 h ago · first seen 12 Sept 2026

paper_01M29X34HGKD6KTZRTVCA5N2AA

Published
12 Sept 2026
T1 · 1 h ago
arXiv
2609.11185
T1 · 1 h ago
Category
cs.AI
T1 · 1 h ago

Abstract

Evidence-based medicine demands strict logical consistency, yet current evaluations of large language models (LLMs) prioritize superficial label matching over genuine reasoning. We introduce LogiMed-RoB, a benchmark grounded in Cochrane Risk of Bias (RoB) 2.0 expert logic, comprising 860 randomized controlled trials (RCTs) and 14,820 queries. It evaluates models under the Hierarchical Logical Consistency (HLC) framework across four dimensions: Atomic Consistency, Domain Consistency, Aggregation Consistency, and Evidential Faithfulness. Experiments on 10 state-of-the-art LLMs reveal a catastrophic Error Compounding Effect: despite the top model reaching 98.88% Atomic Consistency, its end-to-end consistency collapses to 45.13%, with several open-weight architectures plummeting to nearly 0%. We further uncover a systematic evidence-reasoning gap: even when models retrieve high-quality evidence, they fail to deduce correct outcomes in 18.63-40.05% of cases, while Blind Guess Rates reach 48.28%. LogiMed-RoB demonstrates that high outcome accuracy can conceal critical reasoning flaws, underscoring the necessity of white-box logical verification for clinical deployment.

Authors 5

Haihong E, Jiayu Huang, Qianhui Ling, Zemin Kuang, Zichen Tang

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

arXiv id
2609.11185

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

Categories
cs.AI

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

Primary category
cs.AI

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

Published
12 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 1 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

1 h ago

Conflicts

None