Skip to content
AI Atlas
PaperActive

Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents

arxiv.org/abs/2609.11243

quality89

Updated 58 min ago · first seen 12 Sept 2026

paper_01M29X34J849J6ZBQPGG8SCEDM

Published
12 Sept 2026
T1 · 58 min ago
arXiv
2609.11243
T1 · 58 min ago
Category
cs.AI
T1 · 58 min ago

Abstract

Autonomous research agents are increasingly expected to search the literature, analyze experimental evidence, and generate scientific hypotheses. These capabilities require multi-step evidence grounded reasoning that progressively acquires, integrates, and verifies evidence before reaching a conclusion. Existing multimodal benchmarks, however, largely evaluate final-answer accuracy, leaving open whether predictions are actually supported by traceable scientific evidence. We introduce Sci-MMR, a benchmark for multi-step evidence-grounded scientific reasoning built on structured argument graphs linking scientific claims, citation-grounded knowledge, visual evidence, and supporting regions. Sci-MMR comprises 235 multi-hop reasoning tasks spanning four scientific disciplines, with an average of nine figure panels per task. Evaluating eight frontier multimodal models, we find that answer accuracy consistently exceeds complete-evidence recovery rate by more than 20%, revealing a substantial gap that answer-only evaluation is structurally unable to capture. Through controlled interventions, we identify two fundamental bottlenecks. First, evidence acquisition: models struggle to extract complete structured evidence from scientific figures, accounting for 57.2% of failures. While cropping tools yield modest gains (+4.5 points), providing gold evidence improves accuracy by up to 37.0 points, indicating difficulty in assembling complete multi-region evidence. Second, evidence integration: models struggle to translate available evidence into correct conclusions, accounting for 31.8% of failures, while even with gold evidence the strongest model achieves only 69.1% accuracy on the hardest tasks. These findings indicate that current answer-centric benchmarks substantially overestimate the evidence-grounded reasoning capabilities of multimodal research agents

Authors 18

Bicheng Deng, Dingwei Zhu, Enyu Zhou, Han Wang, Jiadong Chen, Jiaqiang Li, Jiazheng Zhang, Lei Bai, Qi Zhang, Senjie Jin, Tao Gui, Xiang Zheng, Xingjun Ma, Yajie Yang, Yang Nan, Yanxin Li, Yuhui Wang, Zhiheng Xi

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 58 min agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 58 min agohigh

arXiv id
2609.11243

Source:arXiv (Atom API + RSS)T1observed 58 min agohigh

Categories
cs.AI

Source:arXiv (Atom API + RSS)T1observed 58 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 58 min agohigh

Primary category
cs.AI

Source:arXiv (Atom API + RSS)T1observed 58 min agohigh

Published
12 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 58 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

58 min ago

Conflicts

None