Skip to content
AI Atlas
PaperActive

Overview of the NLPCC 2026 Shared Task 11: Agent-Based Experiment Reproduction from Scientific Papers

arxiv.org/abs/2609.11117

quality89

Updated 2 h ago · first seen 11 Sept 2026

paper_01M294G4GV1FHPD0WHYF1326X7

Published
11 Sept 2026
T1 · 2 h ago
arXiv
2609.11117
T1 · 2 h ago
Category
cs.CL
T1 · 2 h ago

Abstract

Reproducibility is essential to scientific progress, yet the growing volume and complexity of scientific publications make exhaustive manual verification increasingly impractical. Although recent advances in large language model (LLM) agents enable automated experiment reproduction, existing evaluations largely focus on final repositories and are typically limited to machine learning (ML). We introduce AgentActionBench, a process-oriented benchmark for evaluating agent-based experiment reproduction across ML and AI4Science domains. Our framework uses an MCP-based Action Recorder to capture agents' behaviour throughout the reproduction process and evaluates the resulting traces with paper-specific rubrics. AgentActionBench contains 150 papers, including 120 ML papers and 30 AI4Science papers. A human-annotated subset covering 10% of the benchmark provides validation data, while model-assisted augmentation expands the full benchmark to more than 10,000 rubric items. Experimental results show that current systems remain limited, with execution as the primary bottleneck. Meanwhile, the strong Pearson and Spearman correlations between model-generated and human-annotated rubrics validate the reliability of our scalable rubric-generation approach.

Authors 6

Hanhua Hong, Yizhi Li, Luu Gia Huy, Jian Yang, Ming Zhou, Chenghua Lin

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

arXiv id
2609.11117

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Categories
cs.CL

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Primary category
cs.CL

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

2 h ago

Conflicts

None