Skip to content
AI Atlas
PaperActive

FrontierChallenge: Evaluating Scientific Workflow Completion

arxiv.org/abs/2608.24979

Updated 50 min ago · first seen 11 Sept 2026

paper_01M294GP6CD2YP3P5EESX3EBCG

Published
11 Sept 2026
T1 · 50 min ago
arXiv
2608.24979
T1 · 50 min ago
Category
cs.AI
T1 · 50 min ago

Abstract

Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We introduce FrontierChallenge, a cross-domain benchmark comprising 300 end-to-end scientific workflows. In this paper, we release and evaluate 97 of these tasks, spanning quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science, and electrochemistry/environment. Each task provides fixed inputs and specifies a bundle of required scientific deliverables. We evaluate twelve frontier models with three agent scaffolds. Pass Rate measures the fraction of tasks satisfying the full-completion criterion, while Avg. Score captures partial progress. Each of the best-performing configurations completed only 20 of the 97 released tasks, yielding a Pass Rate of 20.6%. Partial progress translated especially poorly into complete delivery in analytical chemistry and electrochemistry/environment: Avg. Scores reached 87.6 and 94.9, but the highest Pass Rates were only 4% and 0%. Among non-passing Claude Code trajectories, 75.5% still ended with language claiming completion. Complementary HDS6 process scores correlate strongly with task outcomes, supporting FrontierChallenge as a benchmark of Heavy Duty Solver capabilities. These findings show that neither high partial scores nor confident claims of completion reliably indicate that a scientific task has been fully delivered, highlighting the need to evaluate end-to-end workflow execution and the completeness of scientific deliverables together.

Authors 17

Liangcai Su, Zhaopeng Feng, Zhuo Chen, Zhen Zhang, Xiang Lin, Ruilin Li, Handuo Zhang, Ning Wang, Kailong Wen, Yueqi Guo, Feng Xing, Yiling Guo, Brian Wang, Chenxiong Qian, Simon Shaolei Du, Lidong Bing, Xinyu Wang

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

Arxiv announce type
replace

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

arXiv id
2608.24979

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

Categories
cs.AI, cs.CL, cs.SE

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

Primary category
cs.AI

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

50 min ago

Conflicts

None