OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows
Updated 7 h ago · first seen 11 Sept 2026
paper_01M294GK0ZVPPQSGQVAKVX7QC5
- Published
- 11 Sept 2026
- T1 · 7 h ago
- arXiv
- 2609.09203
- T1 · 7 h ago
- Category
- cs.AI
- T1 · 7 h ago
Abstract
Existing benchmarks for autonomous AI scientists evaluate only final outputs---generated code, hypotheses, or papers---yet discard the reasoning process by which those outputs were obtained. This makes it impossible to audit scientific methodology, diagnose failure modes, or distinguish systematic reasoning from fortunate guessing. We present \textbf{OpenDiscoveryTrace}, a public dataset of 558 complete AI scientific agent trajectories that captures how models reason, not just what they produce. Each trajectory records a structured 9-field-per-step trace---including thoughts, tool calls, observations, errors, revision triggers, and self-reported confidence---as models execute 124 scientific tasks spanning drug discovery, materials science, genomics, and scientific literature analysis. The dataset covers seven models: three frontier models (GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro; 124 trajectories each, fully balanced across domains and difficulty levels) and four open-weight models (Qwen2.5-7B, Mistral-7B-v0.3, Phi-3.5-mini, and Qwen2.5-1.5B; 30 each), plus 60 live-retrieval variant trajectories. Pilot analysis on 363 LLM-judged trajectories reveals that process traces expose behavioral differences invisible to output-only evaluation: all three frontier models achieve comparable success rates (84--89%), yet Claude Opus 4.6 produces 30$\times$ more errors than GPT-5.4 (2.5 vs. 0.08 per trajectory, $p < 0.0001$, Cliff's $\delta = 0.613$), with qualitatively different error profiles---66.7% tool misuse for Claude versus 83.6% reasoning errors for GPT-5.4. We define five benchmark tasks with baselines from logistic regression, random forests, LSTMs, and Transformer models. The dataset, trace schema, agent harness, and benchmark definitions are publicly available under CC BY 4.0 to support research on process-level evaluation, scientific agent auditing, and AI governance.
Authors 2
Aayam Bansal, Keertan Balaji
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
- Arxiv announce type
- new
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
- arXiv id
- 2609.09203
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
- Categories
- cs.AI
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
- Primary category
- cs.AI
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
7 h ago
Conflicts
None
No models linked to this paper yet.
- Authors
- Aayam Bansal, Keertan Balaji
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · Published
Publishedpublished_at1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| 11 Sept 2026 | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
New paper: OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.AI | feed | T1· Official | 5 h ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.