Skip to content
AI Atlas
PaperActive

OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis

arxiv.org/abs/2603.20278

quality89

Updated 3 h ago · first seen 11 Sept 2026

paper_01M294G6FE3S1GHZ09NM6E1EKK

Published
11 Sept 2026
T1 · 3 h ago
arXiv
2603.20278
T1 · 3 h ago
Category
cs.IR
T1 · 3 h ago

Abstract

-cross Abstract: Training deep research agents requires long-horizon trajectories that interleave search, evidence aggregation, and multi-step reasoning. However, existing data collection pipelines typically rely on proprietary web APIs, making large-scale trajectory synthesis costly, unstable, and difficult to reproduce. We present OpenResearcher, a reproducible pipeline that decouples one-time corpus bootstrapping from multi-turn trajectory synthesis and executes the search-and-browse loop entirely offline using three explicit browser primitives: search, open, and find, over a 15M-document corpus. Using GPT-OSS-120B as the teacher model, we synthesize over 97K trajectories, including a substantial long-horizon tail with 100+ tool calls. Supervised fine-tuning a 30B-A3B backbone on these trajectories achieves 54.8\% accuracy on BrowseComp-Plus, a +34.0 point improvement over the base model, while remaining competitive on BrowseComp, GAIA, and xbench-DeepSearch. Because the environment is offline and fully instrumented, it also enables controlled analysis, where our study reveals practical insights into deep research pipeline design, including data filtering strategies, agent configuration choices, and how retrieval success relates to final answer accuracy. We release the pipeline, synthesized trajectories, model checkpoints, and the offline search environment at https://github.com/TIGER-AI-Lab/OpenResearcher.

Authors 10

Zhuofeng Li, Dongfu Jiang, Xueguang Ma, Haoxiang Zhang, Ping Nie, Yuyu Zhang, Kai Zou, Jianwen Xie, Yu Zhang, Wenhu Chen

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

Arxiv announce type
replace

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

arXiv id
2603.20278

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

Categories
cs.IR, cs.AI, cs.CL

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

Primary category
cs.IR

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

3 h ago

Conflicts

None