VideoScout: Learning Agentic Active Exploration with Adaptive Reasoning Pacing for Long Video Understanding
Published 16 Sept 2026arXiv:2609.15606
Updated 7 h ago · first seen 15 Sept 2026
paper_01M2JK19DV0NPVW2HJ4TH20W8F
Abstract
Multimodal Large Language Models (MLLMs) have achieved remarkable progress on short video understanding yet remain limited on long videos due to the limited visual context window. Prevailing approaches rely on uniform frame sampling or recent coarse-to-fine agentic zooming, both of which struggle to localize sparse, decisive evidence in sufficiently long videos. We formulate long video understanding as a \textbf{Sequential Evidence Acquisition (SEA)} problem, in which an agent reads the video turn by turn along the temporal axis, deciding at each turn how fast to watch, what evidence to retain, when to revisit uncertain segments, and when to stop and answer. Inspired by this view, we propose \textbf{VideoScout}, a multi-turn reasoning agent that instantiates the SEA paradigm through adaptive reasoning pacing. Specifically, by dynamically controlling the viewing pace, VideoScout enables efficient traversal of long videos within a bounded visual context window, allowing the agent to access more video content while balancing content analysis depth with reading efficiency. To train VideoScout, we construct VideoScout-66K, a set of over 66K high-quality exploration turns from 10K answer-verified trajectories, and adopt a two-stage pipeline: cold-start supervised fine-tuning teaches the agent per-turn output format, while the Decoupled Clip and Dynamic sAmpling Policy Optimization (DAPO) algorithm performs trajectory-level reinforcement learning with a composite reward that jointly considers answer accuracy, output format compliance, and the temporal alignment between the agent's viewing progress and the teacher's answer timing measured by intersection-over-union (IoU). Extensive experiments on long video understanding and reasoning benchmarks demonstrate that our 7B model achieves strong performance compared with existing trained 7B agentic models.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 4
- Property changedPaperVideoScout: Learning Agentic Active Exploration with Adaptive Reasoning Pacing for Long Video Understanding
VideoScout: Learning Agentic Active Exploration with Adaptive Reasoning Pacing for Long Video Understanding: arxiv announce type changed from new to cross
Arxiv announce typenew→crossarxiv - Property changedPaperVideoScout: Learning Agentic Active Exploration with Adaptive Reasoning Pacing for Long Video Understanding
VideoScout: Learning Agentic Active Exploration with Adaptive Reasoning Pacing for Long Video Understanding: published at changed from 2026-09-15T04:00:00+00:00 to 2026-09-16T04:00:00+00:00
Published15 Sept 2026→16 Sept 2026arxiv - Property changedPaperVideoScout: Learning Agentic Active Exploration with Adaptive Reasoning Pacing for Long Video Understanding
VideoScout: Learning Agentic Active Exploration with Adaptive Reasoning Pacing for Long Video Understanding: arxiv announce type changed from cross to new
Arxiv announce typecross→newarxiv - New paperPaperVideoScout: Learning Agentic Active Exploration with Adaptive Reasoning Pacing for Long Video Understanding
New paper: VideoScout: Learning Agentic Active Exploration with Adaptive Reasoning Pacing for Long Video Understanding
arxiv
Sources
Sources 2
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.