When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models
Published 18 Sept 2026arXiv:2609.19671
Updated 4 h ago · first seen 17 Sept 2026
paper_01M2SE0KQW0KC68MG2J7QNBBM2
Abstract
Large Reasoning Models (LRMs) achieve strong performance on complex tasks but exhibit systematic inefficiency: they often overthink easy problems and underthink hard ones. Existing approaches based on uniform length penalties or rigid routing incur an efficiency tax, trading reduced computation on easy instances for accuracy loss on hard instances. We formulate efficient reasoning as an instance-adaptive computation allocation problem and propose When2Think, a post-training framework for hybrid reasoning that dynamically allocates computation based on problem difficulty. Our method introduces Instance-level Difficulty-Aware Control (IDAC), a reward-shaping mechanism that leverages pre-computed reference statistics (accuracy and token usage) to regulate reasoning depth. Combined with verifier-based rewards and batch-wise standardized advantages, IDAC enables stable critic-free optimization without learned reward models or online reference-model queries. When2Think encourages direct answering on easy instances while preserving extended reasoning on hard instances, thereby learning when to use System 1 (NoThink) versus System 2 (Think). Experiments on mathematical benchmarks demonstrate improved accuracy-efficiency trade-offs: on AIME24, Pass@3 increases by 10.0% while token usage is reduced by 27.9% relative to the base model, and on AIME25, When2Think achieves 40.0% Pass@3, outperforming compression and routing-only baselines.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 3
- Property changedPaperWhen2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models
When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models: arxiv announce type changed from cross to new
Arxiv announce typecross→newarxiv - Property changedPaperWhen2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models
When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models: published at changed from 2026-09-17T00:00:00+00:00 to 2026-09-18T04:00:00+00:00
Published17 Sept 2026→18 Sept 2026arxiv - New paperPaperWhen2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models
New paper: When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models
huggingface
Sources
Sources 3
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.