A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and Distillation
Updated 6 h ago · first seen 11 Sept 2026
paper_01M294G5V4A4327GCDHMGD5TR3
- Published
- 11 Sept 2026
- T1 · 6 h ago
- arXiv
- 2605.12227
- T1 · 6 h ago
- Category
- cs.CL
- T1 · 6 h ago
Abstract
Existing approaches to post-train models for long-context tasks face complementary limitations: (i) supervised fine-tuning (SFT) provides stable supervision but suffers from exposure bias; (ii) reinforcement learning methods such as Group Relative Policy Optimization (GRPO) train on model-generated trajectories but struggle with long-horizon credit assignment and sparse rewards; and (iii) on-policy distillation (OPD) provides dense token-level guidance but does not directly optimize task rewards. We study these complementary strategies for long-context alignment and derive a recipe that combines GRPO with OPD-style teacher guidance: the student learns from its own rollouts using outcome-level rewards, while a stronger teacher provides dense token-level regularization in place of the standard reference policy. This is especially useful when process-level supervision is difficult to obtain. To support this study, we introduce LongBlocks, a synthetic multilingual dataset spanning multi-hop reasoning, contextual grounding, and long-form generation. Through controlled ablations, we isolate the roles of cold-start initialization, teacher anchoring, and data mixing, showing that our recipe yields a more stable and effective path to long-context reasoning than GRPO or OPD while preserving short-context capabilities.
Authors 3
Miguel Moura Ramos, Duarte M. Alves, Andr\'e F. T. Martins
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 6 h agohigh
- Arxiv announce type
- replace
Source:arXiv (Atom API + RSS)T1observed 6 h agohigh
- arXiv id
- 2605.12227
Source:arXiv (Atom API + RSS)T1observed 6 h agohigh
- Categories
- cs.CL
Source:arXiv (Atom API + RSS)T1observed 6 h agohigh
Source:arXiv (Atom API + RSS)T1observed 6 h agohigh
- Primary category
- cs.CL
Source:arXiv (Atom API + RSS)T1observed 6 h agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 6 h agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
6 h ago
Conflicts
None
No models linked to this paper yet.
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · Abstract
Abstractabstract1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| Existing approaches to post-train models for long-context tasks face complementary limitations: (i) supervised fine-tuning (SFT) provides stable supervision but suffers from exposure bias; (ii) reinforcement learning methods such as Group Relative Policy Optimization (GRPO) train on model-generated trajectories but struggle with long-horizon credit assignment and sparse rewards; and (iii) on-policy distillation (OPD) provides dense token-level guidance but does not directly optimize task rewards. We study these complementary strategies for long-context alignment and derive a recipe that combines GRPO with OPD-style teacher guidance: the student learns from its own rollouts using outcome-level rewards, while a stronger teacher provides dense token-level regularization in place of the standard reference policy. This is especially useful when process-level supervision is difficult to obtain. To support this study, we introduce LongBlocks, a synthetic multilingual dataset spanning multi-hop reasoning, contextual grounding, and long-form generation. Through controlled ablations, we isolate the roles of cold-start initialization, teacher anchoring, and data mixing, showing that our recipe yields a more stable and effective path to long-context reasoning than GRPO or OPD while preserving short-context capabilities. | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
- New paperPaperA Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and Distillation
New paper: A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and Distillation
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.CL | feed | T1· Official | 4 h ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.