RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
Published 18 Sept 2026arXiv:2609.20784
Updated 4 h ago · first seen 17 Sept 2026
paper_01M2SE0KQTDZJEZZ4XB31FETCR
Abstract
Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision from a self-teacher with privileged task skills, letting a skill-free student internalize them. This recipe, however, is undermined by two findings in agentic tasks: privileged information alone does not always make a teacher reliable, and the benefit of teacher supervision is stage-dependent. We therefore propose RetireOPD (Self-Retiring On-Policy Distillation), which first optimizes a decoupled, skill-conditioned teacher with environment rewards and then trains a skill-free student jointly with RL and OPD. Rather than following a predefined distillation schedule, RetireOPD adopts Adaptive Retirement: the student drops the teacher on its own once their discrepancy stops shrinking and it reaches a target fraction of the teacher's success rate, after which training proceeds with RL alone. Across Qwen2.5 models from 1.5B to 7B, RetireOPD improves ALFWorld success rate over RL baseline by 14.1% to 18.8% and WebShop accuracy by 11.8% to 19.0%, and surpasses its own skill-conditioned teacher in every setting.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 3
- Property changedPaperRetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning: arxiv announce type changed from new to cross
Arxiv announce typenew→crossarxiv - Property changedPaperRetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning: published at changed from 2026-09-17T00:00:00+00:00 to 2026-09-18T04:00:00+00:00
Published17 Sept 2026→18 Sept 2026arxiv New paper: RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
huggingface
Sources
Sources 3
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.