Negative Self-Distillation: Learning to Reason by Avoiding Flaws
Updated 11 min ago · first seen 11 Sept 2026
paper_01M294FR8NZQXY4Y4VP1WPXVNW
- Published
- 11 Sept 2026
- T1 · 21 min ago
- arXiv
- 2609.11699
- T1 · 21 min ago
- Category
- cs.CL
- T1 · 21 min ago
Abstract
On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. However, recent findings indicate that OPSD can severely degrade the performance of LLMs on complex reasoning tasks: By forcing the student to imitate an artificially confident reasoning trace conditioned on privileged information, OPSD inadvertently suppresses expressions of uncertainty and penalizes the exploratory, self-corrective behaviors required to solve challenging problems. To address this, we introduce Negative Self-Distillation (NSD), a new framework that optimizes LLMs by diverging from flawed reasoning rather than imitating privileged solutions. Instead of relying on ground-truth answers or external supervision, NSD uses the model itself to generate a question-specific negative condition (eg, acting as a ``careless reasoner'') and pushes the student's distribution away from this self-generated negative teacher. Naively applying unlearning objectives to achieve this divergence is problematic, as flawed reasoning tokens are confounded with basic linguistic tokens; indiscriminately penalizing both risks catastrophically degrading the model's foundational language capabilities. We resolve this by designing a dynamic gating mechanism that automatically identifies and isolates reasoning-critical tokens, ensuring gradient updates target only behavioral flaws while preserving the model's linguistic priors. Empirically, NSD consistently outperforms OPSD and other label-free, self-bootstrapping reinforcement learning (RL) baselines.
Authors 6
Rongcan Pei, Zhepei Wei, Shuyao Xu, Xinyu Zhu, Wei-Lin Chen, Yu Meng
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 21 min agohigh
- Arxiv announce type
- new
Source:arXiv (Atom API + RSS)T1observed 21 min agohigh
- arXiv id
- 2609.11699
Source:arXiv (Atom API + RSS)T1observed 21 min agohigh
- Categories
- cs.CL, cs.LG
Source:arXiv (Atom API + RSS)T1observed 21 min agohigh
- Github repo
- Prongcan/NSD
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 14 min agomedium
- Hf paper url
- https://huggingface.co/papers/2609.11699
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 14 min agomedium
- Github stars
- 4
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 11 min agomedium
- Hf comments
- 1
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 11 min agomedium
- Upvotes
- 4
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 11 min agomedium
Source:arXiv (Atom API + RSS)T1observed 21 min agohigh
- Primary category
- cs.CL
Source:arXiv (Atom API + RSS)T1observed 21 min agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 21 min agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
14
Source tiers
T1T29 / 5
Freshest observation
11 min ago
Conflicts
2 flagged
No models linked to this paper yet.
- Authors
- Rongcan Pei, Zhepei Wei, Shuyao Xu
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history
Official pageofficial_url1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://arxiv.org/abs/2609.11699 | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Abstractabstract1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. However, recent findings indicate that OPSD can severely degrade the performance of LLMs on complex reasoning tasks: By forcing the student to imitate an artificially confident reasoning trace conditioned on privileged information, OPSD inadvertently suppresses expressions of uncertainty and penalizes the exploratory, self-corrective behaviors required to solve challenging problems. To address this, we introduce Negative Self-Distillation (NSD), a new framework that optimizes LLMs by diverging from flawed reasoning rather than imitating privileged solutions. Instead of relying on ground-truth answers or external supervision, NSD uses the model itself to generate a question-specific negative condition (eg, acting as a ``careless reasoner'') and pushes the student's distribution away from this self-generated negative teacher. Naively applying unlearning objectives to achieve this divergence is problematic, as flawed reasoning tokens are confounded with basic linguistic tokens; indiscriminately penalizing both risks catastrophically degrading the model's foundational language capabilities. We resolve this by designing a dynamic gating mechanism that automatically identifies and isolates reasoning-critical tokens, ensuring gradient updates target only behavioral flaws while preserving the model's linguistic priors. Empirically, NSD consistently outperforms OPSD and other label-free, self-bootstrapping reinforcement learning (RL) baselines. | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Arxiv announce typearxiv_announce_type2
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| new | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
| cross | → | superseded | arXiv (Atom API + RSS)T1 | high | deterministic |
arXiv idarxiv_id1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| 2609.11699 | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Categoriescategories1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| cs.CL, cs.LG | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Github repogithub_repo1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| Prongcan/NSD | → current | current | Hugging Face Hub (public pages, model cards, papers)T2 | medium | deterministic |
Hf paper urlhf_paper_url1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://huggingface.co/papers/2609.11699 | → current | current | Hugging Face Hub (public pages, model cards, papers)T2 | medium | deterministic |
Github starsmetric.github_stars1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| 4 | → current | current | Hugging Face Hub (public pages, model cards, papers)T2 | medium | deterministic |
Hf commentsmetric.hf_comments1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| 1 | → current | current | Hugging Face Hub (public pages, model cards, papers)T2 | medium | deterministic |
Upvotesmetric.upvotes1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| 4 | → current | current | Hugging Face Hub (public pages, model cards, papers)T2 | medium | deterministic |
PDFpdf_url1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://arxiv.org/pdf/2609.11699 | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Primary categoryprimary_category1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| cs.CL | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Publishedpublished_at3conflicting claims
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| 10 Sept 2026 | → current | conflicting | Hugging Face Hub (public pages, model cards, papers)T2 | medium | deterministic |
| 10 Sept 2026 | → current | conflicting | Hugging Face Hub (public pages, model cards, papers)T2 | medium | deterministic |
| 11 Sept 2026 | → current | current | arXiv (Atom API + RSS)T1 | conflicted | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
Negative Self-Distillation: Learning to Reason by Avoiding Flaws: arxiv announce type changed from cross to new
Arxiv announce typecross→newarxivNew paper: Negative Self-Distillation: Learning to Reason by Avoiding Flaws
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.CL | feed | T1· Official | 21 min ago | 1 |
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.LG | feed | T1· Official | 21 min ago | 1 |
| Hugging Face Hub (public pages, model cards, papers) | huggingface.co/papers | listing | T2· Quality secondary | 11 min ago | 2 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.