Skip to content
AI Atlas
PaperActive

Negative Self-Distillation: Learning to Reason by Avoiding Flaws

arxiv.org/abs/2609.11699

Updated 3 h ago · first seen 11 Sept 2026

paper_01M294FR8NZQXY4Y4VP1WPXVNW

Published
11 Sept 2026
T1 · 3 h ago
arXiv
2609.11699
T1 · 3 h ago
Category
cs.CL
T1 · 3 h ago

As of

Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.

Claim history

17 claims · 14 properties2 conflicting

Official pageofficial_url1

Claim history for Official page
ValueValid from → toStatusSourceConfidenceExtractor
https://arxiv.org/abs/2609.11699currentcurrentarXiv (Atom API + RSS)T1highdeterministic

Abstractabstract1

Claim history for Abstract
ValueValid from → toStatusSourceConfidenceExtractor
On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. However, recent findings indicate that OPSD can severely degrade the performance of LLMs on complex reasoning tasks: By forcing the student to imitate an artificially confident reasoning trace conditioned on privileged information, OPSD inadvertently suppresses expressions of uncertainty and penalizes the exploratory, self-corrective behaviors required to solve challenging problems. To address this, we introduce Negative Self-Distillation (NSD), a new framework that optimizes LLMs by diverging from flawed reasoning rather than imitating privileged solutions. Instead of relying on ground-truth answers or external supervision, NSD uses the model itself to generate a question-specific negative condition (eg, acting as a ``careless reasoner'') and pushes the student's distribution away from this self-generated negative teacher. Naively applying unlearning objectives to achieve this divergence is problematic, as flawed reasoning tokens are confounded with basic linguistic tokens; indiscriminately penalizing both risks catastrophically degrading the model's foundational language capabilities. We resolve this by designing a dynamic gating mechanism that automatically identifies and isolates reasoning-critical tokens, ensuring gradient updates target only behavioral flaws while preserving the model's linguistic priors. Empirically, NSD consistently outperforms OPSD and other label-free, self-bootstrapping reinforcement learning (RL) baselines.currentcurrentarXiv (Atom API + RSS)T1highdeterministic

Arxiv announce typearxiv_announce_type2

Claim history for Arxiv announce type
ValueValid from → toStatusSourceConfidenceExtractor
newcurrentcurrentarXiv (Atom API + RSS)T1highdeterministic
crosssupersededarXiv (Atom API + RSS)T1highdeterministic

arXiv idarxiv_id1

Claim history for arXiv id
ValueValid from → toStatusSourceConfidenceExtractor
2609.11699currentcurrentarXiv (Atom API + RSS)T1highdeterministic

Authorsauthors1

Claim history for Authors
ValueValid from → toStatusSourceConfidenceExtractor
Rongcan Pei, Zhepei Wei, Shuyao Xu, Xinyu Zhu, Wei-Lin Chen, Yu MengcurrentcurrentarXiv (Atom API + RSS)T1highdeterministic

Categoriescategories1

Claim history for Categories
ValueValid from → toStatusSourceConfidenceExtractor
cs.CL, cs.LGcurrentcurrentarXiv (Atom API + RSS)T1highdeterministic

Github repogithub_repo1

Claim history for Github repo
ValueValid from → toStatusSourceConfidenceExtractor
Prongcan/NSDcurrentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Hf paper urlhf_paper_url1

Claim history for Hf paper url
ValueValid from → toStatusSourceConfidenceExtractor
https://huggingface.co/papers/2609.11699currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Github starsmetric.github_stars1

Claim history for Github stars
ValueValid from → toStatusSourceConfidenceExtractor
4currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Hf commentsmetric.hf_comments1

Claim history for Hf comments
ValueValid from → toStatusSourceConfidenceExtractor
1currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Upvotesmetric.upvotes1

Claim history for Upvotes
ValueValid from → toStatusSourceConfidenceExtractor
4currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

PDFpdf_url1

Claim history for PDF
ValueValid from → toStatusSourceConfidenceExtractor
https://arxiv.org/pdf/2609.11699currentcurrentarXiv (Atom API + RSS)T1highdeterministic

Primary categoryprimary_category1

Claim history for Primary category
ValueValid from → toStatusSourceConfidenceExtractor
cs.CLcurrentcurrentarXiv (Atom API + RSS)T1highdeterministic

Publishedpublished_at3conflicting claims

Claim history for Published
ValueValid from → toStatusSourceConfidenceExtractor
10 Sept 2026currentconflictingHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic
10 Sept 2026currentconflictingHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic
11 Sept 2026currentcurrentarXiv (Atom API + RSS)T1conflicteddeterministic

Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →