Skip to content
AI Atlas
PaperActive

A Fragility Spectrum for Recursive Language-Model Training

arxiv.org/abs/2609.11149

Updated 21 min ago · first seen 11 Sept 2026

paper_01M294FQMZ0EWCFGZNWFKA4CDT

Published
11 Sept 2026
T1 · 21 min ago
arXiv
2609.11149
T1 · 21 min ago
Category
cs.CL
T1 · 21 min ago

Abstract

Model-generated text is finding its way back into training corpora, and there is plenty of evidence that training on such data over and over collapses output diversity. Prior work has studied the phenomenon itself: which protocols and which data mixtures cause collapse. But different models behave very differently under the same process. We fix one recursive contamination protocol and let 13 publicly released checkpoints form an ecosystem that shares a common corpus for five generations. The unique 4-gram outcome after five generations ranges from 0.187 to 0.940 across checkpoints, a roughly five-fold spread: some models are barely touched, others degenerate into repetitive fragments. Changing the composition of the shared pool or mixing in human text keeps the Spearman correlation of the ordering at 0.91--0.97, and changing the random seed keeps it at 0.93--0.98. Whether a model collapses easily under recursive training is, then, a property of the checkpoint itself, and one that has gone largely unexamined. Parameter scale alone does not explain it, since a three-size ladder within one family is not monotonic in size, and none of the static indicators we tested predicts it either. What does work is cheap: let a model iterate on its own output for two or three generations, and its fragility in the larger ecosystem can be inferred from that alone. Collapse speed also responds to intervention. Tightening top-p, which cuts the low-probability tail at generation time, nearly stops collapse within three generations and stabilizes six checkpoints spanning the whole spectrum together, while data-side filtering slows collapse without stopping it.

Authors 2

Yangze Liu, Zhongyi Han

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 21 min agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 21 min agohigh

arXiv id
2609.11149

Source:arXiv (Atom API + RSS)T1observed 21 min agohigh

Categories
cs.CL, cs.AI, cs.LG

Source:arXiv (Atom API + RSS)T1observed 21 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 21 min agohigh

Primary category
cs.CL

Source:arXiv (Atom API + RSS)T1observed 21 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 21 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

21 min ago

Conflicts

None