Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data
Updated 52 min ago · first seen 11 Sept 2026
paper_01M294FPJHCWSE05SGFZ6ZK7NC
- Published
- 11 Sept 2026
- T1 · 52 min ago
- arXiv
- 2609.11917
- T1 · 52 min ago
- Category
- cs.LG
- T1 · 52 min ago
Abstract
As the supply of human-written text is exhausted, it has become standard practice to repeat language model training data. Prior work has studied data repetition for densely activated Transformers, but the effects of data repetition remains largely unexplored for recently dominant sparse architectures such as Mixture-of-Experts (MoE), despite their increased compute efficiency. We vary data repetition rates across single- and multi-domain data mixes, and across MoE settings, including expert count and granularity. We consistently find, for models ranging from 80M to 1B active (8.5B total) parameters, that MoEs degrade more rapidly under data repetition. This effect increases with sparsity, dictated by total rather than active parameters. While 80M dense models can repeat data over 8x with minimal degradation, MoEs instead begin to suffer at 4x, and deteriorate rapidly, ceding their performance benefits in all-unique data settings to underperform dense models after 32x. We experiment with existing regularization methods as a potential remedy. We find that some methods, such as dropout, can mitigate overfitting. In particular, with strong masking-based regularization, MoEs are able to outperform dense models even when data is repeated more than 64 times. However, no method fully matches the performance of all-unique training data. Finally, we analyze internal mechanisms correlated with MoE overfitting in high repetition regimes, and find that MoE routing universally stabilizes early in training, and that expert specialization correlates with overfitting to repeated data. In sum, our work addresses the adverse interactions between sparsity and data repetition: we present evidence for the core mechanisms of overfitting and its potential remediation, and suggest promising avenues for future methods to reduce over-specialization in model parameters by disrupting memorization patterns.
Authors 5
Atindra Jha, Margaret Li, Jure Leskovec, Percy Liang, Luke Zettlemoyer
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 52 min agohigh
- Arxiv announce type
- cross
Source:arXiv (Atom API + RSS)T1observed 52 min agohigh
- arXiv id
- 2609.11917
Source:arXiv (Atom API + RSS)T1observed 52 min agohigh
- Categories
- cs.LG, cs.CL
Source:arXiv (Atom API + RSS)T1observed 52 min agohigh
Source:arXiv (Atom API + RSS)T1observed 52 min agohigh
- Primary category
- cs.LG
Source:arXiv (Atom API + RSS)T1observed 52 min agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 52 min agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
52 min ago
Conflicts
None
No models linked to this paper yet.
- Authors
- Atindra Jha, Margaret Li, Jure Leskovec
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · Abstract
Abstractabstract1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| As the supply of human-written text is exhausted, it has become standard practice to repeat language model training data. Prior work has studied data repetition for densely activated Transformers, but the effects of data repetition remains largely unexplored for recently dominant sparse architectures such as Mixture-of-Experts (MoE), despite their increased compute efficiency. We vary data repetition rates across single- and multi-domain data mixes, and across MoE settings, including expert count and granularity. We consistently find, for models ranging from 80M to 1B active (8.5B total) parameters, that MoEs degrade more rapidly under data repetition. This effect increases with sparsity, dictated by total rather than active parameters. While 80M dense models can repeat data over 8x with minimal degradation, MoEs instead begin to suffer at 4x, and deteriorate rapidly, ceding their performance benefits in all-unique data settings to underperform dense models after 32x. We experiment with existing regularization methods as a potential remedy. We find that some methods, such as dropout, can mitigate overfitting. In particular, with strong masking-based regularization, MoEs are able to outperform dense models even when data is repeated more than 64 times. However, no method fully matches the performance of all-unique training data. Finally, we analyze internal mechanisms correlated with MoE overfitting in high repetition regimes, and find that MoE routing universally stabilizes early in training, and that expert specialization correlates with overfitting to repeated data. In sum, our work addresses the adverse interactions between sparsity and data repetition: we present evidence for the core mechanisms of overfitting and its potential remediation, and suggest promising avenues for future methods to reduce over-specialization in model parameters by disrupting memorization patterns. | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
- Property changedPaperData Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data
Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data: arxiv announce type changed from new to cross
Arxiv announce typenew→crossarxiv New paper: Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.CL | feed | T1· Official | 52 min ago | 1 |
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.LG | feed | T1· Official | 52 min ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.