Skip to content
AI Atlas
PaperActive

Structural priors for data-efficient language learning

arxiv.org/abs/2609.11505

Updated 20 min ago · first seen 11 Sept 2026

paper_01M294FR0E08S95M9EVHFJXZEX

Published
11 Sept 2026
T1 · 20 min ago
arXiv
2609.11505
T1 · 20 min ago
Category
cs.CL
T1 · 20 min ago

Abstract

Efficient language learning requires methods to reduce the reliance on large data and computational resources. We investigate structural transfer: First training models on non-language data to induce useful priors for natural language. This approach is a form of weight initialization for multilingual language modeling. We evaluate transfer via next-token-prediction loss, weight shifts in the model, and downstream linguistic benchmarks. Several symbolic data types - notably music, probabilistic grammars, and cellular automata - yield lower language-modeling loss than random initialization. These gains coincide with smaller weight shifts during subsequent language training, suggesting that structural transfer positions models in a more favorable region of the parameter space. However, a lower loss does not translate consistently into better downstream linguistic performance, and transfer from non-language data is less efficient than additional language data. We conclude that non-language data can serve as a partial substitute for language data for the training objective of next-token prediction but does not reliably support broader linguistic generalization.

Authors 4

Yana Veitsman, Jonas Mayer Martins, Jonathan Lautenschlager, Lisa Beinborn

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 20 min agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 20 min agohigh

arXiv id
2609.11505

Source:arXiv (Atom API + RSS)T1observed 20 min agohigh

Categories
cs.CL, cs.AI, cs.LG

Source:arXiv (Atom API + RSS)T1observed 20 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 20 min agohigh

Primary category
cs.CL

Source:arXiv (Atom API + RSS)T1observed 20 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 20 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

20 min ago

Conflicts

None