Updated 53 min ago · first seen 11 Sept 2026
paper_01M294FR0E08S95M9EVHFJXZEX
- Published
- 11 Sept 2026
- T1 · 53 min ago
- arXiv
- 2609.11505
- T1 · 53 min ago
- Category
- cs.CL
- T1 · 53 min ago
Abstract
Efficient language learning requires methods to reduce the reliance on large data and computational resources. We investigate structural transfer: First training models on non-language data to induce useful priors for natural language. This approach is a form of weight initialization for multilingual language modeling. We evaluate transfer via next-token-prediction loss, weight shifts in the model, and downstream linguistic benchmarks. Several symbolic data types - notably music, probabilistic grammars, and cellular automata - yield lower language-modeling loss than random initialization. These gains coincide with smaller weight shifts during subsequent language training, suggesting that structural transfer positions models in a more favorable region of the parameter space. However, a lower loss does not translate consistently into better downstream linguistic performance, and transfer from non-language data is less efficient than additional language data. We conclude that non-language data can serve as a partial substitute for language data for the training objective of next-token prediction but does not reliably support broader linguistic generalization.
Authors 4
Yana Veitsman, Jonas Mayer Martins, Jonathan Lautenschlager, Lisa Beinborn
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 53 min agohigh
- Arxiv announce type
- new
Source:arXiv (Atom API + RSS)T1observed 53 min agohigh
- arXiv id
- 2609.11505
Source:arXiv (Atom API + RSS)T1observed 53 min agohigh
- Categories
- cs.CL, cs.AI, cs.LG
Source:arXiv (Atom API + RSS)T1observed 53 min agohigh
Source:arXiv (Atom API + RSS)T1observed 53 min agohigh
- Primary category
- cs.CL
Source:arXiv (Atom API + RSS)T1observed 53 min agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 53 min agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
53 min ago
Conflicts
None
No models linked to this paper yet.
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · arXiv id
arXiv idarxiv_id1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| 2609.11505 | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
Structural priors for data-efficient language learning: arxiv announce type changed from cross to new
Arxiv announce typecross→newarxivNew paper: Structural priors for data-efficient language learning
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.CL | feed | T1· Official | 53 min ago | 1 |
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.LG | feed | T1· Official | 53 min ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.