The information geometry of large language models is shared, learned, and controllable
Updated 34 min ago · first seen 11 Sept 2026
paper_01M294FNZB0JKT6K4S8892MD0Z
- Published
- 11 Sept 2026
- T1 · 35 min ago
- arXiv
- 2609.11063
- T1 · 35 min ago
- Category
- cs.LG
- T1 · 35 min ago
Abstract
Large language models learn similar behaviours, yet it remains unclear what structure they share or how to change one behaviour without disturbing others. The Fisher-Rao geometry of next-token probabilities connects these questions: behaviour determines this geometry up to output-preserving symmetries, whereas activation geometry depends on coordinates. Across transformer, state-space and recurrent models, output geometries agree more strongly than activation geometries, and shared geometry supports semantic-category transfer. Agreement with human word choices increases with predictive accuracy, scale and training, and improves further after model-only calibration. Token probabilities and read-out geometry jointly predict the spectrum and its effective dimension. Controlled language assignments show that geometry follows the language law across architectures. Pretraining corpus statistics predict held-out fact acquisition without recalibration, while randomised experiments show that deeper evidence substantially delays acquisition across every tested architecture and evidence construction. Finally, the geometry prescribes minimum-disturbance local interventions, predicts their relative cost, and supports reusable control: updates learned on donor prompts transfer to unseen prompts while better preserving behaviour on reference prompts than Euclidean control. The same geometric correction improves steering, editing, attribution, dictionary learning and fine-tuning.
Authors 1
Dario Picozzi
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 35 min agohigh
- Arxiv announce type
- cross
Source:arXiv (Atom API + RSS)T1observed 35 min agohigh
- arXiv id
- 2609.11063
Source:arXiv (Atom API + RSS)T1observed 35 min agohigh
- Categories
- cs.LG, cs.CL
Source:arXiv (Atom API + RSS)T1observed 35 min agohigh
Source:arXiv (Atom API + RSS)T1observed 35 min agohigh
- Primary category
- cs.LG
Source:arXiv (Atom API + RSS)T1observed 35 min agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 35 min agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
35 min ago
Conflicts
None
No models linked to this paper yet.
- Authors
- Dario Picozzi
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · Authors
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
- Property changedPaperThe information geometry of large language models is shared, learned, and controllable
The information geometry of large language models is shared, learned, and controllable: arxiv announce type changed from new to cross
Arxiv announce typenew→crossarxiv - New paperPaperThe information geometry of large language models is shared, learned, and controllable
New paper: The information geometry of large language models is shared, learned, and controllable
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.CL | feed | T1· Official | 34 min ago | 1 |
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.LG | feed | T1· Official | 35 min ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.