Characterizing Narrative Content in Web-scale LLM Pretraining Data
Updated 7 h ago · first seen 11 Sept 2026
paper_01M294G5Y40C085PV8MZRH88WQ
- Published
- 11 Sept 2026
- T1 · 7 h ago
- arXiv
- 2606.19468
- T1 · 7 h ago
- Category
- cs.CL
- T1 · 7 h ago
Abstract
The narrative composition of web-scale LLM pretraining corpora remains largely unexplored, even though narrative is a fundamental mode of human communication. We present the first fine-grained study of narrative features in Dolma, a 3-trillion-token open pretraining corpus. Drawing on narrative theory, we design a framework spanning three core narrative elements (agency, setting, and events) operationalized as 11 interpretable dimensions. After curating and hand-annotating a diverse set of 400 passages, we create an LLM-labeled dataset of 25K passages, and finally, we finetune and validate NarraBERT, two RoBERTa-based models for fine-grained narrative prediction. We apply NarraBERT to 13M passages, resulting in a new dataset, NarraDolma. We find that narrative structure is measurable at scale across extremely heterogeneous data and narrative qualities are unequally distributed across pretraining sources, topics, and formats in ways that current data curation practices neither measure nor account for. Our framework, dataset, and analyses provide a foundation for understanding how narrative qualities are distributed in LLM pretraining data.
Authors 4
Teagan Johnson, Elliott Ash, Andrew Piper, Maria Antoniak
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
- Arxiv announce type
- replace
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
- arXiv id
- 2606.19468
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
- Categories
- cs.CL
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
- Primary category
- cs.CL
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
7 h ago
Conflicts
None
No models linked to this paper yet.
- Authors
- Teagan Johnson, Elliott Ash, Andrew Piper
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · Published
Publishedpublished_at1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| 11 Sept 2026 | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
New paper: Characterizing Narrative Content in Web-scale LLM Pretraining Data
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.CL | feed | T1· Official | 5 h ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.