Skip to content
AI Atlas
PaperActive

Characterizing Narrative Content in Web-scale LLM Pretraining Data

arxiv.org/abs/2606.19468

quality89

Updated 5 h ago · first seen 11 Sept 2026

paper_01M294G5Y40C085PV8MZRH88WQ

Published
11 Sept 2026
T1 · 5 h ago
arXiv
2606.19468
T1 · 5 h ago
Category
cs.CL
T1 · 5 h ago

Abstract

The narrative composition of web-scale LLM pretraining corpora remains largely unexplored, even though narrative is a fundamental mode of human communication. We present the first fine-grained study of narrative features in Dolma, a 3-trillion-token open pretraining corpus. Drawing on narrative theory, we design a framework spanning three core narrative elements (agency, setting, and events) operationalized as 11 interpretable dimensions. After curating and hand-annotating a diverse set of 400 passages, we create an LLM-labeled dataset of 25K passages, and finally, we finetune and validate NarraBERT, two RoBERTa-based models for fine-grained narrative prediction. We apply NarraBERT to 13M passages, resulting in a new dataset, NarraDolma. We find that narrative structure is measurable at scale across extremely heterogeneous data and narrative qualities are unequally distributed across pretraining sources, topics, and formats in ways that current data curation practices neither measure nor account for. Our framework, dataset, and analyses provide a foundation for understanding how narrative qualities are distributed in LLM pretraining data.

Authors 4

Teagan Johnson, Elliott Ash, Andrew Piper, Maria Antoniak

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 5 h agohigh

Arxiv announce type
replace

Source:arXiv (Atom API + RSS)T1observed 5 h agohigh

arXiv id
2606.19468

Source:arXiv (Atom API + RSS)T1observed 5 h agohigh

Categories
cs.CL

Source:arXiv (Atom API + RSS)T1observed 5 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 5 h agohigh

Primary category
cs.CL

Source:arXiv (Atom API + RSS)T1observed 5 h agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 5 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

5 h ago

Conflicts

None