Skip to content
AI Atlas
PaperActive

CHRONOBERG: Capturing Language Evolution and Temporal Awareness in Foundation Models

arxiv.org/abs/2509.22360

quality89

Updated 6 h ago · first seen 11 Sept 2026

paper_01M294G5H0PBPDSWYZCZXHBKJ7

Published
11 Sept 2026
T1 · 6 h ago
arXiv
2509.22360
T1 · 6 h ago
Category
cs.CL
T1 · 6 h ago

Abstract

Large language models (LLMs) excel at operating at scale by leveraging social media and various data crawled from the web. Whereas existing corpora are diverse, their frequent lack of long-term temporal structure may however limit an LLM's ability to contextualize semantic and normative evolution of language and to capture diachronic variation. To support analysis and training for the latter, we introduce CHRONOBERG, a temporally structured corpus of English book texts spanning 250 years, curated from Project Gutenberg and enriched with a variety of temporal annotations. First, the edited nature of books enables us to quantify lexical semantic change through time-sensitive Valence-Arousal-Dominance (VAD) analysis and to construct historically calibrated affective lexicons to support temporally grounded interpretation. With the lexicons at hand, we demonstrate a need for modern LLM-based tools to better situate their detection of discriminatory language and contextualization of sentiment across various time-periods. In fact, we show how language models trained sequentially on CHRONOBERG struggle to encode diachronic shifts in meaning, emphasizing the need for temporally aware training and evaluation pipelines, and positioning CHRONOBERG as a scalable resource for the study of linguistic change and temporal generalization. Disclaimer: This paper includes language and display of samples that could be offensive to readers. Open Access: Chronoberg is available publicly on HuggingFace at ( https://huggingface.co/datasets/spaul25/Chronoberg). Code is available at (https://github.com/paulsubarna/Chronoberg).

Authors 7

Niharika Hegde, Subarnaduti Paul, Lars Joel-Frey, Manuel Brack, Kristian Kersting, Martin Mundt, Patrick Schramowski

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

Arxiv announce type
replace

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

arXiv id
2509.22360

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

Categories
cs.CL, cs.AI

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

Primary category
cs.CL

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

6 h ago

Conflicts

None