Measuring the Value of World-Model Updates: A Counterfactual Utility Protocol for Continual Adaptation
Updated 46 min ago · first seen 11 Sept 2026
paper_01M294FNWP8XJS0CH00ZRVBXJN
- Published
- 11 Sept 2026
- T1 · 46 min ago
- arXiv
- 2609.10954
- T1 · 46 min ago
- Category
- cs.LG
- T1 · 46 min ago
Abstract
Continual world models must decide whether new data justify changing the model. Fixed replay schedules and prediction-error triggers specify when to update, but neither reveals the value of an individual update: one deployment run cannot show how the same model would have performed at that moment had it held its parameters. We introduce the fork ledger, which branches a deployment stream at pre-registered decision points into matched update and hold continuations under common random numbers. It evaluates both continuations on the same episodes and records $\Delta R = R_{\mathrm{update}} - R_{\mathrm{hold}}$. Always applying one fixed update mechanism lowers return on all three simulated control tasks: CartPole ($-144.0$; checkpoint-bootstrap $95\%$ CI $[-185.4,-116.1]$, against a converged return near $650$), Walker ($-82.8$; $[-101.1,-61.7]$) and Cheetah ($-18.6$; $[-29.0,-6.6]$). Divergence is an outcome of applying the update, so the estimand counts every attempted fork; restricted to the $693$ of $720$ that did not collapse, CartPole and Walker are unchanged in sign ($-113.4$ and $-82.1$) and Cheetah becomes unresolved ($-3.9$; $[-17.5,+13.0]$). The task is the unit of inference: each contributes $240$ attempted forks over five pretrained checkpoints crossed with two drift directions. The ledger makes counterfactual utility observable for a fixed mechanism, allowing triggers to be judged by the updates they select rather than by surprise detection alone.
Authors 2
Anqi Peter Li, Kaden Kim
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 46 min agohigh
- Arxiv announce type
- new
Source:arXiv (Atom API + RSS)T1observed 46 min agohigh
- arXiv id
- 2609.10954
Source:arXiv (Atom API + RSS)T1observed 46 min agohigh
- Categories
- cs.LG
Source:arXiv (Atom API + RSS)T1observed 46 min agohigh
Source:arXiv (Atom API + RSS)T1observed 46 min agohigh
- Primary category
- cs.LG
Source:arXiv (Atom API + RSS)T1observed 46 min agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 46 min agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
46 min ago
Conflicts
None
No models linked to this paper yet.
- Authors
- Anqi Peter Li, Kaden Kim
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · Categories
Categoriescategories1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| cs.LG | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
- New paperPaperMeasuring the Value of World-Model Updates: A Counterfactual Utility Protocol for Continual Adaptation
New paper: Measuring the Value of World-Model Updates: A Counterfactual Utility Protocol for Continual Adaptation
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.LG | feed | T1· Official | 46 min ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.