Skip to content
AI Atlas
PaperActive

World in World: Explore the World with World Models

arxiv.org/abs/2609.11548

Updated 42 min ago · first seen 11 Sept 2026

paper_01M294H28V2G2MHJ5DVRDX2A7E

Published
11 Sept 2026
T1 · 51 min ago
arXiv
2609.11548
T1 · 51 min ago
Category
cs.CV
T1 · 51 min ago

Abstract

Autoregressive video world models enable interactive, long-horizon exploration, but flexible control remains challenging. Exploring a source video from new viewpoints requires the generated rollout to remain synchronised with the recorded event, place observed content in the requested view, plausibly complete newly exposed regions, and recover previously generated appearance on revisits. Existing methods typically address these requirements through task-specific modules or additional training. We present World in World, a training-free inference-time interface that converts heterogeneous control evidence into camera- and time-labelled clean visual states, which are read through the native self attention of a frozen causal video model. The evidence comprises source-video observations, target-view scene projections, geometry renderings that guide completion of newly exposed subject regions, and retrieved generated states beyond the rolling cache. Each evidence source carries token-level support and its own availability schedule. A correspondence router combines persistent point identities with geometry to establish token correspondences, guiding supported queries towards matching source-video tokens. Evidence-wise attention CFG (EWA) then independently regulates each auxiliary channel's additional contribution using attention responses from the same denoising forward pass. The shared interface supports camera-controlled rerendering, long-horizon revisiting, and human-motion transfer with the same frozen backbone. We evaluate World in World on camera-controlled video rerendering under diverse viewpoint changes, assessing perceptual quality, temporal consistency, and camera-following accuracy.

Authors 3

Chenxi Song, Yanming Yang, Chi Zhang

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

arXiv id
2609.11548

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Categories
cs.CV

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Hf paper url
https://huggingface.co/papers/2609.11548

Source:Hugging Face Hub (public pages, model cards, papers)T2observed 44 min agomedium

Hf comments
0

Source:Hugging Face Hub (public pages, model cards, papers)T2observed 42 min agomedium

Upvotes
5

Source:Hugging Face Hub (public pages, model cards, papers)T2observed 42 min agomedium

PDF

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Primary category
cs.CV

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

12

Source tiers

T1T29 / 3

Freshest observation

42 min ago

Conflicts

2 flagged