Skip to content
AI Atlas
PaperActive

Memory as Plans: World-Action Modeling with Memory-Grounded Planning

arxiv.org/abs/2609.11561

quality59

Updated 4 h ago · first seen 11 Sept 2026

paper_01M294WYDW0E7YQQXXQ2TKGHC9

Published
10 Sept 2026
T2 · 6 h ago
arXiv
2609.11561
T2 · 6 h ago

Abstract

Mainstream robotic policies often adopt a Markovian formulation, but many complex real-world manipulation tasks are inherently non-Markovian, requiring long-horizon memory beyond the current observation. Existing memory mechanisms often rely on language summaries, growing visual windows, or their combinations, and may therefore lose fine-grained visual evidence or face a trade-off between history coverage and execution efficiency. We introduce MaP-WAM, a Memory-as-Plans framework that decomposes memory-dependent world-action modeling into memory-grounded planning and plan-conditioned execution, and uses long-term multimodal episodic context as planning-time evidence rather than repeatedly conditioning the executor on the full history. MaP-WAM represents memory as completed segment records containing language instructions and sparse visual context, and converts this episodic memory into compact plans comprising the next segment-level language plan and corresponding visual guidance. A World-Action-Progress (WAP) model executes each plan over an unknown duration by jointly predicting action chunks and corresponding execution progress at inference time, calibrating predicted progress through plan-observation alignment for adaptive segment transitions and closed-loop context updates. MaP-WAM keeps the executor context length fixed, while structured attention further enables key-value caching in both planning and execution. MaP-WAM achieves state-of-the-art performance on RMBench with an 83.3% success rate and attains 78.0% success on real-robot tasks, while maintaining approximately constant executor inference latency as task history grows.

Authors 8

Sizhe Zhao, Haozhe Xie, Weiyu Zhao, Chenchu Zhang, Huan Wang, Chenyang Wang, Qinglin Liu, Shengping Zhang

Specification

arXiv id
2609.11561

Source:Hugging Face Hub (public pages, model cards, papers)T2observed 6 h agomedium

Github repo
aipixel/MaP-WAM

Source:Hugging Face Hub (public pages, model cards, papers)T2observed 6 h agomedium

Hf paper url
https://huggingface.co/papers/2609.11561

Source:Hugging Face Hub (public pages, model cards, papers)T2observed 6 h agomedium

Github stars
0

Source:Hugging Face Hub (public pages, model cards, papers)T2observed 4 h agomedium

Hf comments
1

Source:Hugging Face Hub (public pages, model cards, papers)T2observed 4 h agomedium

Upvotes
9

Source:Hugging Face Hub (public pages, model cards, papers)T2observed 4 h agomedium

Published
10 Sept 2026

Source:Hugging Face Hub (public pages, model cards, papers)T2observed 6 h agomedium

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

11

Source tiers

T211

Freshest observation

4 h ago

Conflicts

None