Memory as Plans: World-Action Modeling with Memory-Grounded Planning
Updated 1 h ago · first seen 11 Sept 2026
paper_01M294WYDW0E7YQQXXQ2TKGHC9
- Published
- 10 Sept 2026
- T2 · 7 h ago
- arXiv
- 2609.11561
- T2 · 7 h ago
Abstract
Mainstream robotic policies often adopt a Markovian formulation, but many complex real-world manipulation tasks are inherently non-Markovian, requiring long-horizon memory beyond the current observation. Existing memory mechanisms often rely on language summaries, growing visual windows, or their combinations, and may therefore lose fine-grained visual evidence or face a trade-off between history coverage and execution efficiency. We introduce MaP-WAM, a Memory-as-Plans framework that decomposes memory-dependent world-action modeling into memory-grounded planning and plan-conditioned execution, and uses long-term multimodal episodic context as planning-time evidence rather than repeatedly conditioning the executor on the full history. MaP-WAM represents memory as completed segment records containing language instructions and sparse visual context, and converts this episodic memory into compact plans comprising the next segment-level language plan and corresponding visual guidance. A World-Action-Progress (WAP) model executes each plan over an unknown duration by jointly predicting action chunks and corresponding execution progress at inference time, calibrating predicted progress through plan-observation alignment for adaptive segment transitions and closed-loop context updates. MaP-WAM keeps the executor context length fixed, while structured attention further enables key-value caching in both planning and execution. MaP-WAM achieves state-of-the-art performance on RMBench with an 83.3% success rate and attains 78.0% success on real-robot tasks, while maintaining approximately constant executor inference latency as task history grows.
Authors 8
Sizhe Zhao, Haozhe Xie, Weiyu Zhao, Chenchu Zhang, Huan Wang, Chenyang Wang, Qinglin Liu, Shengping Zhang
Specification
- Official page
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 7 h agomedium
- arXiv id
- 2609.11561
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 7 h agomedium
- Github repo
- aipixel/MaP-WAM
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 7 h agomedium
- Hf paper url
- https://huggingface.co/papers/2609.11561
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 7 h agomedium
- Github stars
- 1
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 1 h agomedium
- Hf comments
- 2
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 1 h agomedium
- Upvotes
- 10
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 1 h agomedium
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 7 h agomedium
- Published
- 10 Sept 2026
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 7 h agomedium
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
11
Source tiers
T211
Freshest observation
1 h ago
Conflicts
None
No models linked to this paper yet.
No relations recorded.
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · Abstract
Abstractabstract1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| Mainstream robotic policies often adopt a Markovian formulation, but many complex real-world manipulation tasks are inherently non-Markovian, requiring long-horizon memory beyond the current observation. Existing memory mechanisms often rely on language summaries, growing visual windows, or their combinations, and may therefore lose fine-grained visual evidence or face a trade-off between history coverage and execution efficiency. We introduce MaP-WAM, a Memory-as-Plans framework that decomposes memory-dependent world-action modeling into memory-grounded planning and plan-conditioned execution, and uses long-term multimodal episodic context as planning-time evidence rather than repeatedly conditioning the executor on the full history. MaP-WAM represents memory as completed segment records containing language instructions and sparse visual context, and converts this episodic memory into compact plans comprising the next segment-level language plan and corresponding visual guidance. A World-Action-Progress (WAP) model executes each plan over an unknown duration by jointly predicting action chunks and corresponding execution progress at inference time, calibrating predicted progress through plan-observation alignment for adaptive segment transitions and closed-loop context updates. MaP-WAM keeps the executor context length fixed, while structured attention further enables key-value caching in both planning and execution. MaP-WAM achieves state-of-the-art performance on RMBench with an 83.3% success rate and attains 78.0% success on real-robot tasks, while maintaining approximately constant executor inference latency as task history grows. | → current | current | Hugging Face Hub (public pages, model cards, papers)T2 | medium | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
New paper: Memory as Plans: World-Action Modeling with Memory-Grounded Planning
huggingface
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| Hugging Face Hub (public pages, model cards, papers) | huggingface.co/papers | listing | T2· Quality secondary | 1 h ago | 6 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.