Open-UniMo: Towards Unified Motion-Language Understanding and Generation in the Open World
Published 16 Sept 2026arXiv:2609.14615
Updated 7 h ago · first seen 15 Sept 2026
paper_01M2JK19B07R1PRBRNPDAJQS86
Abstract
Unified motion generation and understanding is crucial for embodied AI systems that can both synthesize and interpret human actions in open-world environments. Existing motion-language models often treat motion as an auxiliary modality of a language model, leading to text-dominated representations and limited cross-modal interaction. Moreover, the next-token prediction paradigm is not naturally suited to long motion sequences, where autoregressive generation may accumulate prediction errors. To address these challenges, we propose Open-UniMo, a unified Large Motion-Language Model (LMLM) trained on million-scale open-world motion-language data. Open-UniMo promotes modality parity by extending Qwen's vocabulary of about 150K text tokens with 64K motion tokens, enabling motion and language to share a unified token space. We further introduce motion-consistent Chain-of-Thought reasoning as an intermediate representation to bridge language semantics and motion dynamics. Open-UniMo is trained with a two-stage pipeline, where supervised fine-tuning establishes CoT-guided bidirectional motion-language mapping and Group Relative Policy Optimization (GRPO) improves semantic alignment while mitigating cumulative errors in autoregressive motion-token generation. To support comprehensive evaluation, we propose Open-MoBench, a VLM-guided benchmark for assessing text-to-motion (T2M) generation, motion-to-text (M2T) understanding, and bidirectional consistency. Extensive experiments show that Open-UniMo achieves state-of-the-art performance on both conventional metrics and Open-MoBench. Furthermore, ablation studies reveal that M2T understanding is not primarily limited by motion-token vocabulary size; instead, coupling M2T with the learnable T2M generation path yields stronger cross-modal representations, demonstrating that generation can facilitate understanding in AR-based motion-language modeling.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 4
- Property changedPaperOpen-UniMo: Towards Unified Motion-Language Understanding and Generation in the Open World
Open-UniMo: Towards Unified Motion-Language Understanding and Generation in the Open World: arxiv announce type changed from new to cross
Arxiv announce typenew→crossarxiv - Property changedPaperOpen-UniMo: Towards Unified Motion-Language Understanding and Generation in the Open World
Open-UniMo: Towards Unified Motion-Language Understanding and Generation in the Open World: published at changed from 2026-09-15T04:00:00+00:00 to 2026-09-16T04:00:00+00:00
Published15 Sept 2026→16 Sept 2026arxiv - Property changedPaperOpen-UniMo: Towards Unified Motion-Language Understanding and Generation in the Open World
Open-UniMo: Towards Unified Motion-Language Understanding and Generation in the Open World: arxiv announce type changed from cross to new
Arxiv announce typecross→newarxiv - New paperPaperOpen-UniMo: Towards Unified Motion-Language Understanding and Generation in the Open World
New paper: Open-UniMo: Towards Unified Motion-Language Understanding and Generation in the Open World
arxiv
Sources
Sources 2
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.