SalsaAgent: A multimodal embodied language model for interactive dance generation
Published 18 Sept 2026arXiv:2605.29219
Updated 4 h ago · first seen 18 Sept 2026
paper_01M2SEHERX30SDKK4C7JH2NX5D
Abstract
Embodied interaction with humanoids depends on bidirectional nonverbal reactivity, coordination, and synchrony to convey cues and move with a partner. For socially interactive embodied agents, reactive motion generation requires expressive full-body motion that remains contextually appropriate while maintaining spatial and temporal synchrony. We present SalsaAgent, a language model that generates expressive, full-body salsa follower motions in reaction to a human leader and music. We formulate partner interaction as nonverbal token passing, extending the vocabulary of a large language model (LLM) to process discrete motion tokens, pairwise relation tokens, and audio tokens. Our method introduces full-body and pairwise-relation tokenizers, aligns language and motion tokens with automatically derived text descriptions of skeleton dynamics, and applies a two-stage token-to-diffusion pipeline. Subjective and objective evaluations show improved motion quality, two-person spatial coordination, and music and partner coordination relative to prior baselines.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 1
New paper: SalsaAgent: A multimodal embodied language model for interactive dance generation
arxiv
Sources
Sources 1
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.