Skip to content
AI Atlas
PaperActive

M3-Former: Multimodal Transformer with Mixture-of-Experts for Long-Term Vessel Trajectory Prediction

arxiv.org/abs/2609.10559

Updated 15 min ago · first seen 11 Sept 2026

paper_01M294FNPGGF0W327BGERR28JE

Published
11 Sept 2026
T1 · 15 min ago
arXiv
2609.10559
T1 · 15 min ago
Category
cs.LG
T1 · 15 min ago

Abstract

To address the challenges of behavioral multimodality, limited semantic utilization, and long-term error accumulation in vessel trajectory prediction, this paper proposes M3-Former, a multimodal trajectory prediction framework enhanced by large language models (LLMs). The proposed framework incorporates vessel static attributes and navigational intent as semantic priors for long-term trajectory modeling. Specifically, a unified multimodal representation space is constructed, in which static semantic information is encoded by a pre-trained LLM and aligned with dynamic trajectory features through self-attention. To jointly capture global route planning and local motion variations, a dual-granularity Mixture-of-Experts (MoE) architecture is introduced, where sequence-level experts model global navigation trends and token-level experts refine fine-grained maneuvering behaviors. In addition, a Steering-Weighted Cross-Entropy loss is designed to alleviate the long-tail distribution of sparse turning samples and improve prediction accuracy in critical maneuvering scenarios. Experiments on a real-world Danish AIS dataset demonstrate that M\textsuperscript{3}-Former consistently outperforms state-of-the-art baselines across prediction horizons from 1 to 4 hours. In the 4-hour prediction task, the proposed method reduces Average Displacement Error (ADE) and Final Displacement Error (FDE) by 4.4\% and 5.1\%, respectively, compared with the strongest baseline. Qualitative and ablation analyses further verify that semantic fusion effectively reduces long-term trajectory drift, while the dual-granularity MoE improves robustness in complex waterways and route-branching scenarios. The proposed framework establishes a semantic-guided hierarchical prediction paradigm, in which high-level navigational intent and local motion dynamics are jointly modeled for robust long-term vessel trajectory forecasting.

Authors 2

Wenzhe Jin, Haina Tang

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 15 min agohigh

Arxiv announce type
cross

Source:arXiv (Atom API + RSS)T1observed 15 min agohigh

arXiv id
2609.10559

Source:arXiv (Atom API + RSS)T1observed 15 min agohigh

Categories
cs.LG, cs.CV

Source:arXiv (Atom API + RSS)T1observed 15 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 15 min agohigh

Primary category
cs.LG

Source:arXiv (Atom API + RSS)T1observed 15 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 15 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

15 min ago

Conflicts

None