EMMI: Edge Multi-Modal Intelligence for Communication-Efficient MLLM Inference via Fused Representation Compression
Updated 35 min ago · first seen 11 Sept 2026
paper_01M294FNZ2EGRDJ67QARD9K24S
- Published
- 11 Sept 2026
- T1 · 35 min ago
- arXiv
- 2609.11058
- T1 · 35 min ago
- Category
- cs.LG
- T1 · 35 min ago
Abstract
Recent advances in multimodal large language mod- els (MLLMs) have opened new opportunities for edge intelligence by enabling reasoning across heterogeneous sensor modalities, such as vision, text, and telemetry data. However, deploying these capabilities on resource-constrained edge platforms remains challenging due to the substantial computational, memory, and communication demands of modern MLLMs. Rather than transmitting raw sensor observations or partitioning neural networks at intermediate layers, Edge Multi-Modal Intelligence (EMMI) communicates a compact representation between edge devices and server resources, enabling communication-efficient edge MLLM inference. To achieve this, EMMI performs modality-specific encoding, cross-modal representation fusion, and learned compression at the edge, transmitting only a compact latent representation to server-side resources for high-capacity MLLM reasoning. This representation-centric design reduces communication overhead, preserves local data privacy, and provides a fixed-size interface between heterogeneous edge devices and server-side MLLMs. Evaluation on a representative multimodal benchmark demonstrates that EMMI can reduce the communication payload by 32x while maintaining comparable downstream accuracy, resulting in up to a 3.4x reduction in estimated end-to-end inference latency under bandwidth-constrained edge conditions.
Authors 2
Motahare Mounesan, Irfan Khan
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 35 min agohigh
- Arxiv announce type
- new
Source:arXiv (Atom API + RSS)T1observed 35 min agohigh
- arXiv id
- 2609.11058
Source:arXiv (Atom API + RSS)T1observed 35 min agohigh
- Categories
- cs.LG, cs.DC
Source:arXiv (Atom API + RSS)T1observed 35 min agohigh
Source:arXiv (Atom API + RSS)T1observed 35 min agohigh
- Primary category
- cs.LG
Source:arXiv (Atom API + RSS)T1observed 35 min agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 35 min agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
35 min ago
Conflicts
None
No models linked to this paper yet.
- Authors
- Motahare Mounesan, Irfan Khan
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · PDF
PDFpdf_url1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://arxiv.org/pdf/2609.11058 | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
- New paperPaperEMMI: Edge Multi-Modal Intelligence for Communication-Efficient MLLM Inference via Fused Representation Compression
New paper: EMMI: Edge Multi-Modal Intelligence for Communication-Efficient MLLM Inference via Fused Representation Compression
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.LG | feed | T1· Official | 35 min ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.