Skip to content
AI Atlas
PaperActive

FluxMoE: Decoupling Expert Residency for High-Performance MoE Serving

arxiv.org/abs/2604.02715

quality89

Updated 4 h ago · first seen 11 Sept 2026

paper_01M294FS4K6918SWM2QWRA9JRY

Published
11 Sept 2026
T1 · 4 h ago
arXiv
2604.02715
T1 · 4 h ago
Category
cs.LG
T1 · 4 h ago

Abstract

Mixture-of-Experts (MoE) models have become mainstream for scaling language models to hundreds of billions of expert parameters. Despite sparse expert activation, existing inference engines keep all experts GPU-resident, crowding out the key-value cache in large-batch, long-output offline workloads. We present FluxMoE, which decouples experts from physical GPU residency and adapts their footprint to available memory through a new \emph{expert paging} abstraction. FluxMoE combines PagedTensor for transparent remapping, a bandwidth-balanced hierarchy spanning losslessly compressed GPU memory and host DRAM, and a budget-aware residency planner. Unlike CPU-GPU co-inference and whole-layer offloading, FluxMoE streams weights on demand while keeping expert computation on GPUs. We implement FluxMoE atop vLLM and evaluate it on three MoE models. For GLM-4.5 on 8$\times$H20 GPUs, FluxMoE delivers up to 7.2$\times$ vLLM's throughput and 79.0\% lower average Time-Per-Output-Token (TPOT), without measurable model-quality loss using lossless compression. For Mixtral-8$\times$7B-Instruct on 2$\times$L40S GPUs, where weight-resident vLLM cannot fit, FluxMoE delivers 4.3$\times$ KTransformers's throughput and 29.1\% lower average TPOT.

Authors 7

Qingxiu Liu, Yongchao He, Runhan Jiang, Zion Wang, Bohan Zhao, Mi Zhang, Patrick P. C. Lee

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

Arxiv announce type
replace

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

arXiv id
2604.02715

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

Categories
cs.LG

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

Primary category
cs.LG

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

4 h ago

Conflicts

None