Phase-Decoupled, Model-Calibrated Power Control for Disaggregated LLM Serving
Updated 35 min ago · first seen 11 Sept 2026
paper_01M294FP0X10TFCM54HSF8W39R
- Published
- 11 Sept 2026
- T1 · 35 min ago
- arXiv
- 2609.11133
- T1 · 35 min ago
- Category
- cs.LG
- T1 · 35 min ago
Abstract
Datacenter GPU power is the binding constraint on LLM serving capacity, and production serving has shifted to prefill/decode (PD) disaggregation. Deploying NVIDIA's Max-Q inference profile on a disaggregated B200 system, we found its realized gain modest (+8.6% tokens/J), model-dependent, and carrying a mean end-to-end latency cost (+5.2%) that throughput-only evaluation does not surface; the profile also applies one setting to prefill and decode GPUs that operate in opposite hardware regimes. We hypothesize that the optimal power setting is a property of the deployed (model, quantization, engine, hardware) combination rather than of the GPU class, that each lane warrants its own profile, and that converting SLO headroom into energy safely requires latency-gated calibration under a runtime SLO guard rather than a fixed recipe. We present a phase-decoupled, model-calibrated controller: the prefill lane runs under an SM-clock window whose floor is a latency guarantee by construction, and the decode lane under a power cap placed by automatic calibration just above a measured throughput/latency cliff. Because a disaggregated decode lane draws flat, memory-bound power, the cap binds continuously, the reactive-overshoot weakness that led POLCA to reject capping is absent, and the GPU's own power manager retains throughput under the cap. On an 8x B200 node serving Qwen3-Coder-480B (FP8) under agentic load, our balanced mode delivers +20.4% tokens/J at +3.5% mean e2e versus +8.6% at +5.2% for Max-Q, a Pareto improvement on both axes. On Qwen3-235B-A22B (NVFP4) every operating mode meets the ITL-p99 SLO in every repetition; both vendor profiles miss it. A decode-actuator A/B shows the calibrated cap beats static clock locks, and a three-day sustained run saves 32.3% of a lane pair's electricity. Both models are MoE; a dense model recovers roughly 5x less, so we scope our claims to MoE serving.
Authors 6
Jae Gon Kim, Donghoon Yoo, Hanyul Ryu, Sungho Ha, Juyeon Lee, Soojung Ryu
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 35 min agohigh
- Arxiv announce type
- new
Source:arXiv (Atom API + RSS)T1observed 35 min agohigh
- arXiv id
- 2609.11133
Source:arXiv (Atom API + RSS)T1observed 35 min agohigh
- Categories
- cs.LG, cs.DC
Source:arXiv (Atom API + RSS)T1observed 35 min agohigh
Source:arXiv (Atom API + RSS)T1observed 35 min agohigh
- Primary category
- cs.LG
Source:arXiv (Atom API + RSS)T1observed 35 min agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 35 min agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
35 min ago
Conflicts
None
No models linked to this paper yet.
- Authors
- Jae Gon Kim, Donghoon Yoo, Hanyul Ryu
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · Primary category
Primary categoryprimary_category1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| cs.LG | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
New paper: Phase-Decoupled, Model-Calibrated Power Control for Disaggregated LLM Serving
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.LG | feed | T1· Official | 35 min ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.