Skip to content
AI Atlas
PaperActive

Component-Aware Differential Privacy for Federated Multilingual Speech-LLMs

arxiv.org/abs/2609.11762

quality89

Updated 4 h ago · first seen 11 Sept 2026

paper_01M294G4TVKQDAMSX5ANXV84FZ

Published
11 Sept 2026
T1 · 4 h ago
arXiv
2609.11762
T1 · 4 h ago
Category
cs.CL
T1 · 4 h ago

Abstract

Per-layer differential privacy (DP) clipping improves gradient fidelity in federated learning by allocating per-matrix clipping budgets proportional to parameter count. We show that this recipe breaks for speech large language models (speech-LLMs), when the acoustic encoder and the language decoder differ by an order of magnitude in update norm. Single-pool per-layer methods suffer \emph{cross-component budget collapse}, dragging word error rate (WER) far from flat global clipping or collapsing training entirely. When the norm imbalance is milder, adaptive single-pool methods partially recover, confirming that collapse severity scales with the inter-component norm ratio. We empirically diagnose the root cause across six per-layer methods and three speech-LLM architectures. We then propose \emph{$\alpha$-split}, a two-pool allocation that normalises encoder and LLM parameters into independent pools, and show that joint $\ell_2$ sensitivity and the original $(\varepsilon,\delta)$-DP guarantee are unchanged. At architecture-calibrated $\alpha$, our method recovers WER utility compared to flat DP, while granting the encoder $4.47{\times}$ tighter per-component noise protection against speaker voice-based gradient-inversion attacks at only $+2.6\%$ LLM noise overhead.

Authors 3

Jordi Luque, Fernando L\'opez, Aleix Sant

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

arXiv id
2609.11762

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

Categories
cs.CL

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

Primary category
cs.CL

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

4 h ago

Conflicts

None