Skip to content
AI Atlas
PaperActive

X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation

arxiv.org/abs/2609.11412

quality59

Updated 2 h ago · first seen 11 Sept 2026

paper_01M294WYDTQJF309BA9AX08EJW

Published
10 Sept 2026
T2 · 6 h ago
arXiv
2609.11412
T2 · 6 h ago

As of

Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.

Claim history

12 claims · 11 properties

Official pageofficial_url1

Claim history for Official page
ValueValid from → toStatusSourceConfidenceExtractor
https://arxiv.org/abs/2609.11412currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Abstractabstract1

Claim history for Abstract
ValueValid from → toStatusSourceConfidenceExtractor
Reducing audio-encoder depth lowers the inference cost of speech large language models, but removing complete blocks perturbs the embeddings consumed by the decoder and can cause deletion and premature end-of-sequence errors. We introduce X-AuT, a progressive framework that selects layer combinations through short behavioral probes and restores the pruned model through representation alignment, cross-scale distillation, scheduled student-policy supervision, and LoRA finetuning. The language-model backbone remains frozen, while attention LoRA adapters and the tied output embedding adapt during distillation. Training uses the highest-agreement tier from a transcript-consistency pipeline, followed by source reweighting during finetuning. On ten public Chinese--English benchmarks, compressing Qwen3-ASR-0.6B from 18 to 16 audio-encoder layers reduces macro-average error from 5.61% to 5.27%. The 14-layer model reaches 5.75% with 20.7% fewer audio-tower parameters. Under the matched recipe, the 1.7B teacher yields 5.55% mean error, compared with 8.45% for self-distillation, and progressive 18rightarrow14 pruning outperforms direct pruning (5.75% vs. 6.73%). These single-run results establish two practical operating points and show that the accuracy effects vary across benchmarks. Project website: https://xpeng-ai.github.io/x-autcurrentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

arXiv idarxiv_id1

Claim history for arXiv id
ValueValid from → toStatusSourceConfidenceExtractor
2609.11412currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Authorsauthors1

Claim history for Authors
ValueValid from → toStatusSourceConfidenceExtractor
Haojun Zhang, Yi Zou, Min Chen, Qize Yu, Lianrui Fan, Xini Ding, Hao Li, Shuchang Zhou, Xianming Liu, Shiyu HuangcurrentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Github repogithub_repo1

Claim history for Github repo
ValueValid from → toStatusSourceConfidenceExtractor
XPENG-AI/X-AuTcurrentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Hf paper urlhf_paper_url1

Claim history for Hf paper url
ValueValid from → toStatusSourceConfidenceExtractor
https://huggingface.co/papers/2609.11412currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Github starsmetric.github_stars1

Claim history for Github stars
ValueValid from → toStatusSourceConfidenceExtractor
7currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Hf commentsmetric.hf_comments2

Claim history for Hf comments
ValueValid from → toStatusSourceConfidenceExtractor
2currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic
1supersededHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Upvotesmetric.upvotes1

Claim history for Upvotes
ValueValid from → toStatusSourceConfidenceExtractor
13currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

PDFpdf_url1

Claim history for PDF
ValueValid from → toStatusSourceConfidenceExtractor
https://arxiv.org/pdf/2609.11412currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Publishedpublished_at1

Claim history for Published
ValueValid from → toStatusSourceConfidenceExtractor
10 Sept 2026currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →