Skip to content
AI Atlas
PaperActive

LILA: Calibration-Free Structured Pruning of Large Language Models via Latent Spectral Geometry

arxiv.org/abs/2609.11163

Updated 15 min ago · first seen 11 Sept 2026

paper_01M294FP1VVTCW15ZGNK0HCF54

Published
11 Sept 2026
T1 · 15 min ago
arXiv
2609.11163
T1 · 15 min ago
Category
cs.LG
T1 · 15 min ago

Abstract

Structured pruning of large language models (LLMs) offers hardware-efficient compression, yet existing methods require calibration data, gradient computation, or large auxiliary policy networks at pruning time. LILA (\emph{Latent-Informed Layer Analysis}) scores neuron importance via the Kolmogorov--Smirnov (KS) distance between empirical singular value distributions of the full and neuron-ablated feed-forward network (FFN) weight matrix, providing a closed-form spectral rule requiring no training, calibration data, or auxiliary network. Without any fine-tuning, LILA surpasses PruneNet (45M-parameter RL policy) by 1.57~pp in zero-shot accuracy on LLaMA-2-7B at 25\% sparsity, and outperforms WikiText-2-calibrated SliceGPT by up to 6.0~pp across all sparsity levels, while preserving the original architecture. After one epoch of LoRA recovery fine-tuning, LILA achieves highly competitive performance, matching the heavily calibrated SliceGPT baseline to within a 0.48~pp margin across LLaMA-2-7B and Phi-2, despite using zero calibration data. A Neural Tangent Kernel analysis confirms a 22$\times$ reduction in functional distortion versus random pruning, providing theoretical grounding for the spectral importance criterion. Finally, extending LILA to dynamically allocate sparsity budgets via KS-scores yields state-of-the-art generative preservation at moderate compression, while uncovering fundamental single-layer architectural bottlenecks at higher compression regimes.

Authors 6

Sankar Behera, Dhruv Singh, Anshika Agnihotri, Raj Kumar Choudhary, Satyadev Ahlawat, Yamuna Prasad

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 15 min agohigh

Arxiv announce type
cross

Source:arXiv (Atom API + RSS)T1observed 15 min agohigh

arXiv id
2609.11163

Source:arXiv (Atom API + RSS)T1observed 15 min agohigh

Categories
cs.LG, cs.CL

Source:arXiv (Atom API + RSS)T1observed 15 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 15 min agohigh

Primary category
cs.LG

Source:arXiv (Atom API + RSS)T1observed 15 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 15 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

15 min ago

Conflicts

None