Skip to content
AI Atlas
PaperActive

Why Does Post-Training Quantization Work?

arxiv.org/abs/2609.11716

Updated 19 min ago · first seen 11 Sept 2026

paper_01M294FPB1VK57J4F3Q9S0P4T7

Published
11 Sept 2026
T1 · 20 min ago
arXiv
2609.11716
T1 · 20 min ago
Category
cs.LG
T1 · 20 min ago

Abstract

Post-training quantization compresses large language models (LLMs) by storing their weights at reduced precision, and each quantized weight introduces an error into the hidden states. Naively, these errors should accumulate with depth and corrupt next-token prediction; randomly initialized models accumulate these discrepancies rapidly, whereas quantized pretrained models accumulate much less hidden-state error and largely maintain downstream task performance, even though they were never trained with quantization noise. This raises the question we address: why does post-training quantization work? Comparing full-precision and quantized forward passes, we identify two mechanisms that characterize pretrained quantization robustness. First, the error a layer newly introduces tends to oppose the error it inherits from the layer's input. The two cancel partially such that the discrepancy between full-precision and quantized passes grows slowly. This counteracting residual interaction develops during pretraining. Our quantitative analysis identifies it as a major factor slowing hidden-error growth. Second, LM-head geometry preferentially preserves the scores and probabilities of high-ranked tokens, which typically represent the model's most confident predictions. Together, these mechanisms explain why quantization error that passes through numerous layers can still produce only small output changes, and we verify the findings across models and quantization settings.

Authors 4

Yuxiang Chen, Michael Beyer, Jun Zhu, Jianfei Chen

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 20 min agohigh

Arxiv announce type
cross

Source:arXiv (Atom API + RSS)T1observed 19 min agohigh

arXiv id
2609.11716

Source:arXiv (Atom API + RSS)T1observed 20 min agohigh

Categories
cs.LG, cs.CL

Source:arXiv (Atom API + RSS)T1observed 20 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 20 min agohigh

Primary category
cs.LG

Source:arXiv (Atom API + RSS)T1observed 20 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 20 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

19 min ago

Conflicts

None