Updated 44 min ago · first seen 11 Sept 2026
paper_01M294FPB1VK57J4F3Q9S0P4T7
- Published
- 11 Sept 2026
- T1 · 44 min ago
- arXiv
- 2609.11716
- T1 · 44 min ago
- Category
- cs.LG
- T1 · 44 min ago
Abstract
Post-training quantization compresses large language models (LLMs) by storing their weights at reduced precision, and each quantized weight introduces an error into the hidden states. Naively, these errors should accumulate with depth and corrupt next-token prediction; randomly initialized models accumulate these discrepancies rapidly, whereas quantized pretrained models accumulate much less hidden-state error and largely maintain downstream task performance, even though they were never trained with quantization noise. This raises the question we address: why does post-training quantization work? Comparing full-precision and quantized forward passes, we identify two mechanisms that characterize pretrained quantization robustness. First, the error a layer newly introduces tends to oppose the error it inherits from the layer's input. The two cancel partially such that the discrepancy between full-precision and quantized passes grows slowly. This counteracting residual interaction develops during pretraining. Our quantitative analysis identifies it as a major factor slowing hidden-error growth. Second, LM-head geometry preferentially preserves the scores and probabilities of high-ranked tokens, which typically represent the model's most confident predictions. Together, these mechanisms explain why quantization error that passes through numerous layers can still produce only small output changes, and we verify the findings across models and quantization settings.
Authors 4
Yuxiang Chen, Michael Beyer, Jun Zhu, Jianfei Chen
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 44 min agohigh
- Arxiv announce type
- cross
Source:arXiv (Atom API + RSS)T1observed 44 min agohigh
- arXiv id
- 2609.11716
Source:arXiv (Atom API + RSS)T1observed 44 min agohigh
- Categories
- cs.LG, cs.CL
Source:arXiv (Atom API + RSS)T1observed 44 min agohigh
Source:arXiv (Atom API + RSS)T1observed 44 min agohigh
- Primary category
- cs.LG
Source:arXiv (Atom API + RSS)T1observed 44 min agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 44 min agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
44 min ago
Conflicts
None
No models linked to this paper yet.
- Authors
- Yuxiang Chen, Michael Beyer, Jun Zhu
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · Official page
Official pageofficial_url1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://arxiv.org/abs/2609.11716 | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
Why Does Post-Training Quantization Work?: arxiv announce type changed from new to cross
Arxiv announce typenew→crossarxivNew paper: Why Does Post-Training Quantization Work?
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.CL | feed | T1· Official | 44 min ago | 1 |
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.LG | feed | T1· Official | 44 min ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.