Skip to content
AI Atlas
PaperActive

Revisiting the Shape Convention of Transformer Language Models

arxiv.org/abs/2602.06471

Updated 51 min ago · first seen 11 Sept 2026

paper_01M294GPTDJANW8HWP0RXCEKMW

Published
11 Sept 2026
T1 · 51 min ago
arXiv
2602.06471
T1 · 51 min ago
Category
cs.CL
T1 · 51 min ago

Abstract

-cross Abstract: The architectural shape of dense Transformers has remained remarkably stable: narrow-wide-narrow feed-forward networks (FFNs) consume most non-embedding parameters. Motivated by theoretical and empirical evidences that residual wide-narrow-wide (hourglass) MLPs remain expressive despite bottlenecks, we revisit whether this architectural convention is necessary for dense language models. We study Hourglass Transformers, which replace the conventional FFN with residual stacks of hourglass sub-MLPs and use hourglass attention to decouple residual-stream width from attention width. This exposes a practical depth-width trade-off: compressing the FFN intermediate dimension allows wider hidden states and fewer layers at matched parameter budgets. Across model scales from 113M to 8B parameters, Hourglass Transformers achieve language-modeling and downstream performance comparable to conventional Transformers, while improving training compute efficiency by $8.7\%$ at matched average downstream accuracy across the 906M, 3B, and 8B scales. After long-context extension, the 8B Hourglass model also outperforms its matched conventional baseline across 4k-64k context lengths. At 64k context, the reduced attention layer count lowers both computation and KV-cache requirements, yielding up to $1.93\times$ faster token decoding and $50\%$ lower KV-cache memory at the 1B scale. These results identify hourglass structures as a practical architecture-efficiency alternative for compute- and latency-conscious Transformer design.

Authors 5

Feng-Ting Liao, Guan-Ting Yi, Tzu-Quan Lin, Meng-Hsi Chen, Da-shan Shiu

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Arxiv announce type
replace

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

arXiv id
2602.06471

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Categories
cs.CL, cs.AI, cs.LG

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Primary category
cs.CL

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

51 min ago

Conflicts

None