Skip to content
AI Atlas
PaperActive

Looped GPT-BERT: Trading Parameters for Computation in Small Language Modeling

arxiv.org/abs/2609.09691

Updated 51 min ago · first seen 11 Sept 2026

paper_01M294GMJAQ7A7YPD2WRK4ZENN

Published
11 Sept 2026
T1 · 51 min ago
arXiv
2609.09691
T1 · 51 min ago
Category
cs.CL
T1 · 51 min ago

Abstract

When training data are limited, increasing parameter count is not the only way to improve language-model performance. A small parameter set, when repeatedly applied, can also deliver comparable performance. We study Looped GPT-BERT in the BabyLM 2026 Strict-small setting, combining GPT-BERT's masked next-token and causal language-modeling objectives with depth-wise parameter sharing. We train on a preprocessed 7.48M-word English corpus and compare objective ratios, non-looped and looped architectures, and loop counts. Our final $4\times12$ model uses four physical layers for twelve recurrent traversals and contains 12.18M parameters. The BabyLM 2026 leaderboard reports an Overall Average of 35.42 and an NLP Average of 48.48. Compared with public BabyLM 10M Strict-small GPT-2 and GPT-BERT baselines, it achieves comparable performance on selected linguistic and downstream metrics, including BLiMP and GLUE, with fewer parameters. The loop ablations show that additional recurrent computation can improve training and preserve strong performance on selected linguistic tasks, whereas poorer performance on other tasks may reveal an inherent limitation of the looped design: using only a few physical layers restricts the model's representational space.

Authors 5

Tingshuo Fan, Hongtao Mu, Tianyu Zhou, Hansen Liu, Tao Ji

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Arxiv announce type
cross

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

arXiv id
2609.09691

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Categories
cs.CL, cs.AI

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Primary category
cs.CL

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

51 min ago

Conflicts

None