SG-Blend: Learning an Interpolation Between Improved Swish and GELU for Robust Neural Representations
Updated 7 h ago · first seen 11 Sept 2026
paper_01M294FRRKBKR6YAM8BA3EPS80
- Published
- 11 Sept 2026
- T1 · 7 h ago
- arXiv
- 2505.23942
- T1 · 7 h ago
- Category
- cs.LG
- T1 · 7 h ago
Abstract
Prevailing activation functions such as Swish and GELU tend toward domain-specific optima, Swish was discovered via neural architecture search on vision benchmarks, while GELU dominates transformer-based language models, and neither offers any mechanism to adapt its gating shape to individual layers. This rigidity is especially consequential in transformer FFN blocks, where LayerNorm, unlike BatchNorm, does not suppress the gradient pathologies that activation choice induces across depth. We propose SG-Blend, a per layer adaptive activation that combines SSwish, a bias-corrected, parametric Swish variant we also introduce, with learnable sharpness \b{eta} and zero-centering bias {\gamma}, with GELU through a per-layer blend coefficient {\alpha}, letting each layer locate its own optimum along the SSwishGELU continuum at a cost of only three additional scalars per FFN block, with \b{eta} initialized to 1.0 and learned freely via backpropagation. On BERT-style IMDB classification (5 seeds), it matches peak accuracy (81.31%) while reducing seed-to-seed variance by 42% relative to GELU. Furthermore, it generalizes to autoregressive pretraining, achieving the lowest validation perplexity (49.10) on WikiText103 among all baselines. Crucially, ablations confirm the interpolation structure itself drives these gains, delivering reliable, top-tier performance. Beyond natural language processing, we demonstrate that SG-Blend generalizes robustly to a wider variety of tasks, extending its efficacy to computer vision and other diverse domains.
Authors 4
Gaurav Sarkar, Syed Affan Daimi, Jay Gala, Subarna Tripathi
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
- Arxiv announce type
- replace
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
- arXiv id
- 2505.23942
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
- Categories
- cs.LG, cs.AI
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
- Primary category
- cs.LG
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 7 h agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
7 h ago
Conflicts
None
No models linked to this paper yet.
- Authors
- Gaurav Sarkar, Syed Affan Daimi, Jay Gala
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · Primary category
Primary categoryprimary_category1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| cs.LG | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
- New paperPaperSG-Blend: Learning an Interpolation Between Improved Swish and GELU for Robust Neural Representations
New paper: SG-Blend: Learning an Interpolation Between Improved Swish and GELU for Robust Neural Representations
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.LG | feed | T1· Official | 5 h ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.