DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation
Updated 1 h ago · first seen 11 Sept 2026
paper_01M294G5ZJTF236B0AZ2BDGMT4
- Published
- 12 Sept 2026
- T1 · 1 h ago
- arXiv
- 2607.07669
- T1 · 9 h ago
- Category
- cs.CL
- T1 · 9 h ago
Abstract
Large language models increasingly understand dialectal English, yet still produce only standard, US-leaning English, leaving dialectal generation, the harder half of the problem, largely unaddressed. We introduce DiaLLM, which continually pretrains three open-weight language model families on the International Corpus of English and applies implicit and explicit post-training paradigms, each combined with three model alignment strategies, giving the first controlled comparison of these components across Australian, Indian, and Northern British English. Our results reveal a robustness-generation gap: benchmarks are shaped by continual pretraining and SFT, while alignment visibly reshapes generation in ways benchmarks do not capture. Explicit variety-targeted adaptation produces output reliably recognised as dialectal and judged more dialectal than broad alignment, yet where human judgement was directly assessed, the method that most aggressively optimises the dialectal reward is not the one judged most dialectal. Independent linguistic analysis corroborates this reward-quality gap, most clearly on two of the three families. No single alignment method dominates, and closing the gap will require richer reward designs and continued investment in dialectal resources. We release all code, checkpoints, and preference datasets.
Authors 6
Jordan Painter, Dipankar Srirag, Adarsh Kappiyath, Diptesh Kanojia, Aditya Joshi, Lu Yin
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 9 h agohigh
- Arxiv announce type
- replace
Source:arXiv (Atom API + RSS)T1observed 9 h agohigh
- arXiv id
- 2607.07669
Source:arXiv (Atom API + RSS)T1observed 9 h agohigh
- Categories
- cs.CL, cs.AI
Source:arXiv (Atom API + RSS)T1observed 9 h agohigh
Source:arXiv (Atom API + RSS)T1observed 9 h agohigh
- Primary category
- cs.CL
Source:arXiv (Atom API + RSS)T1observed 9 h agohigh
- Published
- 12 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 1 h agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
1 h ago
Conflicts
None
No models linked to this paper yet.
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · PDF
PDFpdf_url1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://arxiv.org/pdf/2607.07669 | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
- Property changedPaperDiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation
DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation: published at changed from 2026-09-11T04:00:00+00:00 to 2026-09-12T04:00:00+00:00
Published11 Sept 2026→12 Sept 2026arxiv - New paperPaperDiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation
New paper: DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.AI | feed | T1· Official | 1 h ago | 2 |
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.CL | feed | T1· Official | 1 h ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.