Updated 8 h ago · first seen 11 Sept 2026
paper_01M294G5NCMVXS0V3H4VKFBBHQ
- Published
- 11 Sept 2026
- T1 · 8 h ago
- arXiv
- 2602.17655
- T1 · 8 h ago
- Category
- cs.CL
- T1 · 8 h ago
Abstract
Language Identification (LID) is an important component of many multilingual natural language processing pipelines, where it facilitates corpus curation, training data analysis, and cross-lingual evaluation of large language models. Despite near-perfect performance on high-resource languages, existing systems remain brittle in low-resource and closely related language settings. We introduce UniLID, a simple and efficient LID method based on the UnigramLM tokenization algorithm. In short, to predict a string's language label, we simply ask: under which language's unigram distribution is this string most likely? Our formulation is data- and compute-efficient, supports incremental addition of new languages without retraining existing models, and can naturally be integrated into existing language model tokenization pipelines. Empirical evaluations against widely used baselines, including fasttext, GlotLID-M, and CLD3, show that UniLID achieves competitive performance on standard benchmarks, reaches 69% accuracy with five labeled samples per language and 89% with 25, and delivers large gains on fine-grained dialect identification.
Authors 4
Clara Meister, Ahmetcan Yavuz, Pietro Lesci, Tiago Pimentel
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 8 h agohigh
- Arxiv announce type
- replace
Source:arXiv (Atom API + RSS)T1observed 8 h agohigh
- arXiv id
- 2602.17655
Source:arXiv (Atom API + RSS)T1observed 8 h agohigh
- Categories
- cs.CL
Source:arXiv (Atom API + RSS)T1observed 8 h agohigh
Source:arXiv (Atom API + RSS)T1observed 8 h agohigh
- Primary category
- cs.CL
Source:arXiv (Atom API + RSS)T1observed 8 h agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 8 h agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
8 h ago
Conflicts
None
No models linked to this paper yet.
- Authors
- Clara Meister, Ahmetcan Yavuz, Pietro Lesci
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · Authors
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
New paper: What Language is This? Ask Your Tokenizer
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.CL | feed | T1· Official | 1 h ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.