Skip to content
AI Atlas
PaperActive

What Language is This? Ask Your Tokenizer

arxiv.org/abs/2602.17655

quality89

Updated 7 h ago · first seen 11 Sept 2026

paper_01M294G5NCMVXS0V3H4VKFBBHQ

Published
11 Sept 2026
T1 · 7 h ago
arXiv
2602.17655
T1 · 7 h ago
Category
cs.CL
T1 · 7 h ago

Abstract

Language Identification (LID) is an important component of many multilingual natural language processing pipelines, where it facilitates corpus curation, training data analysis, and cross-lingual evaluation of large language models. Despite near-perfect performance on high-resource languages, existing systems remain brittle in low-resource and closely related language settings. We introduce UniLID, a simple and efficient LID method based on the UnigramLM tokenization algorithm. In short, to predict a string's language label, we simply ask: under which language's unigram distribution is this string most likely? Our formulation is data- and compute-efficient, supports incremental addition of new languages without retraining existing models, and can naturally be integrated into existing language model tokenization pipelines. Empirical evaluations against widely used baselines, including fasttext, GlotLID-M, and CLD3, show that UniLID achieves competitive performance on standard benchmarks, reaches 69% accuracy with five labeled samples per language and 89% with 25, and delivers large gains on fine-grained dialect identification.

Authors 4

Clara Meister, Ahmetcan Yavuz, Pietro Lesci, Tiago Pimentel

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 7 h agohigh

Arxiv announce type
replace

Source:arXiv (Atom API + RSS)T1observed 7 h agohigh

arXiv id
2602.17655

Source:arXiv (Atom API + RSS)T1observed 7 h agohigh

Categories
cs.CL

Source:arXiv (Atom API + RSS)T1observed 7 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 7 h agohigh

Primary category
cs.CL

Source:arXiv (Atom API + RSS)T1observed 7 h agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 7 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

7 h ago

Conflicts

None