Skip to content
AI Atlas
PaperActive

Analyzing Traditional and Neural Approaches to Multilingual Readability Assessment

arxiv.org/abs/2609.10792

quality89

Updated 2 h ago · first seen 11 Sept 2026

paper_01M294G4C5DZJ0DSG8ZNW5WWAR

Published
11 Sept 2026
T1 · 2 h ago
arXiv
2609.10792
T1 · 2 h ago
Category
cs.CL
T1 · 2 h ago

Abstract

Transformer-based models excel at Automatic Readability Assessment (ARA), yet feature-based models remain in active use because their predictions tie back to linguistic properties. This matters because readability labels are subjective and rater-dependent, so high accuracy on noisy ground truth may reflect surface patterns rather than the linguistic structure that defines difficulty. We test whether transformers internalize the same features as traditional models across Arabic, English, French, Hindi, and Russian using the ReadMe++ dataset. Shapley Additive Explanations (SHAP) identify the features driving traditional classifiers, which we then use as TCAV concept sets to probe multilingual XLM-R and language-specific encoders. Transformers recover surface-length, syntactic, and lexical-diversity signals, and reflect the ordinal CEFR structure of the traditional models. Alignment varies by model family, language, and layer, with language-specific encoders tracking traditional models more clearly than XLM-R. High linear separability does not always imply directional influence, limiting linear probing for count-based readability features.

Authors 2

Joshua Wong, Chris Tanner

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

arXiv id
2609.10792

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Categories
cs.CL

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Primary category
cs.CL

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

2 h ago

Conflicts

None