Skip to content
AI Atlas
PaperActive

Toward a Cross-Lingual Romanization Ecosystem for Sinitic Languages: A Paired Mandarin-Cantonese Case Study

arxiv.org/abs/2608.29170

quality89

Updated 6 h ago · first seen 11 Sept 2026

paper_01M294G67ECHMQVECNDCC805YE

Published
11 Sept 2026
T1 · 6 h ago
arXiv
2608.29170
T1 · 6 h ago
Category
cs.CL
T1 · 6 h ago

Abstract

This paper proposes the Sinitic Romanization Ecosystem, a cross-lingual Sinitic romanization design framework with supporting digital infrastructure and a community-driven open-source workflow. The design framework addresses the lack of systematic cross-lingual romanization alignment among Sinitic languages through four design principles: phonetic correspondence for representing similar sounds with similar romanized symbols, historical-phonological correspondence for aligning cognate romanization strings, one-phoneme-one-symbol, and basic Latin-letter use, with a balancing consideration recognizing trade-offs among these principles. For the main paired case study, we develop CantRomZJ1 and MandRomZJ1, Cantonese and Mandarin romanization schemes following the design framework, respectively. We also develop schemes for several other Sinitic languages, including Meixian Hakka, Shanghai Wu, and Nanjing Jianghuai Mandarin, following the same design framework. To bring the romanization schemes into practical use, we develop open-source infrastructure for structured romanization storage, conversion, parsing, dictionary construction, and input-method generation. Finally, we evaluate the design framework through speech-to-romanization experiments based on Meta's Massively Multilingual Speech (MMS) fine-tuning. Compared with the Pinyin+Jyutping baseline, our MandRomZJ1+CantRomZJ1 condition reduces Cantonese WER and CER by 7.80% and 10.61%, respectively. These results suggest that cross-lingual romanization alignment can improve transfer in low-resource Sinitic speech technology.

Authors 4

Zijie Zhang, Tan Lee, Yong Cao, Benyou Wang

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

Arxiv announce type
replace

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

arXiv id
2608.29170

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

Categories
cs.CL

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

Primary category
cs.CL

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 6 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

6 h ago

Conflicts

None