Skip to content
AI Atlas
PaperActive

Combining Synthetic and Real Data for Low-Resource Historical OCR: A Manchu Case Study

arxiv.org/abs/2609.11495

Updated 30 min ago · first seen 11 Sept 2026

paper_01M294FP79YZR5P7ZJGCQVSA94

Published
11 Sept 2026
T1 · 31 min ago
arXiv
2609.11495
T1 · 31 min ago
Category
cs.LG
T1 · 31 min ago

Abstract

Manchu, now critically endangered, was one of the principal languages of the Qing empire (1636-1912), and its extensive archival record is increasingly digitized but remains difficult to search and analyze at scale. Previous work showed that vision-language models (VLMs) trained only on synthetic Manchu word images can reach 87.4% word accuracy on real Qing manuscripts and prints, leaving a substantial synthetic-to-real gap. This study examines how synthetic and real historical training data should be combined for low-resource OCR. Using 60,000 synthetic and 20,306 real historical word images, we evaluate three pretrained VLMs and a compact convolutional recurrent neural network (CRNN) under four regimes: synthetic-only, real-only, joint synthetic-real, and sequential synthetic-to-real training, following a common checkpoint-selection and archival evaluation protocol. Introducing real training images raises the leading configurations to between 95.09% and 96.28% word accuracy, while no synthetic-only configuration exceeds 87.92%. Synthetic supplementation substantially improves all three VLMs, whereas its marginal effect for the CRNN is sensitive to the training objective. Joint and sequential training yield broadly similar archival accuracy under the tested practical pipelines. A compact CRNN also reaches the leading performance range once real images are available, showing that model scale alone does not determine recognition accuracy. Finally, complementary errors among strong recognizers allow voting to raise accuracy to 98.27% without additional training, while an eighteenth-century Manchu dictionary provides a principled rule for adjudicating disagreements.

Authors 2

Yan Hon Michael Chung, Hanlin Wang

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 31 min agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 31 min agohigh

arXiv id
2609.11495

Source:arXiv (Atom API + RSS)T1observed 31 min agohigh

Categories
cs.LG

Source:arXiv (Atom API + RSS)T1observed 31 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 31 min agohigh

Primary category
cs.LG

Source:arXiv (Atom API + RSS)T1observed 31 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 31 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

31 min ago

Conflicts

None