Skip to content
AI Atlas
PaperActive

BaltiVoice: A Speech Corpus and Fine-tuned Whisper ASR System for the Balti Language

arxiv.org/abs/2606.03504

Updated 52 min ago · first seen 11 Sept 2026

paper_01M294GQ47D8WHNB2WW19Y1Z24

Published
11 Sept 2026
T1 · 52 min ago
arXiv
2606.03504
T1 · 52 min ago
Category
cs.CL
T1 · 52 min ago

Abstract

-cross Abstract: We present BaltiVoice, a 16.8-hour read-speech corpus for Balti (ISO 639-3: bft), a Tibetic language spoken in Gilgit-Baltistan, Pakistan, with no prior publicly available ASR resources. The corpus contains 10,060 validated utterances in native Nastaliq script, derived from Mozilla Common Voice recordings. Fine-tuning OpenAI Whisper-small yields a Word Error Rate (WER) of 24.78% and a Character Error Rate (CER) of 8.30% after training for 5 epochs (3,000 steps) on the 538-utterance speaker-disjoint validation set, down from a zero-shot baseline of 159.19% WER and 152.52% CER. A Whisper-base fine-tuned on the same data achieves 44.54% WER and 15.61% CER, confirming that model capacity matters for this low-resource setting. The dataset, fine-tuned model, and a live transcription demo are publicly available on HuggingFace.

Authors 1

Muhammad Ali

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

Arxiv announce type
replace

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

arXiv id
2606.03504

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

Categories
cs.CL, cs.AI

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

Primary category
cs.CL

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 52 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

52 min ago

Conflicts

None