Indic ASR Optimization & Consensus Evaluation
Speech · Evaluation research · 2025Problem
Off-the-shelf Whisper underperforms on Hindi: Devanagari orthography is inconsistent across corpora, and single-model transcripts fail unpredictably — so the evaluation itself is unreliable before the model is.
Approach
Fine-tuned Whisper-small with SpecAugment, label smoothing and cosine LR scheduling under a 5-beam search pipeline; built a Devanagari normalization engine with phonetic reverse-transliteration; then fused five ASR systems through word-level confusion networks with majority voting, so the final transcript is a measured consensus rather than one model's guess.
Results
>48% absolute WER reduction
FLEURS Hindi benchmark vs. baseline Whisper-small
177,000 unique terms
normalized via Devanagari engine + reverse transliteration
5-model consensus
word-level confusion networks, majority voting