Work / Fig. 03 · Speech · Evaluation research · 2025

Indic ASR Optimization & Consensus Evaluation

>48% absolute WER reduction — FLEURS Hindi benchmark vs. baseline Whisper-small.

Fig. 03

Indic ASR Optimization & Consensus Evaluation

Speech · Evaluation research · 2025

Problem

Off-the-shelf Whisper underperforms on Hindi: Devanagari orthography is inconsistent across corpora, and single-model transcripts fail unpredictably — so the evaluation itself is unreliable before the model is.

Approach

Fine-tuned Whisper-small with SpecAugment, label smoothing and cosine LR scheduling under a 5-beam search pipeline; built a Devanagari normalization engine with phonetic reverse-transliteration; then fused five ASR systems through word-level confusion networks with majority voting, so the final transcript is a measured consensus rather than one model's guess.

Results

  • >48% absolute WER reduction

    FLEURS Hindi benchmark vs. baseline Whisper-small

  • 177,000 unique terms

    normalized via Devanagari engine + reverse transliteration

  • 5-model consensus

    word-level confusion networks, majority voting