Evaluating Indic speech recognition honestly
Why word error rate misleads on Indian-language audio, what to measure instead, and how to design the human gate that stands between a model and an official record.
A vendor demonstration of speech recognition is almost always a clean recording of one speaker in a quiet room. The system we run transcribes a legislative chamber: overlapping speech, a public gallery, microphones at varying distances, formal Telugu, colloquial Telugu, English clauses inside Telugu sentences, procedural formulae, and names that appear nowhere in any general lexicon. The gap between those two situations is where every honest evaluation lives.
Word error rate is the wrong headline
WER divides insertions, deletions and substitutions by reference length and treats every word as equal. On Indian-language audio that assumption breaks in three specific ways, and each one moves the number in a direction that flatters the system.
| Measure | What it tells you | Why we keep it |
|---|---|---|
| CER (character) | Degradation independent of tokenisation | The only figure comparable across scripts and languages |
| Entity error rate | Accuracy on names, numbers, dates, places | The measure that predicts whether the record is usable |
| Diarisation error rate | Whether the right speaker is attributed | In a chamber, attributing a sentence to the wrong member is worse than mis-transcribing it |
| Semantic error rate | Whether meaning survived, judged by a reviewer on a sample | Catches the fluent, plausible, wrong output that character measures miss entirely |
| Review minutes per audio hour | The real operating cost | The number the client feels; it decides whether the system is adopted |
Measure against the manual baseline, not against a leaderboard
Every institution already has a process. Before touching a model, measure it: how long does the existing process take, how many errors does it contain, where do they occur, and what happens when one is found. That baseline is the only thing the client cares about, and it is usually not written down anywhere.
- Sample the existing output and score it with the same rubric you will apply to the model. Human transcripts contain errors too, and knowing the rate changes what "good enough" means.
- Time the current end-to-end cycle including review and sign-off, not just the transcription step. Systems are adopted or rejected on cycle time.
- Record where corrections happen today. Those are exactly the places the interface must make easy, and they are rarely where an engineer would guess.
Designing the human gate
Nothing from a model enters an official record unreviewed. That principle is easy to state and mostly implemented badly: a reviewer facing a wall of plausible text approves it, and the gate becomes a rubber stamp that adds latency without adding assurance. A gate works only when it directs attention.
The domain lexicon is most of the win
On institutional audio, the highest-return work is not model selection. It is assembling the vocabulary the institution actually uses — member names with their spellings, constituencies, committee names, procedural phrases, bill and act references — and biasing recognition towards it. This is unglamorous, it does not appear in any benchmark, and it moves entity error rate further than a model upgrade.
{
"entity": "constituency",
"canonical": "Rajahmundry",
"variants": ["Rajamahendravaram", "Rajahmundhry", "రాజమహేంద్రవరం"],
"context_boost": ["assembly", "member for", "constituency"],
"reviewed_by": "clerk-desk",
"last_confirmed": "2026-07-30"
}A sampling plan you can defend
Continuous evaluation on a fixed test set drifts into overfitting; evaluating everything is unaffordable. We hold three sets and are explicit about what each is for.
| Set | Composition | Used for |
|---|---|---|
| Frozen | Sealed at project start, never inspected by anyone tuning the system | Release decisions only |
| Rolling | A stratified sample drawn from the last month of real sessions | Detecting drift as the subject matter and speakers change |
| Adversarial | The hardest passages reviewers have flagged — overlap, poor microphones, dense code-switching | Testing the gate, not the model: does the system flag what it gets wrong? |