Skip to content
HandbookPublic
Applied AI

Evaluating Indic speech recognition honestly

Why word error rate misleads on Indian-language audio, what to measure instead, and how to design the human gate that stands between a model and an official record.

Written forTeams putting speech or language models into a workflow with consequences
Reading time13 min read
Last reviewed2026-08-09

A vendor demonstration of speech recognition is almost always a clean recording of one speaker in a quiet room. The system we run transcribes a legislative chamber: overlapping speech, a public gallery, microphones at varying distances, formal Telugu, colloquial Telugu, English clauses inside Telugu sentences, procedural formulae, and names that appear nowhere in any general lexicon. The gap between those two situations is where every honest evaluation lives.

Word error rate is the wrong headline

WER divides insertions, deletions and substitutions by reference length and treats every word as equal. On Indian-language audio that assumption breaks in three specific ways, and each one moves the number in a direction that flatters the system.

01
Agglutination punishes near-missesTelugu builds long words by suffixation. A single wrong suffix on a compound scores as one substitution regardless of whether the meaning survived, while an English sentence expressing the same content spreads the same error across several tokens. WER is therefore not comparable across languages, and a Telugu figure sitting beside an English figure in a proposal is a category error.
02
Code-switching inflates the denominatorReal speech switches to English for procedural and technical terms. Those tokens are usually the easiest to recognise, so a heavily code-switched passage produces a better WER than a purely Telugu one while being no better understood. Report WER per language segment or it flatters the mix.
03
The words that matter are rareNames, constituencies, bill numbers and dates carry nearly all the consequence, and are a small fraction of tokens. A transcript at 8% WER can be unusable if the errors sit in that fraction, and a transcript at 15% can be fine if they do not.
MeasureWhat it tells youWhy we keep it
CER (character)Degradation independent of tokenisationThe only figure comparable across scripts and languages
Entity error rateAccuracy on names, numbers, dates, placesThe measure that predicts whether the record is usable
Diarisation error rateWhether the right speaker is attributedIn a chamber, attributing a sentence to the wrong member is worse than mis-transcribing it
Semantic error rateWhether meaning survived, judged by a reviewer on a sampleCatches the fluent, plausible, wrong output that character measures miss entirely
Review minutes per audio hourThe real operating costThe number the client feels; it decides whether the system is adopted

Measure against the manual baseline, not against a leaderboard

Every institution already has a process. Before touching a model, measure it: how long does the existing process take, how many errors does it contain, where do they occur, and what happens when one is found. That baseline is the only thing the client cares about, and it is usually not written down anywhere.

  • Sample the existing output and score it with the same rubric you will apply to the model. Human transcripts contain errors too, and knowing the rate changes what "good enough" means.
  • Time the current end-to-end cycle including review and sign-off, not just the transcription step. Systems are adopted or rejected on cycle time.
  • Record where corrections happen today. Those are exactly the places the interface must make easy, and they are rarely where an engineer would guess.

Designing the human gate

Nothing from a model enters an official record unreviewed. That principle is easy to state and mostly implemented badly: a reviewer facing a wall of plausible text approves it, and the gate becomes a rubber stamp that adds latency without adding assurance. A gate works only when it directs attention.

01
Surface uncertainty at token levelHighlight low-confidence spans and every recognised entity, so the reviewer's eye goes to the five per cent that carries the risk rather than reading uniformly.
02
Show the audio at the cursorA reviewer must be able to hear the exact span under the cursor without hunting on a timeline. This single affordance moved review speed more than any model change we made.
03
Make correction cheaper than acceptanceIf fixing a name takes four clicks and accepting takes zero, the gate degrades under time pressure. Inline edit, keyboard-first, with the previous correction offered for the same entity.
04
Record the correction as training signalEvery edit is a labelled example. Captured properly, the domain lexicon and the model improve from the reviewers' work rather than from a separate annotation project.
05
Sign the approvalThe approved version is an event with an actor, a timestamp and a hash of what was approved. Without that, "reviewed by a human" is a claim rather than a fact.

The domain lexicon is most of the win

On institutional audio, the highest-return work is not model selection. It is assembling the vocabulary the institution actually uses — member names with their spellings, constituencies, committee names, procedural phrases, bill and act references — and biasing recognition towards it. This is unglamorous, it does not appear in any benchmark, and it moves entity error rate further than a model upgrade.

json
{
  "entity": "constituency",
  "canonical": "Rajahmundry",
  "variants": ["Rajamahendravaram", "Rajahmundhry", "రాజమహేంద్రవరం"],
  "context_boost": ["assembly", "member for", "constituency"],
  "reviewed_by": "clerk-desk",
  "last_confirmed": "2026-07-30"
}
Lexicon entries carry variants because speech does, and because reviewers should never fix the same name twice.

A sampling plan you can defend

Continuous evaluation on a fixed test set drifts into overfitting; evaluating everything is unaffordable. We hold three sets and are explicit about what each is for.

SetCompositionUsed for
FrozenSealed at project start, never inspected by anyone tuning the systemRelease decisions only
RollingA stratified sample drawn from the last month of real sessionsDetecting drift as the subject matter and speakers change
AdversarialThe hardest passages reviewers have flagged — overlap, poor microphones, dense code-switchingTesting the gate, not the model: does the system flag what it gets wrong?