Automatic transcription has quietly become very good. On a clean recording of one person speaking clearly, modern systems are competitive with a human typist and enormously faster. On a phone recording of four people in a busy cafe, they produce something close to nonsense. The gap between those two outcomes is almost entirely about the recording, not the software.
What the accuracy numbers mean
Transcription quality is usually quoted as word error rate — the percentage of words wrong, missing or invented. Human transcribers on decent audio sit around 4–5%; nobody is perfect.
| Error rate | What that feels like |
|---|---|
| Under 5% | Excellent. Occasional name or technical term to fix. |
| 5–10% | Good. Clearly readable, needs a proofread. |
| 10–20% | Usable as a rough draft; noticeable editing required. |
| Over 20% | Often faster to retype than to fix. |
A 5% error rate sounds small but means roughly one wrong word per sentence — which is why review is never optional for anything you intend to publish.
What actually breaks transcription
Distance from the microphone
The single biggest factor, and the most fixable. Sound falls off sharply with distance while room noise stays constant, so a speaker across the room is quieter relative to everything else. A cheap microphone close to someone beats an expensive one far away, every time.
Overlapping speech
Two people talking at once is genuinely hard. The models are trained mostly on one voice at a time, and interruptions frequently lose both halves.
Background noise and music
Steady noise like a fan is handled reasonably. Music with vocals is the worst case, because the system cannot tell which voice it should be transcribing.
Proper nouns and jargon
Names of people, companies, products and places are the most common errors even in otherwise excellent transcripts. The model is choosing the most probable word, and your colleague's surname is not probable. Industry terminology fails the same way.
Accents and speaking style
Accuracy varies with how well an accent is represented in the training data. Fast speech, trailing off at the end of sentences, and heavy filler words all reduce accuracy too.
Six things that improve results
- Get closer to the microphone. Thirty centimetres instead of three metres will do more than any setting.
- Record somewhere quiet. Close windows, turn off fans, avoid rooms with a strong echo — hard walls create reflections that smear consonants.
- One person at a time. If you are running the conversation, let people finish.
- Speak at a normal, unhurried pace. Rushing drops consonants, which is exactly what recognition depends on.
- Clean the audio first. Running noise reduction before transcription measurably helps on a noisy recording.
- Tell it the language. Selecting the right language explicitly beats leaving it to auto-detect, especially on short clips.
Reviewing efficiently
Do not read the transcript cold. Play the audio and follow along — errors that look plausible on the page are obvious the moment you hear what was actually said. Check names first, since they are both the most likely to be wrong and the most conspicuous when they are. Then search for homophone pairs: their/there, to/too, its/it's. Recognition picks by probability, so it gets these wrong in ways a human never would.
Where automatic transcription is genuinely good enough
- Searchable archives. A 90% accurate transcript makes hours of recordings findable, which is transformative even with errors.
- Personal notes. Nobody else reads them.
- First drafts of subtitles. Correcting is far faster than typing from scratch.
- Getting the gist of a long recording you have not time to listen to.
Where it is not enough on its own: anything published, legal or medical records, and quotations attributed to a named person.
Common questions
Should I pick the language or let it detect?
Pick it when you know it. Auto-detection works from a short sample and can misfire on brief or noisy clips, and a wrong guess ruins the whole transcript.
Will it identify who is speaking?
No. This produces a transcript of what was said, not a labelled record of who said it. Adding speaker labels is a separate step you would do by hand.
Does it handle two languages in one recording?
Poorly. Transcription assumes a single language throughout, so code-switching mid-sentence tends to produce garbled output for the second language.
Why did it invent a sentence that was never said?
Speech models predict likely word sequences, so during silence or unintelligible audio they can occasionally generate plausible-sounding text from nothing. It is uncommon but real, and another reason to review against the audio rather than trusting the page.