Automatic speech recognition (ASR) converts spoken audio into text. Modern systems use deep learning and handle accents and noise far better than older approaches.
How It Works
Audio is converted into a representation such as a spectrogram, which shows which frequencies are present over time. A neural network — often an encoder–decoder transformer — maps that representation to text. Open models such as Whisper were trained on hundreds of thousands of hours of audio.
Choosing a Model
- Languages and accents your users speak.
- Accuracy versus speed: larger models are more accurate but slower.
- Streaming or batch: live captions need low latency.
- Deployment: cloud API or on-device/self-hosted for privacy.
- Extras: timestamps, punctuation, speaker diarisation (who spoke when).
Measuring Accuracy
Word error rate (WER) counts substitutions, deletions and insertions relative to a correct transcript. Lower is better. Normalise text (case, punctuation, numbers) before comparing, and test on your own audio.
Improving Results
- Better audio at the source: closer microphones, less background noise.
- Trim long silences, where some models hallucinate text.
- Provide the language if known, and vocabulary hints where supported.
Risks
Accuracy varies across accents, languages and speech impairments, so check performance for all the people who will use the system. Recordings are personal data: get consent and handle them securely.