Skip to content

Speech to text with Whisper

Transcribe and translate audio with OpenAI's open Whisper model, handle long recordings, and measure accuracy with word error rate.

Free on glitchdata intermediate 4 lessons 56 min

What you'll learn

  • Explain how Whisper turns audio into text
  • Transcribe and translate recordings in Python
  • Handle long audio and timestamps
  • Measure accuracy with word error rate

About this course

Whisper is an open speech recognition model that runs on your own machine. This course uses the Whisper tiny model from the hub to transcribe and translate audio, then covers the practical parts: long recordings, timestamps, choosing a model size, and measuring accuracy honestly.

Install with pip install transformers torch jiwer, plus ffmpeg for reading audio files.

Before you start

  • Basic Python

Course content

4 lessons · 56 min

  1. 1
    How speech recognition works

    From sound waves to spectrograms to text.

    Free preview 12 min
  2. 2
    Transcribing and translating audio

    Run Whisper from Python on a recording.

    16 min
  3. 3
    Long recordings and timestamps

    Chunking, timestamps for subtitles, and preparing audio.

    14 min
  4. 4
    Measuring accuracy with word error rate

    Compare transcripts to a reference and read the result.

    14 min

What learners say

Sign in and enrol to leave a review.

No reviews yet — be the first once you have worked through it.

Related guides

All Applied AI guides →

More in Applied AI

Applied AI intermediate

Forecasting time series

Trend, seasonality, honest baselines and time-based testing, using NASA temperatures and bike rentals from the hub.

4 lessons 1 hr Free