Computer vision with pretrained models
Classify images with a pretrained ResNet-50 in ONNX Runtime, then adapt a pretrained network to your own categories.
Transcribe and translate audio with OpenAI's open Whisper model, handle long recordings, and measure accuracy with word error rate.
Whisper is an open speech recognition model that runs on your own machine. This course uses the Whisper tiny model from the hub to transcribe and translate audio, then covers the practical parts: long recordings, timestamps, choosing a model size, and measuring accuracy honestly.
Install with pip install transformers torch jiwer, plus ffmpeg for reading audio files.
4 lessons · 56 min
From sound waves to spectrograms to text.
Run Whisper from Python on a recording.
Chunking, timestamps for subtitles, and preparing audio.
Compare transcripts to a reference and read the result.
Sign in and enrol to leave a review.
No reviews yet — be the first once you have worked through it.
2 min read
2 min read
2 min read
2 min read
Classify images with a pretrained ResNet-50 in ONNX Runtime, then adapt a pretrained network to your own categories.
Trend, seasonality, honest baselines and time-based testing, using NASA temperatures and bike rentals from the hub.