Prompting large language models
How LLMs work, how to write prompts that get reliable results, and how to check what comes back.
Build test sets, grade outputs with code, people and model-based graders, and catch regressions before your users do.
"It looked good in the demo" is how most LLM projects fail in production. This course shows how to measure an LLM application properly: what to test, how to grade open-ended output, and how to keep checking as prompts, models and data change.
The approach is provider-neutral and applies equally to chatbots, RAG systems, extraction pipelines and agents.
4 lessons · 1 hr
From vibes to measurable success criteria.
Representative examples, hard cases and expected results.
Three ways to score outputs, and how to combine them.
Regression testing, monitoring in production and reading the results.
Sign in and enrol to leave a review.
No reviews yet — be the first once you have worked through it.
2 min read
2 min read
2 min read
2 min read
How LLMs work, how to write prompts that get reliable results, and how to check what comes back.
Turn text into vectors, build a semantic search index, and ground an LLM's answers in your own documents.
How LLM agents call tools in a loop, how to design tools they use well, and how to keep agents safe and reliable.