Skip to content

Evaluating LLM applications

Build test sets, grade outputs with code, people and model-based graders, and catch regressions before your users do.

Free on glitchdata advanced 4 lessons 1 hr

What you'll learn

  • Write success criteria that can actually be measured
  • Build a representative test set, including hard cases
  • Choose between code-based, human and model-based grading
  • Run evaluations continuously to catch regressions

About this course

"It looked good in the demo" is how most LLM projects fail in production. This course shows how to measure an LLM application properly: what to test, how to grade open-ended output, and how to keep checking as prompts, models and data change.

The approach is provider-neutral and applies equally to chatbots, RAG systems, extraction pipelines and agents.

Before you start

  • Prompting large language models

Course content

4 lessons · 1 hr

  1. 1
    Why "it looked good" isn't enough

    From vibes to measurable success criteria.

    Free preview 12 min
  2. 2
    Building a test set

    Representative examples, hard cases and expected results.

    16 min
  3. 3
    Grading: code, people and model graders

    Three ways to score outputs, and how to combine them.

    18 min
  4. 4
    Evaluating continuously

    Regression testing, monitoring in production and reading the results.

    14 min

What learners say

Sign in and enrol to leave a review.

No reviews yet — be the first once you have worked through it.

More in Generative AI