Skip to content

Lesson 1 of 4

Free preview

Why "it looked good" isn't enough

From vibes to measurable success criteria.

12 min 3-question quiz 3 guides to read next

LLM outputs vary from run to run, work well on typical inputs and fail on unusual ones, and change whenever you edit a prompt or the provider updates a model. Trying a handful of examples by hand can't reveal any of that.

Evaluation starts with success criteria that are:

  • Specific: not "good summaries" but "summaries under 100 words that include every figure mentioned in the source and contain no facts that aren't in it".
  • Measurable: you can check them automatically, or with a clear rubric a person can apply consistently.
  • Tied to the use case: a support bot might be judged on correct answers, correct refusals, tone and whether it escalates when it should.

Split criteria into:

  • Must-haves — failures here are serious (leaking personal data, giving a wrong price, missing a safety warning). Aim for zero.
  • Quality measures — scores you want to raise over time (helpfulness, concision, style).

Write these down before you tune a single prompt. They decide what "better" means for everything that follows.

Check your understanding

3 questions · pass with 2 correct

1. Which success criterion is well written?
2. What are 'must-haves'?
3. Why isn't trying a handful of examples by hand enough?

You'll see your score; enrol to have it count towards your certificate.

Enrol for free to save your progress, unlock every lesson and earn a certificate.

Sign in to enrol

Further reading

Guides that go deeper on this lesson.

  • Evaluating LLM Outputs

    How to measure the quality of language model outputs with test sets, code checks, human review and model-based grading.

    2 min read

  • Measuring LLM Quality With Benchmarks

    What public benchmarks measure, why leaderboard scores can mislead, and how to complement them with your own tests.

    1 min read

  • Defining Good Metrics and KPIs

    How to choose metrics that reflect real goals, define them precisely and avoid the traps of vanity metrics and gaming.

    2 min read