Skip to content

Labelling Data for Machine Learning

How to create high-quality labels: clear guidelines, measuring agreement, handling ambiguity and reducing cost.

Editorial team 2 min read

Supervised models learn from labels, so label quality sets the ceiling on model quality.

Write Labelling Guidelines

  • Define each label clearly, with examples.
  • Include borderline and ambiguous cases, and how to decide them.
  • Explain what to do when no label fits or the input is unreadable.
  • Update the guidelines as new edge cases appear.

Measure Agreement

Have multiple people label the same sample and measure inter-annotator agreement (for example Cohen's kappa). Low agreement means the task or guidelines are unclear — and the model can't be more consistent than the humans.

Quality Control

  • Include items with known answers to check labellers.
  • Review a random sample of labels regularly.
  • Resolve disagreements with an expert or discussion.
  • Give labellers feedback.

Reduce Cost

  • Active learning: have the model pick the examples it's least sure about for labelling.
  • Pre-labelling: a model suggests labels that people confirm or correct.
  • LLM-assisted labelling: can speed up work, but check its accuracy against human labels before trusting it.
  • Weak supervision: combine rules and heuristics to label at scale, accepting some noise.

Look After Labellers

Labelling can involve distressing content or repetitive work. Provide fair pay, breaks, support and clear expectations.

Version Labels

Keep track of which guideline version produced which labels, so you can re-label consistently when definitions change.

More in Data for AI

All Data for AI guides →
Data for AI Guide · 1 min

How to Read a Dataset Card

The questions a dataset card should answer — what one row is, where the data came from, its licence and its quirks — before you use it.

Data for AI 1 min read 5 Apr 2026

Data for AI Guide · 2 min

Data Quality Dimensions

Accuracy, completeness, consistency, timeliness, validity and uniqueness: a framework for checking whether data is fit for purpose.

Data for AI 2 min read 4 Apr 2026

Data for AI Guide · 2 min

Handling Missing Data

Why data goes missing, how to find out, and the options — dropping, imputing, flagging — with their trade-offs.

Data for AI 2 min read 3 Apr 2026

Data for AI Guide · 2 min

Detecting and Handling Outliers

How to spot unusual values, decide whether they are errors or genuine extremes, and treat them appropriately.

Data for AI 2 min read 2 Apr 2026