Supervised models learn from labels, so label quality sets the ceiling on model quality.
Write Labelling Guidelines
- Define each label clearly, with examples.
- Include borderline and ambiguous cases, and how to decide them.
- Explain what to do when no label fits or the input is unreadable.
- Update the guidelines as new edge cases appear.
Measure Agreement
Have multiple people label the same sample and measure inter-annotator agreement (for example Cohen's kappa). Low agreement means the task or guidelines are unclear — and the model can't be more consistent than the humans.
Quality Control
- Include items with known answers to check labellers.
- Review a random sample of labels regularly.
- Resolve disagreements with an expert or discussion.
- Give labellers feedback.
Reduce Cost
- Active learning: have the model pick the examples it's least sure about for labelling.
- Pre-labelling: a model suggests labels that people confirm or correct.
- LLM-assisted labelling: can speed up work, but check its accuracy against human labels before trusting it.
- Weak supervision: combine rules and heuristics to label at scale, accepting some noise.
Look After Labellers
Labelling can involve distressing content or repetitive work. Provide fair pay, breaks, support and clear expectations.
Version Labels
Keep track of which guideline version produced which labels, so you can re-label consistently when definitions change.