A model learns from its data. If the data doesn't represent the real world it will be used in, the model will fail — often for the people least represented.
Common Gaps
- Groups: under-representation of certain ages, regions, languages, accents or skin tones.
- Conditions: images only in good lighting; speech only from quiet rooms.
- Time: data from a period that no longer reflects current behaviour.
- Channels: data from one source, such as web users only, when the model will serve phone users too.
Checking Representativeness
- Describe the population and conditions the model will face.
- Compare the dataset's make-up against that description.
- Evaluate model performance separately for important groups and conditions.
- Pay special attention to small groups, where performance is often worst.
Improving Coverage
- Collect more data from under-represented groups and conditions.
- Partner with communities to collect data respectfully and with consent.
- Use stratified sampling when building datasets.
- Use augmentation carefully to broaden conditions.
- Reweight examples during training.
Document Limitations
When gaps can't be fixed, state them in the dataset and model cards, and restrict use accordingly.
Revisit Over Time
Populations and behaviour change. Periodically compare production data with training data and refresh datasets.