Skip to content

Consent and Data Rights for Training Data

Where training data comes from, what rights and licences apply, and how to respect creators and data subjects.

Editorial team 2 min read

The data used to train AI raises questions about consent, copyright and fairness to the people and creators behind it.

Sources of Training Data

  • Data you collect directly from users or customers.
  • Licensed datasets.
  • Public datasets with published licences.
  • Web-scraped content.
  • Synthetic data.

Personal Data

If training data contains personal information, privacy law applies: you need a lawful basis, must respect purpose limits, and may need to handle deletion requests. Data collected for one purpose can't automatically be used to train models for another.

Datasets and content come with licences that may restrict commercial use, require attribution or prohibit redistribution. The legal position on training AI with copyrighted material is contested and varies by country. Record the source and licence of every dataset you use.

Respecting Creators

Consider whether creators expected their work to be used this way, honour opt-outs where offered (such as robots.txt and other machine-readable signals), and prefer licensed or openly licensed content.

Good Practice

  • Maintain a data inventory with sources, licences and consent basis.
  • Prefer datasets with clear documentation and licences.
  • Remove personal data you don't need.
  • Provide ways for people to opt out where appropriate.
  • Get legal advice for significant commercial uses.

Why It Matters

Beyond legal risk, respecting data rights builds trust with users, customers and creators.

More in Responsible AI

All Responsible AI guides →
Responsible AI Guide · 2 min

What Is Responsible AI?

The principles behind responsible AI — fairness, transparency, privacy, safety, accountability — and how to turn them into practice.

Responsible AI 2 min read 30 Apr 2026

Responsible AI Guide · 2 min

Understanding Bias in AI Systems

Where bias in AI comes from — data, labels, design and deployment — and why removing sensitive attributes isn't enough.

Responsible AI 2 min read 29 Apr 2026

Responsible AI Guide · 2 min

Fairness Metrics for Machine Learning

Demographic parity, equal opportunity, equalised odds and calibration: what each measures and why they can't all be satisfied at once.

Responsible AI 2 min read 28 Apr 2026

Responsible AI Guide · 2 min

Privacy in AI Projects

How to handle personal data responsibly when building AI: minimisation, purpose limits, de-identification and the risks of models leaking data.

Responsible AI 2 min read 27 Apr 2026