AI systems often depend on personal data. Privacy law and user trust both require handling it with care.
Principles
- Lawful basis and purpose: have a legitimate reason for using the data, and don't repurpose it without checking that's allowed.
- Minimisation: collect and use only what's needed.
- Retention limits: delete data when no longer needed.
- Security: protect data in storage and transit, with access controls.
- Transparency: tell people how their data is used.
- Rights: support access, correction and deletion requests.
De-Identification
Removing names isn't enough: combinations of ordinary fields (postcode, birth date, sex) can re-identify people. Techniques include aggregation, generalisation (age bands instead of birth dates), suppressing rare combinations and adding noise. Assess re-identification risk rather than assuming it's zero.
Models Can Leak Data
Models may memorise and reproduce training examples, especially rare ones. Language models trained on raw text can repeat personal information. Mitigations include removing personal data before training, deduplication, and privacy-preserving training techniques such as differential privacy.
Using Third-Party AI Services
Check where data is processed, whether it's retained or used for training, and what contractual protections exist before sending personal data to external APIs.
Privacy Impact Assessments
For projects involving significant personal data or new uses of it, run a privacy impact assessment early — it's often legally required and always useful.
This guide is general information, not legal advice.