Fine-tuning adapts models to your needs — and creates specific security risks.
Risks
- Poisoned training data planting backdoors or harmful behaviours.
- Sensitive data memorisation: fine-tuned models may reproduce personal or confidential training examples.
- Safety degradation: fine-tuning can weaken a model's safety training, even with benign-looking data.
- Unauthorised access to training jobs, data and resulting models.
Controls
- Curate data: review sources, remove secrets and unnecessary personal data, and check for manipulation.
- Restrict access: who can submit data, launch jobs and deploy models.
- Version everything: datasets, configurations and model outputs, with traceability.
- Evaluate safety after tuning, not just task performance: rerun safety and security tests.
- Test for memorisation: probe whether the model reproduces training records.
- Protect outputs: store fine-tuned models securely; they may encode sensitive information.
Provider Fine-Tuning
When using a provider's fine-tuning service, check data retention, isolation and use terms.
Consider Alternatives
For adding knowledge, retrieval often avoids embedding sensitive data in model weights.