Skip to content

Securing Fine-Tuning Pipelines

Protecting the data, access and outputs of model fine-tuning against poisoning, leakage and misuse.

Editorial team 1 min read

Fine-tuning adapts models to your needs — and creates specific security risks.

Risks

  • Poisoned training data planting backdoors or harmful behaviours.
  • Sensitive data memorisation: fine-tuned models may reproduce personal or confidential training examples.
  • Safety degradation: fine-tuning can weaken a model's safety training, even with benign-looking data.
  • Unauthorised access to training jobs, data and resulting models.

Controls

  • Curate data: review sources, remove secrets and unnecessary personal data, and check for manipulation.
  • Restrict access: who can submit data, launch jobs and deploy models.
  • Version everything: datasets, configurations and model outputs, with traceability.
  • Evaluate safety after tuning, not just task performance: rerun safety and security tests.
  • Test for memorisation: probe whether the model reproduces training records.
  • Protect outputs: store fine-tuned models securely; they may encode sensitive information.

Provider Fine-Tuning

When using a provider's fine-tuning service, check data retention, isolation and use terms.

Consider Alternatives

For adding knowledge, retrieval often avoids embedding sensitive data in model weights.

More in AI security

All AI security guides →
AI security Guide · 1 min

Introduction to AI Security

What AI security covers — attacks on models, data and AI applications — and how it differs from traditional security.

AI security 1 min read 29 Jun 2025

AI security Guide · 1 min

The OWASP Top 10 for LLM Applications

An overview of the widely used list of the most critical security risks for applications built on language models.

AI security 1 min read 28 Jun 2025

AI security Guide · 1 min

Jailbreaks: How They Work and How to Defend

How people try to get models to bypass their safety training, common techniques, and layered defences.

AI security 1 min read 27 Jun 2025

AI security Guide · 1 min

Indirect Prompt Injection

How attackers hide instructions in web pages, emails and documents that AI systems read, and why it's so dangerous for agents.

AI security 1 min read 26 Jun 2025