Skip to content

Model Extraction and Theft

How attackers steal models by copying weights or imitating them through queries, and how to protect model assets.

Editorial team 1 min read

Trained models embody expensive data, compute and expertise. Attackers may try to steal them.

Direct Theft

Stealing model weights through compromised servers, insider access, leaked credentials or insecure storage. For frontier models, weight security is a major concern.

Extraction Through Queries

Querying a model many times and training a copy on the responses — sometimes called distillation or model stealing. The copy may approximate the original's behaviour for a fraction of the cost.

Risks

  • Loss of competitive advantage.
  • Stolen models used without safety safeguards.
  • Copies used to develop attacks against the original.

Protecting Weights

  • Strict access control and multi-person approval.
  • Encryption at rest and in transit.
  • Isolated, monitored infrastructure.
  • Insider-risk programmes.

Limiting Extraction

  • Rate limits and quotas per user.
  • Monitoring for high-volume, systematic querying.
  • Terms of service prohibiting training on outputs.
  • Returning only necessary outputs, not full probability distributions.

Balance

Protections must be weighed against legitimate use; excessive restrictions frustrate real customers.

More in AI security

All AI security guides →
AI security Guide · 1 min

Introduction to AI Security

What AI security covers — attacks on models, data and AI applications — and how it differs from traditional security.

AI security 1 min read 29 Jun 2025

AI security Guide · 1 min

The OWASP Top 10 for LLM Applications

An overview of the widely used list of the most critical security risks for applications built on language models.

AI security 1 min read 28 Jun 2025

AI security Guide · 1 min

Jailbreaks: How They Work and How to Defend

How people try to get models to bypass their safety training, common techniques, and layered defences.

AI security 1 min read 27 Jun 2025

AI security Guide · 1 min

Indirect Prompt Injection

How attackers hide instructions in web pages, emails and documents that AI systems read, and why it's so dangerous for agents.

AI security 1 min read 26 Jun 2025