A backdoored model behaves normally on ordinary inputs but does something specific when a hidden trigger appears.
Examples
- An image classifier that labels any picture with a small sticker as "authorised".
- A language model that produces insecure code when a certain phrase appears.
- A sentiment model that always rates reviews containing a trigger word as positive.
How Backdoors Are Planted
- Poisoning training data with trigger examples.
- Modifying a model directly and redistributing it.
- Compromised fine-tuning pipelines.
Why They're Hard to Detect
The model passes standard tests because triggers are absent from test data. Research has shown some backdoors can persist through additional safety training.
Reducing Risk
- Obtain models from trusted sources and verify checksums or signatures.
- Prefer providers with documented training processes.
- Control access to training and fine-tuning pipelines.
- Test models with varied inputs, including unusual tokens and patterns.
- Monitor production behaviour for anomalies.
- Use interpretability and detection tools where available.
Downloading Open Models
Treat third-party models like third-party code: vet the source, scan files, and prefer safe serialisation formats.