AI security protects AI systems, and the organisations using them, from attack and misuse. It combines traditional security with new threats specific to machine learning.
What's Different
- Inputs are instructions: language models treat text as potential commands, so data can become an attack vector.
- Behaviour is learned: models can be manipulated through their training data.
- Outputs are probabilistic: the same input doesn't always produce the same output, which complicates testing.
- Models are valuable assets: weights and training data can be stolen.
Main Threat Areas
- Prompt injection and jailbreaks.
- Data poisoning and backdoors.
- Adversarial examples.
- Model theft and extraction.
- Privacy attacks that reveal training data.
- Supply-chain risks in models, datasets and libraries.
- Insecure agent and tool integrations.
What Stays the Same
Access control, least privilege, input validation, logging, patching and incident response all still apply — and many AI incidents come from ordinary security failures around AI systems.
Frameworks
Resources such as the OWASP Top 10 for LLM Applications, MITRE ATLAS and the NIST AI Risk Management Framework help structure the work.