Securing AI systems: the threat landscape
What actually changes when a model joins a system, the attacks that follow from it, and how to threat model an AI feature before you ship it.
Attacks on models themselves: evasion, poisoning and backdoors, model theft, and what the data remembers.
Before language models, the field of attacking machine learning already had a decade of results — and most of them still apply, including to the models inside today's AI products.
This course covers the four families of attack on a model: making it wrong at inference, corrupting what it learns, stealing it, and extracting the data it was trained on. It is deliberately practical about which defences survive contact with a real adversary.
4 lessons · 1 hr
Small, deliberate changes to an input that flip a model's answer.
Corrupting what a model learns, and planting behaviour that only appears on a trigger.
Extraction, inversion and membership inference: what an API gives away.
Adaptive attacks broke most published defences; here is what survived.
Sign in and enrol to leave a review.
No reviews yet — be the first once you have worked through it.
1 min read
1 min read
1 min read
1 min read
What actually changes when a model joins a system, the attacks that follow from it, and how to threat model an AI feature before you ship it.
How injection works, why filtering fails, and the design patterns that actually contain it.
Access control across a retrieval index, the confused deputy problem in tool use, and keeping an agent inside its blast radius.