Skip to content

Adversarial machine learning

Attacks on models themselves: evasion, poisoning and backdoors, model theft, and what the data remembers.

Free on glitchdata advanced 4 lessons 1 hr

What you'll learn

  • Explain adversarial examples and why they transfer between models
  • Distinguish poisoning from backdoors, and know where each enters
  • Describe model extraction and membership inference, and what they cost an owner
  • Judge which published defences are worth deploying

About this course

Before language models, the field of attacking machine learning already had a decade of results — and most of them still apply, including to the models inside today's AI products.

This course covers the four families of attack on a model: making it wrong at inference, corrupting what it learns, stealing it, and extracting the data it was trained on. It is deliberately practical about which defences survive contact with a real adversary.

Before you start

  • How machine learning works, or equivalent
  • Comfort with the idea of training and evaluation

Course content

4 lessons · 1 hr

  1. 1
    Evasion: adversarial examples

    Small, deliberate changes to an input that flip a model's answer.

    Free preview 15 min
  2. 2
    Poisoning and backdoors

    Corrupting what a model learns, and planting behaviour that only appears on a trigger.

    16 min
  3. 3
    Stealing models, and the data inside them

    Extraction, inversion and membership inference: what an API gives away.

    15 min
  4. 4
    Which defences are worth deploying

    Adaptive attacks broke most published defences; here is what survived.

    14 min

What learners say

Sign in and enrol to leave a review.

No reviews yet — be the first once you have worked through it.

More in AI security

AI security intermediate

Securing RAG, tools and agents

Access control across a retrieval index, the confused deputy problem in tool use, and keeping an agent inside its blast radius.

4 lessons 1 hr 3 min Free