Security red teaming puts your AI application under realistic attack before real attackers do.
Define Scope and Goals
What would a successful attack look like? Leaking another customer's data, triggering an unauthorised action, extracting the system prompt, generating prohibited content, or running up costs.
Attack Surfaces
- The chat interface itself.
- Documents, emails and web pages the system reads.
- Tool inputs and outputs.
- File uploads, including images.
- APIs and authentication.
Techniques
- Direct and indirect prompt injection.
- Jailbreak attempts.
- Data exfiltration through links, images or tool calls.
- Privilege escalation through tools.
- Denial of service with huge or looping inputs.
Automated and Manual
Automated tools generate many attack variants; skilled humans find creative, context-specific attacks. Use both.
Report and Fix
Record reproduction steps, impact and suggested fixes. Prioritise fixes that reduce capability or exposure, not just prompt tweaks.
Regression
Add successful attacks to your test suite and rerun after changes. Repeat red teaming when the system changes significantly.