Skip to content

Securing AI Agents

Practical controls for agents that use tools and act autonomously: least privilege, isolation, approval and monitoring.

Editorial team 1 min read

Agents raise AI security stakes: a manipulated model that can only chat produces bad text, while a manipulated agent can take harmful actions.

Least Privilege

Give agents only the tools, data and permissions needed for their task. Use scoped, short-lived credentials rather than broad keys.

Isolation

Run agents in sandboxes with restricted file system and network access. Keep secrets outside the sandbox.

Break Dangerous Combinations

The highest risk arises when an agent can access private data, process untrusted content and communicate externally at once. Remove one where possible.

Human Approval

Require confirmation for irreversible, financial or external actions, with clear descriptions of what will happen.

Validate Actions

Check tool arguments against policies in code — not just instructions in the prompt.

Monitor

Log every action; alert on unusual patterns such as bulk data access or unexpected destinations.

Identity

Give each agent its own identity, so its actions are distinguishable from users' and can be audited and revoked.

Test

Red-team agents with injection attempts through every content source they read.

More in AI security

All AI security guides →
AI security Guide · 1 min

Introduction to AI Security

What AI security covers — attacks on models, data and AI applications — and how it differs from traditional security.

AI security 1 min read 29 Jun 2025

AI security Guide · 1 min

The OWASP Top 10 for LLM Applications

An overview of the widely used list of the most critical security risks for applications built on language models.

AI security 1 min read 28 Jun 2025

AI security Guide · 1 min

Jailbreaks: How They Work and How to Defend

How people try to get models to bypass their safety training, common techniques, and layered defences.

AI security 1 min read 27 Jun 2025

AI security Guide · 1 min

Indirect Prompt Injection

How attackers hide instructions in web pages, emails and documents that AI systems read, and why it's so dangerous for agents.

AI security 1 min read 26 Jun 2025