Skip to content

Protecting System Prompts

Why system prompts leak, what shouldn't be in them, and how to reduce the impact of extraction.

Editorial team 1 min read

System prompts often contain valuable instructions. Users frequently try to extract them — and often succeed.

Why Prompts Leak

Models can be persuaded to repeat or paraphrase their instructions through direct requests, role-play, translation tricks or gradual probing. There's no reliable way to guarantee a prompt stays secret.

What Not to Put in Prompts

  • API keys, passwords and tokens.
  • Internal URLs and infrastructure details.
  • Personal data.
  • Access-control logic ("admins can see everything; user IDs starting with 9 are admins").
  • Anything whose disclosure would cause real harm.

Reducing Impact

  • Assume the prompt will be seen; write it accordingly.
  • Enforce permissions and business rules in code, not in the prompt.
  • Keep sensitive logic server-side.

Reducing Leakage

  • Instruct the model not to reveal its instructions (helps, but isn't reliable).
  • Output filters that detect verbatim prompt disclosure.
  • Monitor for extraction attempts.

Intellectual Property

If prompts are commercially valuable, accept that determined users may reconstruct them. Competitive advantage usually lies in data, product and execution rather than prompt secrecy.

More in AI security

All AI security guides →
AI security Guide · 1 min

Introduction to AI Security

What AI security covers — attacks on models, data and AI applications — and how it differs from traditional security.

AI security 1 min read 29 Jun 2025

AI security Guide · 1 min

The OWASP Top 10 for LLM Applications

An overview of the widely used list of the most critical security risks for applications built on language models.

AI security 1 min read 28 Jun 2025

AI security Guide · 1 min

Jailbreaks: How They Work and How to Defend

How people try to get models to bypass their safety training, common techniques, and layered defences.

AI security 1 min read 27 Jun 2025

AI security Guide · 1 min

Indirect Prompt Injection

How attackers hide instructions in web pages, emails and documents that AI systems read, and why it's so dangerous for agents.

AI security 1 min read 26 Jun 2025