System prompts often contain valuable instructions. Users frequently try to extract them — and often succeed.
Why Prompts Leak
Models can be persuaded to repeat or paraphrase their instructions through direct requests, role-play, translation tricks or gradual probing. There's no reliable way to guarantee a prompt stays secret.
What Not to Put in Prompts
- API keys, passwords and tokens.
- Internal URLs and infrastructure details.
- Personal data.
- Access-control logic ("admins can see everything; user IDs starting with 9 are admins").
- Anything whose disclosure would cause real harm.
Reducing Impact
- Assume the prompt will be seen; write it accordingly.
- Enforce permissions and business rules in code, not in the prompt.
- Keep sensitive logic server-side.
Reducing Leakage
- Instruct the model not to reveal its instructions (helps, but isn't reliable).
- Output filters that detect verbatim prompt disclosure.
- Monitor for extraction attempts.
Intellectual Property
If prompts are commercially valuable, accept that determined users may reconstruct them. Competitive advantage usually lies in data, product and execution rather than prompt secrecy.