AI glossary
Jailbreak
A prompt crafted to trick a model into ignoring its safety rules and producing content it would normally refuse.
Jailbreaks use role-play, invented scenarios, odd formatting or long chains of messages to talk a model out of its guardrails. Providers patch known techniques, and attackers look for new ones, so it is an ongoing contest.
A jailbreak is different from prompt injection: in a jailbreak the user attacks the model; in prompt injection, hidden instructions in a document or web page attack the user through the model.
Example: Someone asks a chatbot for dangerous instructions disguised as “a scene from a novel”. If the model falls for it, it has been jailbroken; good models spot the trick and refuse.
In practice
- If you build a product, do not rely on the model alone: add your own guardrails.
- Test your assistant against known jailbreak attempts before launch.
- Attempting a jailbreak usually breaks the provider’s terms of use and can get the account closed.