AI glossary
Red teaming
Deliberately attacking an AI system, the way an adversary would, to find dangerous behaviour and security holes before release.
Red teams try to make a model misbehave: extract dangerous information, bypass its guardrails, leak data or get an agent to do something harmful. What they find goes into fixes, extra training and the safety reports published with new models.
It combines in-house experts, outside specialists and automated attacks generated by other models. For agents, red teaming focuses heavily on prompt injection.
Example: Before launching an agent that reads emails, a team sends it messages with hidden instructions to see whether it obeys. They find it forwards a file when it should not, and add a mandatory confirmation.
In practice
- If you launch an AI product, do your own red teaming, even on a small scale.
- Test the most likely kinds of abuse for your case, not just generic ones.
- Repeat the tests every time you change the model or the prompt.
