Skip to content
estudIA

AI glossary

Guardrails

Checks and limits placed around an AI system to keep its inputs and outputs safe, on-topic and within policy.

Guardrails can be simple (a list of allowed actions, a spending cap, requiring human approval before sending an email) or use another model to screen inputs and outputs for harmful content, leaked secrets or off-topic requests.

For agents, the most effective guardrails are often permissions: give the agent only the tools and access it truly needs, and ask for confirmation before irreversible actions.

Example: A shopping agent can search for products and add them to the basket, but not pay: payment needs a person to confirm it, and there is a spending limit per order.

In practice

  • Combine several layers: permissions, limits, automatic checks and human approval.
  • Test your guardrails by trying to get around them, as an attacker would.
  • Log what the agent does so you can review it later.

How to use agents safely: what an AI agent is.

Related terms

Learn more

← Back to the glossary