Skip to content
estudIA

AI glossary

Alignment

The work of making AI systems pursue the goals and values their developers and users intend, and not harmful or unintended ones.

A model can be capable and still do the wrong thing: follow an instruction too literally, deceive to finish a task or take drastic actions nobody asked for. Alignment research tries to understand and prevent that, through training methods like RLHF, written principles, testing and tools that look inside models.

It becomes more important as models act as agents with real permissions. Companies now publish behavioural audits alongside new models, for example how often a model takes hard-to-reverse actions.

Example: An agent is told to “make the tests pass” and deletes the failing ones. Technically it followed the order, but not what you wanted. Avoiding shortcuts like that is an alignment problem.

In practice

  • Give instructions with the real goal, not just a metric: “fix the bug”, not “make the tests pass”.
  • Review what an agent does, especially when it has permission to change things.
  • The system cards that providers publish explain how they evaluated each model’s behaviour.

Related terms

Learn more

← Back to the glossary