Skip to content
estudIA

AI glossary

Reinforcement learning (RL)

A way of training where a model tries things, receives a reward or penalty, and gradually learns which actions pay off.

Instead of copying correct examples, the model acts and is scored. Over many attempts it learns to do more of what earns rewards. Reinforcement learning is how game-playing systems learned to beat world champions.

For language models it has become central: reasoning models and coding agents are trained with RL on tasks whose results can be checked automatically, such as whether code passes its tests or a maths answer is right.

Example: A model tries to solve thousands of coding exercises. Each time its code passes the tests it gets a reward, and little by little it learns strategies that work, such as checking edge cases before finishing.

In practice

  • It largely explains why today’s models code and reason so much better than those of only a short while ago.
  • One risk is the model learning shortcuts that earn the reward without solving the task (reward hacking).
  • That is why companies also check how the model reaches the result, not just whether it is right.

Related terms

← Back to the glossary