Skip to content
estudIA

AI glossary

RLHF

Training a model with reinforcement learning where the reward comes from human judgements of which answers are better.

RLHF stands for reinforcement learning from human feedback. After pretraining, a model knows a lot but is not yet a helpful assistant. In RLHF, people compare pairs of answers and pick the better one. Those choices train a reward model, which then scores the language model’s answers during reinforcement learning.

RLHF is a big part of why chat assistants follow instructions, decline harmful requests and write in a helpful tone. Variants replace some human feedback with AI feedback guided by written principles.

Example: People are shown two answers to “how do I relax before an exam?”: one helpful and kind, the other curt and generic. They pick the first. With thousands of choices like this, the model learns which kind of answer people prefer.

In practice

  • It explains why assistants tend to be polite and careful.
  • It can also make them eager to please, inclined to agree with you: if you want blunt criticism, ask for it.
  • Variants such as constitutional AI use written principles to guide part of the ratings.

Related terms

← Back to the glossary