AI glossary
RLHF
Training a model with reinforcement learning where the reward comes from human judgements of which answers are better.
RLHF stands for reinforcement learning from human feedback. After pretraining, a model knows a lot but is not yet a helpful assistant. In RLHF, people compare pairs of answers and pick the better one. Those choices train a reward model, which then scores the language model’s answers during reinforcement learning.
RLHF is a big part of why chat assistants follow instructions, decline harmful requests and write in a helpful tone. Variants replace some human feedback with AI feedback guided by written principles.
Example: People are shown two answers to “how do I relax before an exam?”: one helpful and kind, the other curt and generic. They pick the first. With thousands of choices like this, the model learns which kind of answer people prefer.
In practice
- It explains why assistants tend to be polite and careful.
- It can also make them eager to please, inclined to agree with you: if you want blunt criticism, ask for it.
- Variants such as constitutional AI use written principles to guide part of the ratings.