AI glossary
Synthetic data
Training data created by AI models or programs rather than collected from people, used to teach skills where real examples are scarce.
As good human-written text becomes harder to find, labs generate data: worked maths problems, code with tests, conversations, translations. A strong model can produce examples for a weaker one (see distillation), and programs can check the answers automatically.
The risk is a feedback loop: if models train mostly on their own output, errors and blandness can accumulate. Careful filtering and verification are key.
Example: To teach a model to solve equations, thousands of exercises are generated by a program that knows the answer. The model’s answers are checked automatically and only the correct ones are kept.
In practice
- It helps when real data is scarce or private.
- Mixing it with real data and reviewing it stops the model from picking up bad habits.
- If you generate data with a commercial model, check its terms of use.