Skip to content
estudIA

AI glossary

Synthetic data

Training data created by AI models or programs rather than collected from people, used to teach skills where real examples are scarce.

As good human-written text becomes harder to find, labs generate data: worked maths problems, code with tests, conversations, translations. A strong model can produce examples for a weaker one (see distillation), and programs can check the answers automatically.

The risk is a feedback loop: if models train mostly on their own output, errors and blandness can accumulate. Careful filtering and verification are key.

Example: To teach a model to solve equations, thousands of exercises are generated by a program that knows the answer. The model’s answers are checked automatically and only the correct ones are kept.

In practice

  • It helps when real data is scarce or private.
  • Mixing it with real data and reviewing it stops the model from picking up bad habits.
  • If you generate data with a commercial model, check its terms of use.

Related terms

← Back to the glossary