AI glossary
Evals (evaluations)
A set of test cases with expected results used to measure how well an AI system performs on a specific task.
An eval is like a test suite for prompts and AI features. You collect realistic inputs, define what a good answer looks like, and score the outputs — automatically, with another model as a judge, or by hand. Each time you change the prompt, model or settings, you rerun the eval and see whether things improved.
Even 20–50 well-chosen cases catch most regressions.
Example: Before switching models, a team runs its 40 ticket-classification test cases with the current model and the new one. The new one gets more right but fails two billing cases, so they adjust the prompt before switching.
In practice
- Include hard and unusual cases, not just typical ones.
- Decide beforehand what a good answer looks like: a clear criterion makes scoring easier.
- If you use a model as a judge, check a sample by hand to make sure it scores well.
How to improve a prompt step by step: prompting techniques.
