AI glossary
Benchmark
A standard test used to compare AI models, such as a set of coding problems or exam questions with known answers.
Benchmarks give a common yardstick: SWE-bench for fixing real software bugs, for example, or various maths and science exams. They are useful for spotting broad trends, but a high score does not guarantee a model will do well on your task. Test data can leak into training, and leaderboards can be gamed.
For real decisions, build a small eval with examples from your own work.
Example: Google says Gemini 4 Argon scores 77.9% on DeepSWE v1.1, and xAI reports 71.0% for Grok 4.7. These are each company’s own figures, measured with its own set-up, so treat them as a guide, not a verdict.
In practice
- Check exactly what the test measures and whether it resembles what you do.
- Be wary of comparisons with different set-ups: more reasoning effort, more attempts or other tools.
- Wait for independent results before drawing conclusions.
We cover those figures in the Gemini 4 Argon story.

