Skip to content
estudIA

AI glossary

Benchmark

A standard test used to compare AI models, such as a set of coding problems or exam questions with known answers.

Benchmarks give a common yardstick: SWE-bench for fixing real software bugs, for example, or various maths and science exams. They are useful for spotting broad trends, but a high score does not guarantee a model will do well on your task. Test data can leak into training, and leaderboards can be gamed.

For real decisions, build a small eval with examples from your own work.

Example: Google says Gemini 4 Argon scores 77.9% on DeepSWE v1.1, and xAI reports 71.0% for Grok 4.7. These are each company’s own figures, measured with its own set-up, so treat them as a guide, not a verdict.

In practice

  • Check exactly what the test measures and whether it resembles what you do.
  • Be wary of comparisons with different set-ups: more reasoning effort, more attempts or other tools.
  • Wait for independent results before drawing conclusions.

We cover those figures in the Gemini 4 Argon story.

Related terms

Learn more

← Back to the glossary