AI glossary
Inference
Running a trained model to get an output — every time you send a prompt and receive an answer.
Training builds the model once; inference is using it, millions of times. Inference cost is what you pay per token in an API, and inference speed is how fast tokens appear on screen.
It depends on model size, hardware, how long the input and output are, and how much the model reasons before answering.
Example: When you ask a chatbot something, the model runs inference: it processes your message and generates the answer token by token. If a thousand people ask at once, a thousand inferences are running on the provider’s servers.
In practice
- Cost grows with input and output tokens: shorter prompts and more concise answers save money.
- Small models, prompt caching and batch processing make inference cheaper.
- With an open-weights model you can run inference on your own hardware.
Prices per million tokens for each model: model comparison.