AI glossary
Latency
How long you wait for a model’s answer: the time to the first token plus the time to generate the rest.
Two numbers matter: time to first token (how quickly the answer starts appearing) and output speed (tokens per second). Larger models and long reasoning increase latency; smaller models and streaming the answer as it is generated reduce the wait.
For a chat, a few seconds is fine. For voice assistants or autocomplete, every hundred milliseconds counts, which is why those use fast, small models.
Example: In a voice assistant, if the answer takes three seconds to start, the conversation feels broken. For a report you will spend ten minutes reading, waiting thirty seconds does not matter.
In practice
- Show the answer as it is generated (streaming): the wait feels much shorter.
- Use small models or less reasoning effort where speed matters.
- With long prompts, prompt caching also cuts the time to the first word.