Skip to content
estudIA

AI glossary

Latency

How long you wait for a model’s answer: the time to the first token plus the time to generate the rest.

Two numbers matter: time to first token (how quickly the answer starts appearing) and output speed (tokens per second). Larger models and long reasoning increase latency; smaller models and streaming the answer as it is generated reduce the wait.

For a chat, a few seconds is fine. For voice assistants or autocomplete, every hundred milliseconds counts, which is why those use fast, small models.

Example: In a voice assistant, if the answer takes three seconds to start, the conversation feels broken. For a report you will spend ten minutes reading, waiting thirty seconds does not matter.

In practice

  • Show the answer as it is generated (streaming): the wait feels much shorter.
  • Use small models or less reasoning effort where speed matters.
  • With long prompts, prompt caching also cuts the time to the first word.

Related terms

← Back to the glossary