AI glossary
Output limit (max output tokens)
The maximum number of tokens a model can generate in a single response, separate from how much it can read.
The context window is how much the model can take in; the output limit is how much it can write back in one go. Output limits are usually much smaller: 64K or 128K tokens are common in 2026, while Gemini 4 Argon raised it to 1 million.
It matters for long documents, big code changes and translations. If you hit it, the answer is cut off and you need to ask the model to continue or split the task.
Example: You ask for a 300-page manual to be translated in one go and the answer stops halfway. The context window did not fill up: the output limit was reached.
In practice
- For long texts, ask for the work in parts, for example chapter by chapter.
- In the API, set the maximum output tokens with some headroom so the answer is not cut off.
- Current limits: 128,000 tokens on Claude Opus 5.5 and GPT-6 Astra, and 64,000 on Claude Haiku 4.5.
The jump to a million output tokens, in the Gemini 4 Argon story.