Skip to content
estudIA

AI glossary

Rate limit

The maximum number of requests or tokens you may send to an API in a period of time, such as per minute.

Providers limit usage to share capacity fairly and prevent abuse. If you exceed the limit the API returns an error (often HTTP 429) and you must wait and retry. Limits usually grow as you spend more or move to a higher usage tier.

In practice: retry with increasing waits, spread work over time and use batch processing for large jobs that are not urgent. Chat subscriptions have their own usage limits, often reset every few hours.

Example: You fire off 500 requests at once to translate a website. Past a certain number, the API returns a 429 error (“too many requests”); your program waits a few seconds and retries the ones that failed.

Related terms

← Back to the glossary