Skip to content
estudIA

AI glossary

Batch processing (batch API)

Sending many requests at once to be processed in the background within hours, usually at about half the normal price.

If you do not need an answer right away — classifying thousands of tickets, summarising an archive, running an evaluation — you can submit everything as a batch. The provider runs it when it has spare capacity and you collect the results later.

Anthropic, OpenAI and Google charge half price for requests sent in batches. Combined with prompt caching, it is one of the easiest ways to cut API costs.

Example: A company wants to classify 50,000 old emails. Instead of sending them one by one, it submits them as a batch overnight and has the results at half price the next morning.

In practice

  • Use it for anything that can wait a few hours: Google, for example, aims to finish batches within 24 hours.
  • Each request in the batch has its own ID: use it to match answers to your data, because they may come back in a different order.
  • Test the prompt with a few normal requests first: a mistake in a batch is repeated thousands of times.

Related terms

← Back to the glossary