Skip to content
estudIA

AI glossary

Multimodal

Able to work with more than one type of data — for example text, images, audio, video or PDFs — in the same model.

A multimodal model can, for instance, read a screenshot and explain an error, describe a chart, transcribe a meeting or answer questions about a PDF with tables and images. Most leading models now accept images and documents as input; some also take audio and video.

Input and output modalities can differ: many models read images but only write text, while image generators do the opposite.

Example: You take a photo of a crumpled invoice and ask the assistant to pull out the amount, date and tax number into a table. It reads the image and gives you the data.

In practice

  • Images count as tokens too and can make a request more expensive.
  • Check what it reads from blurry images or small print: it can get it wrong.
  • Not every model supports everything: DeepSeek V4-Pro, for example, cannot read images. We note it in the model comparison.

Related terms

Learn more

← Back to the glossary