AI glossary
Multimodal
Able to work with more than one type of data — for example text, images, audio, video or PDFs — in the same model.
A multimodal model can, for instance, read a screenshot and explain an error, describe a chart, transcribe a meeting or answer questions about a PDF with tables and images. Most leading models now accept images and documents as input; some also take audio and video.
Input and output modalities can differ: many models read images but only write text, while image generators do the opposite.
Example: You take a photo of a crumpled invoice and ask the assistant to pull out the amount, date and tax number into a table. It reads the image and gives you the data.
In practice
- Images count as tokens too and can make a request more expensive.
- Check what it reads from blurry images or small print: it can get it wrong.
- Not every model supports everything: DeepSeek V4-Pro, for example, cannot read images. We note it in the model comparison.


