AI glossary
Computer use
The ability of an AI agent to operate software through its graphical interface: looking at the screen, moving the pointer, clicking and typing.
Most tools give a model access through an API. Computer use is the fallback for everything without one: the model receives screenshots, decides what to do and sends mouse and keyboard actions. It can fill in forms, use old desktop apps or navigate websites.
It is slower and less reliable than an API, and risky if the agent can reach sensitive accounts, so it usually runs in an isolated virtual machine. OpenAI’s Agents API and Anthropic’s models both offer it.
Example: An agent has to download invoices from a supplier portal that has no API. It opens the browser, signs in, goes to the invoices section and downloads the month’s PDFs, looking at the screen at each step.
In practice
- Use it only when there is no API or connector: it is slower and fails more often.
- Run it in a virtual machine or a separate browser, without your personal sessions open.
- Watch out for prompt injection: a web page can show instructions aimed at the agent.
OpenAI’s Agents API with computer use features in the DevDay 2026 story.
