Skip to content
estudIA

AI agents2 min read

Agents That Use the Browser and the Computer

How agents that browse and operate software for you work (Claude in Chrome, the Agents API, Muse), what they're good for and what precautions to take.

Until recently, an AI assistant could only talk to you. Browser and computer use agents go a step further: they see the screen, move the pointer, click and type, as you would. That lets them fill in forms, compare products across websites or operate software that has no API.

How they work

The agent works in a loop, like any agent:

  1. It receives a screenshot (or the page content).
  2. The model, which is multimodal, decides the next action: click here, type this, scroll.
  3. The tool performs the action and takes another screenshot.
  4. It repeats until it finishes or needs your confirmation.

Examples that already exist

  • Claude in Chrome: an extension for desktop Chrome, available on Claude’s paid plans, that reads pages, clicks, types and moves between tabs from a side panel.
  • OpenAI’s Agents API: includes a computer use tool so developers can build agents that operate software through its interface.
  • Muse, from Meta: a personal agent that handles errands in the background (emails, bookings, purchases), with a second agent watching what goes out to the internet. US only for now; see this story.

What they’re good for

  • Repetitive tasks on websites without an API: copying data from a portal into a spreadsheet, downloading invoices from several sites.
  • Comparing: prices, features or terms across several pages.
  • Filling in long forms with data you already have.
  • Testing websites: development teams use them to check an app works the way a person would use it.

Their limits

  • They’re slow compared with a direct integration: each step means looking at the screen and thinking.
  • They make mistakes with confusing interfaces, pop-ups or design changes.
  • They use more tokens, because they process images at every step.

If there’s an API or an MCP connector for what you want to do, it’s usually faster, cheaper and more reliable than an agent that clicks.

Precautions

An agent that browses reads content you don’t control, and that’s the main risk: prompt injection. A website can hide instructions to trick it.

  • Start with sessions without important accounts: don’t let it browse with your bank or email open until you trust it.
  • Confirmation before paying, sending or deleting. Any serious product asks for it; don’t turn it off.
  • Supervise the first few times: watch what it does, step by step.
  • Narrow tasks: “compare these three hotels” beats “plan my holiday”.

More measures in agent safety.

Frequently asked questions

Can an agent buy or book things for me?

Technically yes, and some products offer it. Always do it with confirmation before paying and start with low-risk tasks.

Why are they slower than an API?

Because they work like a person: they look at the screen, decide, click and wait for the page to load. Each step is a model call with a screenshot.

Glossary terms

Sources

Related articles