Skip to content

01 · Magazine

Browser Use: AI-Driven Browser Automation

Browser Use explained: how the Python library lets an AI agent control a browser, getting started with code, use cases, limitations and alternatives.

· Reviewed

Browser Use is an open-source Python library that lets an AI agent operate a web browser: opening pages, clicking, filling in forms and reading content. You describe the task in natural language, and a language model decides step by step what happens in the browser. Besides the library, there is a command-line tool that gives existing agents such as Claude Code or Codex browser access, and a paid cloud offering with hosted browsers. This article explains how it works, how to get started, where it is used and where its limits are.

Fact sheet (as of October 2026)

Characteristic Details
License MIT
GitHub more than 110,000 stars (as of October 2026)
Language Python 3.11 or later
Installation uv add browser-use or pip install browser-use
Models OpenAI, Anthropic, Google, Browser Use's own models, local models via Ollama
Ways to use it Python library, CLI for existing agents, hosted cloud API
Cloud Hosted browsers with proxies and CAPTCHA handling, usage-based billing

How Browser Use works

  1. Browser Use starts a Chromium browser (locally or in the cloud) and controls it via the Chrome DevTools Protocol.
  2. The current page state is prepared for the model: a simplified list of the interactive elements from the DOM, optionally supplemented with a screenshot.
  3. The model decides on the next action, such as "click element 12" or "type text into field 4".
  4. Browser Use executes the action, reads the new state and repeats the loop until the task is done or a step limit is reached.

The approach is therefore based mainly on the DOM rather than on pure image recognition. That makes it fast and relatively inexpensive, as long as the page has a usable DOM.

Getting started

uv init --python 3.12
uv add browser-use

Add the model provider's API key to .env, then:

import asyncio
from browser_use import Agent, ChatOpenAI  # or ChatAnthropic, ChatGoogle, ChatBrowserUse

async def main():
    agent = Agent(
        task="Open the Deutsche Bahn website and find the next connection from Cologne to Berlin.",
        llm=ChatOpenAI(model="<model-name>"),
    )
    history = await agent.run()
    print(history.final_result())

asyncio.run(main())

With Browser(use_cloud=True) and a BROWSER_USE_API_KEY, the browser runs in the Browser Use cloud instead. You can register your own functions via Tools, for example to write results directly to a database.

Typical use cases

  • Filling in forms: Transferring data from a spreadsheet into web portals that have no API.
  • Data comparison and research: Collecting prices, availability or contact details across several sites.
  • Testing: Describing user flows in natural language and having them checked.
  • Agents with browser access: Through the CLI, a coding agent gets the ability to open web pages, for example to check a web application it is building.

Strengths

  • Quick start: A few lines of code for a working agent.
  • Free choice of model, including local models.
  • Large community with many examples and integrations into agent frameworks.
  • Usable locally and in the cloud, with the option to reuse an existing Chrome profile for logged-in sessions.

Limitations

  • Reliability: On complex or unusual pages, the agent gets lost or gives up. For recurring, critical processes, conventional automation with Playwright is often more robust.
  • Cost per run: Every step is a model call. Long processes become expensive and slow.
  • Bot detection: Many sites block automated browsers. The cloud offers countermeasures, but there is no guarantee.
  • Legal aspects: Website terms of use, data protection and copyright apply to agents too. Before you deploy one, clarify whether you are allowed to retrieve the data automatically.
  • Security: Web pages can contain instructions that manipulate the agent (prompt injection). Do not give the agent access to accounts it does not need for the task.

Browser Use compared

Browser Use Skyvern Stagehand
Approach DOM-based agent Vision models plus Playwright Playwright-like API with AI commands
Language Python Python, TypeScript client TypeScript, Python, Go
License MIT AGPL-3.0 MIT
Strength Open-ended tasks with little code Sites with changing layouts, workflows Combining fixed code with AI steps

You can find the direct comparison under Browser Use vs. Skyvern, and an overview of Stagehand in the Stagehand profile.

Who is Browser Use for?

  • Python developers who quickly need an agent with browser access
  • Teams using agent frameworks who want to add browser tasks as a tool
  • Prototypes and internal tools where occasional failures are acceptable

Frequently asked questions

Is Browser Use free?

The library is licensed under MIT and free of charge. You pay for model calls with your provider or use a local model. Only the hosted cloud with its own browsers, proxies and CAPTCHA handling is billed by usage.

Which model do I need for Browser Use?

Browser Use works with models from OpenAI, Anthropic and Google, with Browser Use's own models and with local models via Ollama. Test with your real pages, because the success rate depends heavily on how well the model understands page structures.

How is it different from Playwright or Selenium?

With Playwright or Selenium you write every click as code. Browser Use lets a language model decide what happens next. That is more flexible on unknown pages, but slower, more expensive and less predictable. For fixed, critical processes, classic automation usually remains the better choice.

Can Browser Use work with logged-in accounts?

Yes. You can reuse an existing Chrome profile with logged-in sessions. Only give the agent access to accounts it needs for the task, because websites can contain hidden instructions.

Sources

Last reviewed:

Market notes

Get the next analysis by email

New articles and changes in the directory, at most once a week.