Which LLMs actually hold up running a browser-use agent in production
This list is for people who already have a browser agent working on a laptop and now need it to run unattended. Maybe you’re on the browser-use library, maybe on something similar. The question stops being “which model is smartest” and becomes “which model and provider combination doesn’t break at 3am”. I’m Xavier Fok, I run agents from Singapore, and most of my production pain has come from the seam between a model and the framework around it, not from the model’s raw ability.
That seam is where this list focuses. The reader question that prompted it is browser-use issue 2242. I’m not going to paraphrase that thread, because bug threads change faster than articles do. Open it, read the latest comments, and check whether your model and provider pair is named there. What I can give you is the method for answering it yourself, plus how I rank the families once you’ve done that check.
One honest limit up front. I didn’t run a fresh benchmark for this piece and I’m not publishing any numbers I can’t source. The ranking below is my operator judgment, built from running these families behind agent loops and from the way their APIs are shaped. Anything about prices or API behaviour is marked as of October 2026, and you should confirm it on the vendor page before you commit a budget.
how I picked
I scored every candidate on the same short list, and I put open integration problems ahead of capability claims.
- open integration bugs: for each model and provider pair, I look at the framework’s issue tracker for open reports touching that provider’s client. A model with a great demo and three open tool-call parsing bugs loses to a duller model with none.
- structured output reliability: browser agents emit an action every step, usually as a JSON schema or a tool call. If the provider’s schema mode is strict, a malformed action is rare. If it’s loose, you write retry code.
- vision handling: most browser agents send a screenshot or a DOM snapshot with every step. I care about how the provider handles image size, how many images per request, and whether it fails loudly or silently when you exceed limits.
- cost per step, not per token: an agent run is dozens of steps, each with a growing context. A cheap model that needs twice the steps isn’t cheap.
- rate limits and outage behaviour: can I get a stable quota, and does the API return clear errors I can back off on?
I dropped anything I couldn’t reach through a normal API with a documented client. Hobby wrappers and unofficial proxies of consumer chat apps are out, partly for reliability and partly because they usually break the vendor’s terms.
the picks
Anthropic Claude Sonnet 5.5
Sonnet 5.5 (claude-sonnet-5-5) is where I start for a new browser agent. Anthropic documents computer use and tool use as first-class features, and the computer use tool docs are the clearest primary source on how screenshots, coordinates and actions are meant to flow. When a framework’s Anthropic client is well maintained, you get a model that was built around the screenshot-then-act loop.
My caution is the usual one. Check the tracker for your framework version and the Anthropic client together. Open reports about tool-call formatting or message ordering are the kind of thing that turns a good model into a flaky agent, and you’d rather learn that before deploy.
- pros: tool use is a documented, central API feature, so action output is predictable
- pros: strong at multi-step tasks where the page changes under it
- pros: vendor docs for the agent pattern are unusually clear
- cons: per-step cost adds up on long runs with image-heavy context
- cons: your framework’s Anthropic client version matters, so pin it and test upgrades
Pricing: per-token, input and output billed separately. As of October 2026, check Anthropic’s pricing page for current Sonnet rates before budgeting.
Link: Anthropic computer use tool docs
Anthropic Claude Opus 5.5
Opus 5.5 (claude-opus-5-5) is the model I escalate to, not the one I run by default. I use it for the steps that fail on Sonnet: ambiguous pages, long forms with dependencies, or recovery after the agent has gone off course. Splitting a run so the expensive model only handles the hard branch keeps the bill sane.
Because it shares the same API shape as Sonnet, the integration surface is the same too. That’s a real operational advantage. One client, one set of known bugs, two capability tiers. You can swap the model string per step without rewriting your adapter.
- pros: highest ceiling in the Anthropic family for hard, ambiguous pages
- pros: same client and schema path as Sonnet, so no second integration to maintain
- pros: good as a fallback model when a cheaper model fails a step
- cons: slow and expensive if used for every step
- cons: easy to overspend by forgetting it’s set as the default
Pricing: per-token, premium tier. As of October 2026, confirm on Anthropic’s pricing page.
Link: browser-use repository
Anthropic Claude Haiku 5.5
Haiku 5.5 (claude-haiku-5-5) is the cheap, fast tier, and it earns its place on the simple half of an agent’s work. Think navigation, reading a table, clicking an obvious button, or classifying a page before a bigger model acts. Same API family as the other two, so the same adapter works.
I wouldn’t hand it a whole long-horizon task and walk away. Used as a router or a first-pass worker, it cuts cost. Used as the only brain on a complicated workflow, it tends to need more retries, and retries erase the savings.
- pros: low cost and low latency per step
- pros: shares the Anthropic client path, so no extra integration risk
- pros: good fit for page classification and simple extraction
- cons: more likely to need retries on complex multi-step tasks
- cons: weaker recovery when a page behaves unexpectedly
Pricing: per-token, lowest tier of the family. As of October 2026, check Anthropic’s pricing page.
Link: see playwright vs puppeteer for AI browser agents for the driver layer that sits under whichever model you pick.
OpenAI GPT flagship
OpenAI’s current flagship GPT models are the other obvious default. The strongest practical argument for them is ecosystem gravity. Almost every agent framework tested its OpenAI client first, so you are least likely to be the first person to hit a bug there. Structured outputs with strict JSON schema is a documented API feature, and that is exactly what an action-per-step loop wants.
The catch is that “OpenAI” is not one integration. Chat Completions, the Responses API, and Azure-hosted OpenAI each have slightly different behaviour, and a framework may support them unevenly. When I check the tracker, I search for the specific endpoint I’m going to use, not the vendor name.
- pros: most heavily tested client in most frameworks, so fewer surprises
- pros: strict structured outputs suit per-step action schemas
- pros: broad quota tiers and mature account tooling
- cons: Chat Completions, Responses and Azure variants behave differently, so test the one you will ship
- cons: model naming and defaults shift, so pin an exact model ID
Pricing: per-token. As of October 2026, see OpenAI’s API pricing page for the model you choose. I’m not copying figures here because they move.
Link: prompt injection for browser agents, explained, which matters for any model that reads untrusted page text.
Google Gemini
Gemini is the pick when cost per step and long context are the main constraints. Its Flash tier is built for high-volume, low-cost calls, and a browser agent that sends a screenshot every step is exactly that workload. Google publishes current rates on its Gemini API pricing page, which is the place to check before you do any maths.
The integration question is whether your framework talks to Gemini through Google’s own SDK, an OpenAI-compatible endpoint, or a wrapper library. Each path has had its own quirks around schema handling. Pick one path, pin the version, and read the open issues for that path specifically before you rely on it.
- pros: low per-step cost on the Flash tier suits long, image-heavy runs
- pros: large context window helps when you keep history in the prompt
- pros: pricing and limits are documented on one official page
- cons: schema handling differs between SDK, compatibility endpoint and wrappers
- cons: smaller body of community debugging than the two bigger vendors
Pricing: per-token, with free and paid tiers. As of October 2026, confirm on the official Gemini pricing page.
Link: Gemini API pricing
Open-weight models (Qwen family) behind an OpenAI-compatible server
If you need data to stay on your own hardware, or your volume makes per-token pricing painful, open-weight vision-language models served locally are the realistic route. I’d look at the Qwen family first because it ships vision-capable variants, and serve it through vLLM or Ollama using the OpenAI-compatible API. That keeps your framework adapter identical to the hosted OpenAI path.
This is the pick with the most integration risk. Strict schema mode depends on your inference server, not the model, and tool-call parsing varies by chat template. You’re taking on GPU cost, uptime and upgrades. For a single operator running a small fleet, I’d only do it when privacy or volume truly demands it.
- pros: no per-token bill and full control over where data goes
- pros: the OpenAI-compatible endpoint reuses the same framework adapter
- pros: you decide when the model changes, so no surprise deprecations
- cons: schema and tool-call behaviour depend on your inference server and chat template
- cons: you own GPU capacity, monitoring and failover
Pricing: no per-token fee, but you pay for hardware or rented GPUs. As of October 2026, that cost depends entirely on your setup.
Link: how to sandbox a computer use agent, worth reading before you run any self-hosted agent against the open web.
comparison table
| pick | price | primary strength | primary weakness |
|---|---|---|---|
| Claude Sonnet 5.5 | per-token, mid tier, check Anthropic pricing (as of October 2026) | documented tool use and screenshot loop | cost on long image-heavy runs |
| Claude Opus 5.5 | per-token, premium tier, check Anthropic pricing | best fallback for hard, ambiguous pages | too costly as a default |
| Claude Haiku 5.5 | per-token, lowest Anthropic tier | cheap and fast for simple steps | more retries on complex tasks |
| OpenAI GPT flagship | per-token, check OpenAI pricing (as of October 2026) | most tested client, strict structured outputs | endpoint variants behave differently |
| Google Gemini | per-token with free tier, check Gemini pricing | low cost per step on Flash, long context | schema handling varies by SDK path |
| Qwen on vLLM or Ollama | hardware or GPU rental cost | data stays local, no per-token fee | schema and parsing depend on your server |
how to choose
Start from the tracker, not the leaderboard. Before you pick a model, search the open issues of your framework for the exact client you will use: the vendor, the endpoint, and the version. If the thread behind issue 2242 touches your pair, treat that as a reason to test harder or switch, and re-check it the day you deploy. A pair with zero open reports isn’t proven safe, but it’s a better bet than one with an active, unresolved report.
Run a two-tier setup if cost matters. A cheap model such as Haiku or Gemini Flash handles navigation and simple reads, and an expensive one such as Opus takes over when a step fails twice. Keep both on the same provider if you can, because one client means one set of bugs. The same logic applies to how you run many sessions at once, which I covered in the best ways to orchestrate many Playwright browser sessions for a fleet of agents.
Pin everything. Pin the framework version, the provider SDK version and the exact model ID. A floating alias that changes underneath you looks identical to a framework regression in your logs. When you upgrade, do it on a branch with a fixed set of replay tasks, and keep checkpointing so a crash doesn’t lose a long run. For that, see self-hosted LangGraph checkpointing vs LangGraph Cloud for long tool calls and the related LangGraph durability sync vs async piece. If your agent loops, how to stop a LangGraph agent from looping until it hits the recursion limit covers the guard rails.
Last, think about where the browser appears to come from. Your own sites and consenting users’ accounts sometimes behave differently by region, and if you are testing a Singapore-facing flow on your own property, a real local connection helps you see what a local user sees. Singapore Mobile Proxy sells real Singapore mobile IPs with sticky sessions, and it is Singapore-only. That’s only relevant if your testing is genuinely Singapore-specific. If not, skip it, and keep all of this within each site’s terms. For the proxy layer in general, see residential vs mobile proxies for browser agents. The rest of our writing on agents is in the blog index.
verdict / top pick
If I had to ship one thing this week, it would be Claude Sonnet 5.5 as the default model, with Opus 5.5 as the escalation path and Haiku 5.5 for cheap steps. The reason isn’t a benchmark. It’s that one vendor, one client and one documented tool-use path gives me the fewest places for an integration bug to hide, and the vendor’s own docs describe the exact loop a browser agent runs.
OpenAI’s flagship is a close second, and for some teams it’s first, since the framework clients are so heavily exercised. Gemini is my choice when per-step cost dominates. The open-weight route is for privacy or volume, and only if you accept running the infrastructure.
Whatever you pick, open the issue tracker for your exact pair on deploy day. That ten-minute check has saved me more than any model upgrade.
Written by Xavier Fok
disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-10-08.