← all articles

Best ways to orchestrate many Playwright sessions for an agent fleet

This list is for people who already have one Playwright agent working and now need twenty, or two hundred, running at once. I’m an operator in Singapore, and I run browser agents against my own properties and against sites where I have the user’s consent. The question I keep hearing is some version of “how do I run lots of Playwright sessions in parallel without it turning into a mess?” If you arrived from playwright issue 18892, I’m answering that kind of question directly. Short answer: for a fleet of agents, run your own pool of Playwright servers or containers, put a queue in front, and treat the test runner’s sharding and Selenium Grid as special cases. The rest of this piece explains why.

There’s no single winner for everyone. It depends on whether you already own Selenium infrastructure, whether your agents are long-lived or short-lived, and how much glue code you’ll maintain. I ranked on integration pain, meaning the work from “works on my laptop” to “works for a fleet”, and I say plainly where each option hurts.

I haven’t run formal benchmarks, so there are no throughput or memory numbers here. Measure on your own hardware with your own pages. Anything that changes over time is dated as of October 2026.

how I picked

  • integration pain: how many moving parts sit between your agent code and a running browser, and how many of them are yours to maintain.
  • fit for agents rather than tests: agents hold a session open for minutes or hours, need state between steps, and fail in messy ways. a tool built for short, independent test files scores lower.
  • failure handling: what happens when a browser crashes, a node dies, or a session hangs. can you kill and restart one session without touching the rest.
  • state and identity: can each agent keep its own cookies, storage and network identity, and can you persist that between runs.
  • cost shape: licence fees, compute, and the engineering hours. I only quote prices I can point to, and I’ll say when the answer is “free software, you pay for machines.”

the picks

Your own browser pool (Playwright run-server plus connect)

This is my top pick for most agent fleets. You run Playwright’s server on a set of machines, each agent calls browserType.connect with a WebSocket endpoint, and a small allocator in front decides which endpoint each agent gets. The Playwright API docs for browserType.connect describe the client side. The server side is npx playwright run-server, which exposes a browser over a WebSocket.

What you get is a clean split. Your agent code stays a normal Playwright script. The browsers live somewhere else, so your agent process can be small and the heavy Chromium processes can sit on beefier boxes. You decide the pool size, the machine layout, and the retry policy. When a browser wedges, you kill that one endpoint and the allocator hands out another.

The integration pain is real but it’s all yours to design. You write the allocator, the health checks and the cleanup of orphaned browsers. You also have to keep client and server Playwright versions matched, because connect expects them to line up. I pin one Playwright version in a single image and ship it to both sides. If you want agents to survive restarts, read my notes on how to persist browser sessions for AI agents, because a remote browser doesn’t remember anything you don’t save.

  • pros: full control over placement, versions and retries
  • pros: agent code stays plain Playwright, so there’s no new API to learn
  • pros: easy to add a per-agent proxy or isolated profile at the allocator
  • cons: you build and run the allocator, health checks and garbage collection yourself
  • cons: version pinning between client and server is your job

Pricing: the software is free (Apache 2.0). you pay for the machines and the engineering time.

Playwright’s built-in sharding

Sharding is what the Playwright test runner gives you out of the box. You run npx playwright test --shard=1/4 on four machines and each one takes a quarter of the suite. The official sharding docs cover the flags and the CI patterns, including merging reports from the shards afterwards.

I’m including it because people reach for it first, and for the right workload it’s excellent. If your “agents” are really deterministic scripts written as Playwright tests, such as nightly checks of your own site’s checkout flow, sharding is the least painful way to spread them across machines.

Where it breaks down is real agents. Sharding divides a known list of tests among a fixed number of machines at launch. It doesn’t give you a live pool you can hand a new long-running task to. A model-driven agent that runs for forty minutes and holds state doesn’t fit the “split the files, run, exit” shape. You can force it, but you end up writing your own scheduler on top, at which point the first pick is cleaner.

  • pros: zero extra infrastructure beyond your CI or a few machines
  • pros: well documented and maintained by the Playwright team
  • pros: great for scheduled, deterministic checks
  • cons: assumes a fixed batch of tests, not a rolling stream of agent tasks
  • cons: no live session reuse or per-agent identity out of the box

Pricing: free. you pay for CI minutes or machines.

Selenium Grid

If your company already runs Selenium Grid, this is the pick that saves you the most procurement and ops pain. Grid 4 is a hub-and-node system with a router, distributor and session map, and the Selenium Grid documentation explains how to deploy it standalone, as hub and nodes, or fully distributed. Playwright can talk to it. The Playwright Selenium Grid page describes pointing Playwright at a Grid through the SELENIUM_REMOTE_URL environment variable.

The appeal is obvious. Someone has already solved node registration, session queuing, timeouts and dashboards. If your platform team owns the Grid, you get capacity without building a pool.

The catch is that Grid was designed around WebDriver sessions, and Playwright is not a WebDriver client. As I read the docs, Playwright’s support here is aimed at Chromium-family browsers and comes with limitations compared with using Playwright’s own browser builds, so check the current page before you commit. Session lifetimes are another thing to check: Grid has timeouts built for test sessions, and a long-lived agent can run into them. Test your longest realistic agent run before you promise anything to anyone.

  • pros: reuse an existing, battle-tested scheduler and ops setup
  • pros: central place for capacity, queuing and visibility
  • cons: Playwright through Grid is a compatibility layer, not the native path, so expect gaps
  • cons: session timeouts and queue behaviour were designed for short test sessions

Pricing: free and open source. cost is nodes and whoever operates them.

One container per session with the official Playwright image

This is the isolation-first pick. Each agent session gets its own container built from Microsoft’s official Playwright image, and the container is thrown away when the session ends. The Playwright Docker docs cover the image and the recommended run flags for Chromium, including --init and --ipc=host, which exist to avoid zombie processes and memory-sharing crashes inside the container.

I like this for fleets that touch sensitive pages or run model-driven actions. If an agent goes sideways, the blast radius is one container. It pairs naturally with the sandbox thinking in my piece on how to sandbox a computer use agent, and it’s a sensible partner to the defences in prompt injection for browser agents, explained, since a compromised session can’t see its neighbours.

The cost is startup overhead and density. Spinning up a fresh container and browser per task is slower than reusing a warm one, and each container carries its own browser memory. You’ll need to measure how many fit per host. If each agent needs its own network identity, see residential vs mobile proxies for browser agents and the walkthrough on SOCKS5 vs HTTP proxies for a Playwright agent that needs authentication.

One narrow case where an outside service helps: if you’re testing your own Singapore-facing site as a local mobile user and need a real Singapore mobile IP with a sticky session, Singapore Mobile Proxy sells exactly that. It’s Singapore-only, so it’s no use if your agents need to appear anywhere else.

  • pros: strongest isolation per session, and clean teardown
  • pros: official image keeps browser dependencies and versions consistent
  • pros: per-session network and profile settings are straightforward
  • cons: cold start per task and per-container memory limit your density
  • cons: you still need something to schedule and clean up containers

Pricing: the image is free. compute is your cost. Docker’s own licensing for large companies is a separate matter, so check Docker’s current terms (as of October 2026) if you run Docker Desktop at work. plain Docker Engine on Linux servers is a different case.

Kubernetes jobs with a queue in front

Once you’re past a handful of hosts, I’d put the container-per-session idea on Kubernetes and drive it from a queue. A producer pushes tasks, a controller or worker launches one Job or pod per task from the Playwright image, and the pod reports back and exits. Capacity becomes a cluster autoscaling question instead of a “which host has room” question.

What this buys you is scheduling you don’t have to write. Resource limits stop one greedy Chromium from starving the node, restart policies handle crashed pods, and network policies let you enforce which agents can reach which sites. If you also use LangGraph for the agent logic, the orchestration of the browser fleet and the checkpointing of the agent are separate concerns, and my piece on self-hosted LangGraph checkpointing vs LangGraph Cloud for long tool calls covers the second half.

The pain here is operational weight. If you don’t already run Kubernetes, standing it up for browser agents is a project, not a weekend. Pod startup latency, image pulls on fresh nodes, and shared-memory settings for Chromium all need tuning. Pick this when you already have the cluster or when the scale genuinely demands it.

  • pros: scheduling, limits, restarts and autoscaling are solved problems
  • pros: policy controls over what each agent workload can reach
  • cons: heavy to operate if you don’t already run it
  • cons: startup latency and image distribution need real tuning

Pricing: Kubernetes itself is free. managed services (GKE, EKS, AKS) charge for the control plane and nodes, and those prices change, so check your provider’s current page (as of October 2026).

Long-lived workers with persistent contexts

The last pick is a different shape. Instead of throwing browsers away, you run long-lived worker processes, each owning a small number of Playwright persistent contexts or saved storage states, and each tied to one agent identity. A queue assigns tasks to the worker that owns that identity. Logged-in state, cookies and local storage carry across tasks without re-authenticating every time.

This is the right fit when your agents act on behalf of consenting users or on your own accounts, and re-logging in per task would be slow or would trigger extra verification on the site’s side. It’s cheaper per task than a cold container because the browser is already warm. The full pattern, including what to save and where, is in how to persist browser sessions for AI agents.

The downsides are the usual ones for stateful systems. A corrupted profile now affects every future task for that identity, and you need a plan for in-flight tasks when a worker dies. I’d combine this with the first pick: workers connect to pooled browsers, and the persisted state lives in your own store rather than on any one machine. If a given agent also needs a particular network identity, the guide on how to route a Playwright agent through a mobile proxy shows the wiring.

  • pros: fast per task, since browsers and logins stay warm
  • pros: natural fit for one-agent-one-identity setups
  • cons: state corruption and stale sessions become your problem
  • cons: harder to scale down cleanly, and restarts need care

Pricing: free software. cost is always-on compute plus storage for saved state.

comparison table

pick price primary strength primary weakness
own browser pool (run-server and connect) free software, you pay for machines full control, plain Playwright code you build the allocator and health checks
built-in sharding free, CI minutes or machines zero extra infrastructure for test batches wrong shape for long-running agents
Selenium Grid free, nodes and ops time reuses existing scheduler and ops Playwright support is a compatibility layer
container per session image free, compute cost strongest isolation and clean teardown cold starts and lower density
Kubernetes jobs plus queue free software, managed cluster fees vary scheduling and autoscaling solved heavy to operate from scratch
long-lived persistent workers free, always-on compute warm browsers and kept logins state corruption and harder scale-down

how to choose

If you already run Selenium Grid and your agents are mostly short tasks, start there and see how far it takes you. Run your longest realistic agent session through it before you rely on it. If session timeouts or missing Playwright features bite, you’ll know in an afternoon and you can move to your own pool with little wasted work.

If your “agents” are really scripted checks of your own site, use sharding and stop there. Don’t stretch it into a scheduler for open-ended agent tasks. The moment you catch yourself writing code to dispatch new work into running shards, switch to a pool with a queue.

If your agents are model-driven, long-running, or touch pages where a misbehaving session could do damage, go container-per-session, and add Kubernetes only when you have enough hosts that placement becomes painful. Pair it with clear rules about which sites each agent may visit, and keep the work to consenting users, your own properties, and each site’s terms.

If identity continuity matters more than isolation, for example agents that operate a user’s own logged-in account with their permission, use persistent workers and put the saved state in a store you control. Whichever way you go, decide early where session state lives and who can read it, because that decision is hard to change later.

verdict / top pick

For most teams running a real fleet of agents, my top pick is your own browser pool using Playwright’s server and connect, with a queue in front and containers as the unit of isolation. It keeps your agent code as ordinary Playwright, it puts the allocator and failure handling in code you can read, and it grows into Kubernetes later without a rewrite. The price is that you build the glue, and I’d rather own that glue than fight a compatibility layer.

Selenium Grid is the runner-up if you already have it, and sharding is the right answer only for scheduled test-style checks. If you want more background before committing, Playwright vs Puppeteer for AI browser agents covers the tool choice underneath all of this, and the blog index has the rest of my agent ops notes.

Written by Xavier Fok

disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-10-06.

free download
Why did my agent get blocked? A triage checklist

The checks we run, in order, when a browser or phone agent starts failing: network, fingerprint, behaviour, account. Leave your email and we will also tell you when we publish a new field note, a few times a month at most.

from the team behind this site
A Singapore mobile IP for browser agents

Singapore Mobile Proxy runs real mobile IPs on SingTel, StarHub and M1, with sticky sessions so one task keeps one IP. Singapore only: a fit for SEA or location-agnostic work, the wrong tool if you need a US IP.

see plans →
from the team behind this site
A real Android phone for phone-use agents

cloudf.one hosts real Android phones in Singapore on dedicated hardware, each with a persistent Singapore mobile IP. For agents that need an actual device and a stable carrier identity.

get a phone →
read on
More from The Agent Ops Report

Blocks, sessions, retries, traces, cost per task and phone-use agents. Browse all articles →