← all articles

Playwright vs Puppeteer for AI browser agents

i run browser agents out of Singapore, and the library underneath them matters more than the model on top. when an agent fails at 3am, it is almost never because the LLM picked a bad plan. it is because a click landed before the page was ready, a login expired, or a container lost its Chromium. so this comparison is about the boring layer: Playwright from Microsoft and Puppeteer from the Chrome team at Google.

both are free, open source libraries that drive a real browser from code. both can sit behind an LLM loop, a LangGraph node, or an MCP server. the difference is in how much of the operational work each one does for you. Playwright ships with auto-waiting, isolated browser contexts, saved auth state and a trace viewer. Puppeteer is leaner, closer to Chrome, and asks you to build more of that yourself.

my verdicts, short version. for general agent work with logins, retries and debugging, Playwright wins. for a Chrome-only script, a PDF or screenshot service, or a Node codebase that already runs on Puppeteer, Puppeteer is fine and sometimes simpler. for long-running multi-account agents, Playwright again, mostly because of session handling and traces. details below. this is as of October 2026, so check the docs before pinning a version.

TL;DR comparison table

Playwright Puppeteer
Maintainer Microsoft Google Chrome team
Pricing free, Apache 2.0 licence free, Apache 2.0 licence
Languages JavaScript/TypeScript, Python, Java, .NET JavaScript/TypeScript (Node)
Browsers Chromium, Firefox, WebKit Chrome and Chromium, Firefox via WebDriver BiDi
Auto-waiting built into actions and locators manual waits and selectors, newer locator API helps
Session handling contexts, storageState, persistent contexts userDataDir, cookie APIs, browser contexts
Tracing trace viewer with DOM snapshots, network, screenshots Chrome performance traces, screenshots, your own logging
Support GitHub issues, Discord, docs site, active release cadence GitHub issues, docs site, Chrome team releases
Target user teams building agents, testing and scraping own properties across browsers Node developers who want Chrome control with a small API

on pricing: there is no licence fee for either one. the real costs are compute, proxies, and LLM tokens, which i cover below.

Playwright at a glance

Playwright came out of Microsoft in 2020, built in large part by engineers who had previously worked on Puppeteer. it drives Chromium, Firefox and WebKit through one API, and it has official bindings for JavaScript/TypeScript, Python, Java and .NET. the Python binding matters for agent work, because most agent frameworks, LangGraph included, are Python first.

the features that matter to an operator:

  • auto-waiting: actions like click and fill wait for the element to be attached, visible, stable and enabled before they fire. this removes a whole class of flaky failures.
  • locators: role, text and label based locators that re-resolve on each action, so a re-rendered DOM does not leave you holding a stale handle.
  • browser contexts: cheap isolated profiles inside one browser process. one context per user or per task is the normal pattern.
  • storageState: save cookies and localStorage to a JSON file, load it into a new context. the Playwright authentication docs describe the pattern.
  • trace viewer: a recorded timeline with DOM snapshots, network calls, console output and screenshots per action. more on this below.
  • Playwright MCP: Microsoft maintains an MCP server that exposes browser control to LLM clients, which is why a lot of agent tooling defaults to Playwright now.

the downsides are real. the install pulls down its own browser builds, which is heavy in Docker images and awkward behind corporate proxies. i have written up two of those failures already: the SSL certificate error during playwright install behind a proxy and the “executable doesn’t exist” launch failure in Docker. the WebKit build is also not the same as Safari on an iPhone, so do not treat it as a mobile Safari test.

Puppeteer at a glance

Puppeteer was released by the Chrome team in 2017 and is the older project. it is a Node library. it controls Chrome or Chromium, mostly over the Chrome DevTools Protocol, and it has supported Firefox through WebDriver BiDi since 2024. the Puppeteer documentation is short and readable, and the API surface is smaller than Playwright’s.

what it does well:

  • it is close to Chrome. new DevTools capabilities tend to show up in Puppeteer quickly, because the same organisation builds both.
  • the API is small. a new developer can read most of it in an afternoon.
  • it is good at the jobs it was built for: screenshots, PDFs, rendering pages server-side, crawling your own site, and running Chrome extension or DevTools experiments.

what it does less well for agents:

  • Node only for the official library. if your agent stack is Python, you either run a Node sidecar or use an unofficial port, and i would not build a production agent on an unofficial port.
  • more manual waiting. you can get reliable behaviour with careful waitForSelector and network idle logic, but you write and maintain that logic. Puppeteer has added a locator API, which narrows the gap, but the default habits in most Puppeteer code you will find online are still manual waits.
  • session persistence is do-it-yourself beyond the basics.
  • no equivalent of the Playwright trace viewer out of the box.

Head-to-head

Reliability

this is the axis that decides most agent projects. an agent loop multiplies every small flake. if a step has a two percent chance of failing for timing reasons (illustrative, not measured), a twenty step task fails often enough that you will spend your week on retries.

Playwright’s auto-waiting is the main reason it feels steadier. every action checks that the target is actionable before acting, and failures come with a clear message about which check did not pass. in Puppeteer you can reach the same place, but each selector, each click and each navigation needs a deliberate wait, and agents that generate code or tool calls on the fly are bad at remembering to add them.

navigation is the other pain point. both libraries can hang when a page never fires a load event, because of a long-polling request or a hung third party script. i wrote up the Playwright side in the page.goto timeout with no load event ever firing. the same root cause bites Puppeteer, and the same fix applies: wait for a specific element or domcontentloaded instead of full load, and cap the wait.

browser coverage also counts. if your target site renders differently in Firefox or WebKit, Playwright lets you test that with one API. Puppeteer’s Firefox support exists but is newer and narrower in practice. for an agent that only ever needs Chrome, this does not matter.

a caveat on both: neither library fixes a bad agent loop. if your graph spins until it hits a limit, that is an orchestration problem. see how to stop a LangGraph agent from looping until it hits the recursion limit and what actually survives a crash in LangGraph durability modes. the browser library is only one layer.

edge: Playwright.

Session handling

agents that do anything useful usually sit behind a login, and logging in on every run is slow, costs tokens, and trips the site’s own security checks on your own account. you want to log in once, save the state, and reuse it. this is legitimate work on accounts the user has consented to, inside the site’s terms.

Playwright gives you three layers:

  • storageState: dump cookies and localStorage from a context, load into another. simple and portable.
  • persistent contexts: launchPersistentContext with a user data directory keeps the full profile on disk, including things storageState does not capture well.
  • one browser, many contexts: isolate users or tasks without spawning a browser process each time.

Puppeteer gives you:

  • userDataDir on launch, which persists the Chrome profile on disk. this works, and for a single long-lived profile it is perfectly good.
  • cookie APIs to read and set cookies. localStorage and sessionStorage you have to evaluate and restore yourself.
  • browser contexts for isolation, which work but come with less ceremony around saving and restoring a whole auth state.

the practical difference is glue code: with Playwright one JSON file usually does it, with Puppeteer you build the save and restore routine. i have watched injected cookies fail to stick in both, usually because of domain, path or timing mistakes. i covered that in how to make injected storage state cookies actually stick in a browser agent, and the broader pattern is in how to persist browser sessions for AI agents.

one more operator note on networks. some agent work is location dependent, for example checking how your own Singapore store or app renders for local users. in that case the browser’s exit IP matters. both libraries accept a proxy at launch, and Playwright also lets you set it per context, which is handy when different tasks need different exits. i wrote the Playwright setup in how to route a Playwright agent through a mobile proxy. if you need real Singapore mobile IPs with sticky sessions for that kind of own-property testing, Singapore Mobile Proxy sells them. it is Singapore only, so it is no help if you need other countries.

edge: Playwright, clearly on ergonomics. Puppeteer is adequate if you keep one profile on disk.

Tracing and debugging

when a run fails at night, you want to see what the browser saw. this is where the gap is largest.

Playwright’s trace viewer records each action with before and after DOM snapshots, the network log, console messages and screenshots. you open the resulting zip and scrub through the run. the trace viewer docs show how to turn it on per context. for an agent, i keep tracing on for failed runs and throw away the zip for passing ones, because traces get large.

Puppeteer can write a Chrome performance trace, take screenshots, and log network and console events through its event API. that is useful for performance analysis. it is not the same as a replayable, action by action record of what an agent did. to get close you write your own logging wrapper: screenshot after every tool call, dump the URL and a DOM excerpt, store it with the run id. i have done that, it works, and it is a day or two of work that Playwright users skip.

one thing neither library gives you is model level tracing. you still need your own logging of prompts, tool calls and model outputs. the browser trace tells you what the page did, not why the agent chose the click.

edge: Playwright, by a wide margin.

Cost

licence cost is zero for both, so the comparison is about what you pay to run them.

  • compute: both launch a full browser, and a headless Chromium instance eats real memory. Playwright contexts are cheaper than separate browsers, so packing many tasks into one browser process helps. Puppeteer can do similar with its browser contexts, but the save and restore tooling is yours to build.
  • image size: Playwright’s image with three browser engines is big. if you only install Chromium, the gap to Puppeteer shrinks, and you can install only the browser you need.
  • engineering time: this is the biggest number and the hardest to measure, so i will not invent one. Playwright’s built-in waiting, state saving and tracing replace code you would otherwise write and maintain for Puppeteer. on a one-off script that saving is small. on an agent fleet it adds up.
  • LLM tokens: the library does not change your token bill directly, but how you feed the page to the model does. Playwright MCP uses accessibility tree snapshots rather than screenshots by default, which are usually smaller than images.
  • proxies and phones: network cost sits outside both libraries. check current vendor pricing rather than trusting a blog post.

edge: tie on licence, Playwright on total cost of ownership for agents, Puppeteer on minimal footprint for single-purpose Chrome jobs.

Language and ecosystem fit

if your agent lives in Python, Playwright is the natural choice, since the Python package is official and tracks the Node one. if your whole stack is Node, both are fine, and the question becomes whether you want the extra features. if you are in Java or .NET, Playwright is the only one of the two with an official binding.

Puppeteer has more legacy tutorials, which stops mattering when an LLM writes the code, since current models know both APIs. older examples use renamed or removed methods, so pin your version and read the changelog before upgrading.

edge: Playwright for polyglot teams, tie for Node-only teams.

Sandboxing and isolation

a browser agent that can click anything is a risk surface. both libraries launch the browser as a normal process, so isolation is your job: run it in a container, drop privileges, limit network egress, and keep the agent’s credentials scoped. neither library ships a sandbox for the agent itself. my notes on that are in how to sandbox a computer use agent. note that running Chromium as root in Docker often forces --no-sandbox, which weakens the browser’s own protection, so use a non-root user where you can.

edge: tie.

Use-case verdicts

An LLM agent that logs into a user’s accounts and works through tasks

winner: Playwright. this is the common agent shape: load saved auth state, act, recover from a page that changed, log what happened. storageState, auto-waiting and the trace viewer line up with exactly those needs. Puppeteer can do it, and i know people running it well, but they have written the glue that Playwright ships.

A screenshot, PDF or page-render service

winner: Puppeteer. if the job is “give me a URL and get back a PDF or PNG from Chrome”, the small API and Chrome-first design fit. there are no sessions to manage and no multi-step flows to debug. Playwright does this too, but you gain little from its extras.

Cross-browser checks on your own web app

winner: Playwright. it drives Chromium, Firefox and WebKit from one API, and the test runner is built for this. if an agent is helping you triage regressions across engines, Playwright is the practical pick. Puppeteer’s Firefox support is newer and i would not rely on it for parity checks.

A long-running fleet of agents, each with its own profile

winner: Playwright, with a caveat. contexts keep per-agent state isolated without a process per agent, and traces make failures diagnosable after the fact. the caveat is memory: a fleet still needs careful limits, periodic browser restarts and a plan for crashes, and that is true of both libraries. if each agent needs a real Singapore mobile identity, that is a network and device question, not a library question, and whatever you use should stay within each site’s terms and your users’ consent.

Who should pick Playwright

  • you build agents in Python, Java or .NET, or you want one library across languages
  • your agents log in, hold state across runs, or juggle several users
  • you want to debug failures after the fact without writing your own recorder
  • you need Firefox or WebKit as well as Chromium
  • you would rather pay in disk space than in engineering time

Who should pick Puppeteer

  • your stack is Node and already runs Puppeteer in production with no pain
  • you only need Chrome, and the job is narrow: rendering, screenshots, PDFs, simple crawls of your own properties
  • you care about keeping the dependency small and want the API to fit in your head
  • you are happy to write your own waiting, state and logging layers, or the job does not need them

Verdict overall

for AI browser agents, i pick Playwright. the gap comes down to three things that show up at 3am: auto-waiting that cuts timing flakes, storageState and contexts that make sessions easy to save and isolate, and a trace viewer that tells you what actually happened. none of that is glamorous. all of it saves time once you run more than a handful of agents.

Puppeteer is not a bad choice. it is a good, focused library, and for Chrome-only rendering jobs or an existing Node codebase i would not rewrite it. if you are starting an agent project today without that history, start on Playwright, pin the version, keep traces for failed runs, and keep your model logs separate. if you want more on the operational side, the blog index has the rest of the agent ops notes.

Written by Xavier Fok

disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-10-01.

free download
Why did my agent get blocked? A triage checklist

The checks we run, in order, when a browser or phone agent starts failing: network, fingerprint, behaviour, account. Leave your email and we will also tell you when we publish a new field note, a few times a month at most.

from the team behind this site
A Singapore mobile IP for browser agents

Singapore Mobile Proxy runs real mobile IPs on SingTel, StarHub and M1, with sticky sessions so one task keeps one IP. Singapore only: a fit for SEA or location-agnostic work, the wrong tool if you need a US IP.

see plans →
from the team behind this site
A real Android phone for phone-use agents

cloudf.one hosts real Android phones in Singapore on dedicated hardware, each with a persistent Singapore mobile IP. For agents that need an actual device and a stable carrier identity.

get a phone →
read on
More from The Agent Ops Report

Blocks, sessions, retries, traces, cost per task and phone-use agents. Browse all articles →