← all articles

Why a local browser agent times out 30 seconds into startup on a host

the report shows up in the same shape every few weeks. an agent runs perfectly on a laptop, someone containerises it, pushes it to a hosting platform, and the first request dies with a timeout almost exactly 30 seconds in. the question in browser-use issue 2941 is this class of problem, and it is the one I get asked most by people moving agents off their own machines. I have not reproduced that reporter’s exact setup, so I will not pretend to diagnose their stack. what I can do is lay out the mechanism that produces a 30 second wall on almost every platform I have run browsers on, and give you a way to budget for it.

the number is not a coincidence. playwright’s browser launch timeout defaults to 30000 ms, and many agent frameworks sit on top of playwright or a similar launch path with the same default. so a “30 seconds into startup” error usually means chromium was asked to start, did not report ready inside its window, and the library gave up. the interesting question is why chromium takes 4 to 10 times longer on a hosting platform than on your laptop, and why it is often fine on the second attempt.

the stakes are practical. a launch timeout on a laptop costs you a retry. on a platform it costs you a failed deploy health check, a restart loop, a billed cold start per attempt, and sometimes a user-facing job that vanishes. if your agent is a long-running worker, a bad startup budget also compounds with idle limits, because the platform will happily freeze or stop your container between jobs and make you pay the launch cost again. this article walks through the cold start and idle behaviour of typical platforms, what it does to a browser launch event, and how I budget for it.

background and prior art

chromium is not a small process. a headless launch spawns a browser process, a GPU or software-rendering process, a network service, a zygote and renderer processes per tab, and it touches a profile directory, shared memory, and a lot of shared libraries. on your laptop most of that is warm in the page cache, you have several cores, and your disk is an NVMe drive. on a hosting platform you usually have a fraction of a vCPU, a cold root filesystem pulled over a network, a small memory limit, and a small /dev/shm. the same binary does a very different amount of waiting.

hosting platforms also apply their own clocks on top. heroku documents a 60 second boot timeout (error R10) for web processes that fail to bind to their port. google cloud run lets you set minimum instances precisely because scale-to-zero means the next request pays for a cold container start. several free or hobby tiers (render’s free web services are the one I see most) spin an idle service down after a period of inactivity and take a noticeable amount of time to bring it back. none of these numbers are about browsers. they are generic platform clocks, and a browser launch event lands inside them with very little spare room. fees, limits and behaviour on all of these change, so treat any figure I give as “as of October 2026” and check the current docs before you size anything.

the core mechanism

think of the failure as a stack of budgets, each of which can be eaten before your launch call even starts.

the clocks that are running

when a request or a deploy triggers your agent on a cold platform, several timers can be running at once:

  • platform start window: how long the platform lets the container take to become healthy (heroku’s 60 seconds is one example, cloud run startup probes are another)
  • request timeout: how long the platform waits for a response on the request that triggered the cold start
  • library launch timeout: playwright’s 30000 ms default for the browser to start, plus separate navigation and action timeouts
  • your own orchestration timeout: a queue visibility timeout, a client http timeout, a langgraph node timeout, whatever wraps the job

these are nested, not sequential. a cold start spends from the outer budgets before your code runs a single line. if the platform needs 12 seconds to pull the image and start the container, and your app imports its dependencies for 6 more, your browser launch begins with the outer clocks already well into their windows. the 30 second launch timeout then starts fresh, which is why the visible error says 30 seconds even though the whole event has been going for longer.

what chromium is waiting on

inside the launch window, chromium is typically waiting on one of a small number of things:

  • cpu: a platform that gives you a fraction of a vCPU, or throttles cpu outside of request handling, makes process startup, library loading and the first renderer slow. cloud run’s cpu allocation modes are a clear documented example: by default cpu is only allocated while a request is being processed, so work you start outside a request (like a background warm-up) can be starved.
  • disk: the first read of a large layer on a lazily-pulled root filesystem is slow. chromium plus its fonts and libraries is a lot of small reads.
  • memory: a chromium launch plus one page can use several hundred MB. if the platform’s limit is small, the kernel’s oom killer or the platform’s memory enforcement kills a child process and the launch hangs or fails in a confusing way.
  • /dev/shm: docker gives containers a 64 MB shared memory mount by default, and chromium uses shared memory heavily. playwright’s docker docs recommend --ipc=host locally for exactly this reason, and on platforms where you cannot set it you need to work around it, typically with --disable-dev-shm-usage so chromium writes to /tmp instead.
  • sandbox setup: the chromium sandbox needs user namespaces or a setuid helper. in a locked-down container it either fails fast with a clear error or, in some configurations, stalls.
  • dns and network: if your launch path or your first page load needs name resolution through a slow platform resolver, that time is spent inside a window you thought was about the browser.

none of these show up on a laptop, which is why “works locally” tells you almost nothing about the platform.

why it is intermittent

the single most confusing part of this failure is that it is often not deterministic. the first launch after a deploy or a scale-from-zero times out. you retry, it works in 6 seconds. that is the page cache and the image layer cache doing their job. the second launch reads from warm memory, the platform has finished pulling layers, and the cpu burst credit (on platforms that have one) has not run out yet. so a naive retry hides the problem in testing and reveals it in production, where the first request after any idle period is a cold one.

idle limits make this recurring rather than one-off. if the platform stops or freezes your container after some minutes of no traffic, every quiet period resets the cache state and you pay the cold launch again. an agent that is called a few times an hour on a free or scale-to-zero tier will hit the cold path a large share of the time.

the budget

here is how I write the budget down before I deploy. I treat it as a sum of worst-case stages, and I require the sum to fit inside the smallest outer clock with margin.

cold_total = image_pull + container_start + app_import + browser_launch + first_navigation

must be < min(platform_start_window, request_timeout, orchestrator_timeout)

then I split the work so that the stages that are cold-start-only (image pull, container start, imports, browser launch) are separated from the stages that are per-job (navigation, agent steps). the point is that the outer clock for a per-job request should never include the cold-start-only stages. you do that by moving the browser launch out of the request path, which the next section shows with numbers.

worked examples

the figures below are budget sketches with round numbers I chose to show the arithmetic. they are not benchmarks and I am not claiming they are what any specific platform delivers. measure your own image on your own tier and replace the numbers.

example 1: launch inside the request on a scale-to-zero service

the setup: an http endpoint receives a job, launches chromium, runs the agent, returns the result. the platform has a 60 second boot window (think heroku-style) and the playwright default 30000 ms launch timeout.

# naive: launch happens inside the first request
@app.post("/run")
async def run(job: Job):
    async with async_playwright() as p:
        browser = await p.chromium.launch()   # default timeout 30000 ms
        page = await browser.new_page()
        ...

the sketch for a cold path, with made-up round numbers:

  • image pull and container start: 15 s
  • python imports (agent framework, llm client, playwright): 6 s
  • chromium launch on a throttled fraction of a vCPU with a cold disk: 25 s
  • first navigation: 4 s

that adds to 50 s. it fits a 60 s outer window only if nothing is slower than assumed, and the launch alone is close to its own 30 s limit. when the disk is colder than expected and launch takes 31 s, you get the exact error from the issue: a launch timeout roughly 30 seconds into the browser stage. the fix is not to raise one number. it is to stop letting one request carry all four stages.

example 2: launch at process start with a readiness gate

the same service, restructured. the browser launches once when the worker boots, and the platform only routes jobs after a readiness check passes.

import asyncio
from contextlib import asynccontextmanager
from fastapi import FastAPI
from playwright.async_api import async_playwright

state = {"browser": None, "ready": False}

@asynccontextmanager
async def lifespan(app: FastAPI):
    pw = await async_playwright().start()
    state["browser"] = await pw.chromium.launch(
        timeout=120_000,  # cold start budget, not the 30 s default
        args=["--disable-dev-shm-usage"],
    )
    # prove the browser can actually render before saying ready
    page = await state["browser"].new_page()
    await page.goto("about:blank")
    await page.close()
    state["ready"] = True
    yield
    await state["browser"].close()
    await pw.stop()

app = FastAPI(lifespan=lifespan)

@app.get("/healthz")
async def healthz():
    return {"ready": state["ready"]}

now the cold-start-only stages (container start, imports, browser launch) sit inside the platform’s start window, where a longer budget is normal, and the per-job request only pays for navigation and agent steps. in the same made-up numbers, the start window sees 15 + 6 + 25 = 46 s of a 60 s allowance, and a job request sees maybe 4 s of navigation. the launch timeout is raised deliberately to 120000 ms because it is now a startup budget, not a request budget. two caveats: with a platform that has a hard start window, 46 of 60 is thin, so you also want to shrink the image and the import time. and if the platform only starts your container in response to a request (true scale-to-zero), the request that triggers the cold start still waits for all of it. which brings us to the third case.

example 3: scale-to-zero with a warm instance

the same worker on a platform with a minimum-instances setting. with cloud run’s min instances set to 1, one container stays provisioned, so the next request does not wait for image pull and container start. you pay for that idle instance (check the current pricing page for the rate in your region, as of October 2026 I would not quote a figure from memory), and in exchange the cold path mostly disappears for the first job after a quiet period.

the sketch: if you set min instances to 1 and also keep the browser launched in the lifespan hook from example 2, a job arriving after an hour of silence sees roughly the per-job cost only. if instead you only set min instances but still launch chromium per request, you removed 15 s of container start but kept the 25 s browser launch in the request path. I have seen people declare victory after the first change and then meet the same timeout on a busier day when two requests race to launch two browsers inside one small instance.

a related option on cloud run is startup cpu boost, which temporarily gives more cpu during container start. it helps imports and the first launch, and it is a setting worth reading about in the current docs rather than assuming. it does not help if your browser launches lazily after startup finishes.

edge cases and failure modes

the idle freeze after a warm start

you did everything right, the browser is launched at boot, and jobs work. then you leave it alone for 20 minutes and the next job fails. on platforms that throttle cpu outside request handling, or that freeze the container between requests, the browser process can be alive on paper but its connection is dead or stale. the counter-strategy is to treat the browser handle as disposable. before each job, run a cheap liveness check (open a blank page, evaluate 1+1), and if it fails, relaunch inside a generous timeout and log that you did. do not trust that “launched at boot” means “usable now”.

two jobs, one small instance

concurrent requests on an instance sized for one browser will each try to hold a chromium plus a page. memory climbs, the oom killer takes a child process, and the symptom looks like a launch or navigation timeout rather than a memory problem. counter-strategy: cap concurrency per instance at what your memory limit supports with headroom, use a semaphore in the worker, and scale out with more instances instead of more tabs. I covered the fleet-level version of this in the best ways to orchestrate many playwright browser sessions for a fleet of agents, and the single-instance rule is the same: count memory per session, not per request.

/dev/shm and sandbox flags that hide the real error

people respond to a launch timeout by adding every chromium flag they find in a forum thread. some of these (--disable-dev-shm-usage, and --no-sandbox only where your container genuinely cannot create the sandbox and you control what pages the agent loads) address real constraints. others are cargo cult and make debugging harder. counter-strategy: turn on playwright’s debug logging for the launch (DEBUG=pw:browser) in the platform environment, read what chromium actually printed, and add one flag at a time. note that --no-sandbox weakens isolation, which matters more once an agent is browsing pages it does not control. see prompt injection for browser agents explained for why I do not casually remove isolation on a box that loads arbitrary content.

the retry that doubles the bill

the instinct is to wrap launch in a retry. with a 30 second timeout and three retries you can spend 90 seconds failing, hold a platform worker the whole time, and exceed an outer request timeout so the client retries too. now you have several half-started browsers fighting over the same cpu. counter-strategy: one launch attempt with a realistic startup timeout, a distinct and loud log line when it fails, and let the platform’s restart policy handle the retry at the container level where the cache is shared. if you must retry in-process, kill the previous browser first and cap total elapsed time, not just attempt count.

the orchestrator timeout you forgot

your langgraph node, queue consumer or client sdk has its own timeout, and it often defaults to something short. a cold launch that fits the platform window can still blow a 30 second client http timeout on the request that triggered it. counter-strategy: make the trigger request return quickly with a job id, run the agent asynchronously, and poll or receive a callback. if the agent is a graph with long tool calls, the persistence question matters too, and self-hosted langgraph checkpointing vs langgraph cloud for long tool calls goes through what survives a restart. if a restart makes the agent loop instead of resume, how to stop a langgraph agent from looping until it hits the recursion limit is the companion read.

the egress ip changes when you move

this one is not a timeout, but it shows up in the same migration. on your laptop the agent uses your home or office connection. on a platform it egresses from a datacenter range in whatever region you picked. for work on your own properties or with a consenting user’s account, this can change what pages serve, or trigger a site’s own login protections. if your legitimate task needs a Singapore consumer mobile connection, say testing how your own region-specific site renders for local users, a service like Singapore Mobile Proxy sells real Singapore mobile IPs with sticky sessions. it is Singapore-only, so it only helps if Singapore is the location you need, and it is a networking choice, not a fix for launch timeouts. keep the two problems separate when you debug.

what we learned in production

the main lesson from running browser workers on managed platforms is that the 30 second timeout is a symptom of putting a startup cost in a request path. every time I have seen it, the permanent fix was structural: launch once at boot, gate readiness on a real page render, size the instance for the browser instead of for the python process, and give startup its own generous budget while keeping per-job budgets tight. raising the launch timeout alone only moves the failure to the next outer clock. I also stopped trusting local timing. I now measure cold launch on the target platform after a real idle period, write that number in the repo next to the deploy config, and re-measure when the base image or tier changes.

the second lesson is about being honest on cost. keeping a warm instance, or a larger memory tier, costs money every hour whether jobs arrive or not. for a low-volume agent that runs a few times a day, a scale-to-zero tier with an asynchronous trigger and a patient client can be the right answer, and you accept that the first job after quiet is slow. for anything a user is waiting on, I pay for the warm instance. I choose per workload, and I do not claim either choice is free. your model choice affects how long the steps themselves take once the browser is up, which is its own budget, covered in which llms actually hold up running a browser-use agent in production. if you are still deciding between drivers, playwright vs puppeteer for ai browser agents covers the tradeoffs, and the blog index has the rest of the operator notes.

a short checklist I keep next to deploy configs:

  • measure cold browser launch on the real platform after a real idle period
  • set the launch timeout deliberately for startup, and keep request timeouts separate
  • launch at boot and gate readiness on a real render, not just an open port
  • cap concurrency by memory per browser session
  • liveness check the browser before each job and relaunch with a log line if it fails
  • return a job id fast from the trigger request and run the agent asynchronously
  • recheck platform limits and pricing every quarter, since they change

references and further reading

Written by Xavier Fok

disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-10-09.

free download
Why did my agent get blocked? A triage checklist

The checks we run, in order, when a browser or phone agent starts failing: network, fingerprint, behaviour, account. Leave your email and we will also tell you when we publish a new field note, a few times a month at most.

from the team behind this site
A Singapore mobile IP for browser agents

Singapore Mobile Proxy runs real mobile IPs on SingTel, StarHub and M1, with sticky sessions so one task keeps one IP. Singapore only: a fit for SEA or location-agnostic work, the wrong tool if you need a US IP.

see plans →
from the team behind this site
A real Android phone for phone-use agents

cloudf.one hosts real Android phones in Singapore on dedicated hardware, each with a persistent Singapore mobile IP. For agents that need an actual device and a stable carrier identity.

get a phone →
read on
More from The Agent Ops Report

Blocks, sessions, retries, traces, cost per task and phone-use agents. Browse all articles →