Self-hosted LangGraph checkpointing vs LangGraph Cloud for long tool calls
In this article Vendor A is self-hosted LangGraph checkpointing: the open source library with a Postgres or SQLite checkpointer that you run on your own box. Vendor B is LangGraph Cloud, the managed deployment from LangChain that runs your graph and its persistence for you. The question I keep getting is narrow. what happens to a tool call that runs past a few minutes, and can I trust the framework not to run it twice?
That question comes from langchain-ai/langgraph issue 7417, which reports a silent re-execution of a long tool call. I’ll answer it directly here. I have not reproduced that bug, and as of October 2026 I can’t tell you from the outside which versions or deployment types it affects, so read the thread for the current status. What I can tell you is how checkpointing is documented to behave, why a re-run of a long call is a design property you have to plan for on both sides, and how to make a re-run harmless.
My short verdicts. For a long call that is not safe to repeat (a payment, an email, a ticket), neither product saves you, and the fix is idempotency in your own tool. For a long call with a few minutes of browser work behind a proxy, I’d lean self-hosted because you control the process, the timeouts and the egress IP. For a team that does not want to run Postgres and workers, LangGraph Cloud is the saner choice, provided you still write the tool defensively. I’m Xavier Fok, I run agents from Singapore, and this is written from that operator seat.
One note on the structure. This site’s comparison template was built for proxy vendors, so the head-to-head axes below (IP pool, rotation, geo, success rate, speed, price per GB, sessions, concurrency) are mapped onto this pair. Where an axis is really about the egress IP of your tool call, I say so, because that is where long tool calls and proxies actually meet.
TL;DR comparison table
| Self-hosted checkpointing (Vendor A) | LangGraph Cloud (Vendor B) | |
|---|---|---|
| Pricing | library is free, you pay for your servers and Postgres | managed, billed by LangChain, check the LangChain pricing page for current terms |
| Where the long call runs | in your own worker process | in the managed deployment’s workers |
| Persistence | PostgresSaver or SqliteSaver you operate | managed Postgres, handled for you |
| Re-run risk on interruption | documented, make tools idempotent | same graph semantics, same advice applies |
| Egress IP | yours, fixed, can sit behind a sticky proxy | decided by the platform, check current docs |
| Support | community, GitHub issues, your own on-call | vendor support tiers, check your plan |
| Target user | teams with ops skills and strict control needs | teams that want to ship without running infrastructure |
I’d rather leave a cell vague than invent a number, so where the table says “check”, I mean it. Pricing and support terms for hosted products change.
Vendor A at a glance
Self-hosted checkpointing means you install LangGraph, pick a checkpointer, and compile your graph with it. In production that is usually the Postgres checkpointer (langgraph-checkpoint-postgres). After each step the graph state is written as a checkpoint under a thread id. If the process dies, you resume the thread and the graph continues from the last saved state. The LangGraph persistence docs describe checkpoints, threads and how state is replayed.
What you own in this setup:
- the worker process that actually runs the tool call
- the Postgres instance and its backups
- the timeouts at every layer: tool, HTTP client, reverse proxy, container
- the egress IP, since the call leaves from your machine or whatever proxy you wire in
That last point matters more than it sounds. If your tool drives a browser, you decide the network path, and I’ve written up how to do that in how to route a Playwright agent through a mobile proxy.
The cost is operational. When a worker is OOM-killed at minute six of a seven minute tool call, nobody pages anyone except you.
Vendor B at a glance
LangGraph Cloud is the managed route. You push your graph, and the platform runs it, stores checkpoints, and exposes an API for threads and runs. You don’t operate the database or the task queue. That is a real saving, and for a small team it is probably the right default.
What you give up is visibility and control of the layers underneath. You don’t pick the worker’s network path, and you don’t choose the process lifecycle. When a long call is interrupted, you are relying on the platform’s retry and resume behaviour as it currently ships. I’m not going to describe internals I can’t verify. Check the current LangChain deployment docs for run timeouts, retry policy and how interrupted runs are recovered, and date what you read, because this is exactly the sort of behaviour that changes between releases.
Pricing is on the vendor’s page. I won’t quote per-run or per-node figures from memory, because I’d be guessing and they move.
head-to-head
IP pool size
For a pure LangGraph comparison there is no pool. The honest reframe: the egress IP of a long tool call is one IP, and it matters when the call is retried.
With self-hosted, that IP is whatever your worker uses, and you can pin it. With Cloud, you take whatever the platform provides, and I’d check the current docs on egress and whether you can put a proxy in the path from your tool code. If your tool calls a site that ties a login or a cart to an IP, a retry from a different IP is a different visitor from that site’s point of view. Winner on control: self-hosted.
rotation control
Rotation is the thing you want off for a long call. A session that starts on one IP and finishes on another can break a login mid-flow, and a silent re-execution makes that worse because the second run may come from a fresh IP with no cookies.
Self-hosted lets you choose a sticky session in your own proxy config. On Cloud you control rotation only to the extent your tool code can route its own traffic. In both cases, persist the browser state outside the graph. I covered the mechanics in how to persist browser sessions for AI agents and the common failure in how to make injected storage state cookies actually stick in a browser agent. Winner: self-hosted, by control, not by magic.
geo coverage
LangGraph has no geo in the proxy sense. Where your tool call exits is where your worker runs, or where your proxy sits.
Self-hosted: you can run the worker in Singapore, or anywhere you rent a box. Cloud: the region options are set by the platform, so check current docs. If a target site serves a local version based on IP, run the worker or the proxy in the right country. This is a data-location question, not a quality one, and it is not legal advice either way. Winner: self-hosted for flexibility, Cloud for not having to think about it.
connection success rate
I have no benchmark comparing success rates between the two and I won’t invent one. What I can offer is a different measure that I do care about: how often does a long call finish exactly once?
For both, that depends mostly on your tool. A call that is idempotent finishes “effectively once” even if it runs twice. A call that is not idempotent can finish twice, and issue 7417 is a report of exactly that shape of problem. The LangGraph durable execution docs say that on resume, execution restarts from a starting point and that side effects should be wrapped so they are not repeated, and that operations should be idempotent. Read that as the baseline contract on both sides. Winner: tie, because the contract is the same and the safe pattern is yours to write.
speed
Self-hosted writes a checkpoint to a database you probably run close to the worker, so the overhead is small and you can measure it. Cloud adds a network hop and whatever the platform does around a run.
There is a second speed factor that matters for long calls: durability mode. LangGraph lets you choose how eagerly checkpoints are written (the options are named sync, async and exit). Eager writes cost a little time and survive crashes better. I wrote about the trade in LangGraph durability, sync vs async, and what actually survives a crash. If you pick exit on a long call, a crash mid-call can lose more than you expect. Winner: self-hosted for measurability, but I have no numbers to put on it.
pricing per GB
Neither is priced per GB. The relevant cost for self-hosted is a server plus Postgres, and for Cloud it’s the vendor’s own billing model as listed on their pricing page as of October 2026.
The cost I’d actually track is the cost of a duplicate: a long tool call that runs twice burns twice the compute, twice the LLM tokens around it, and if you route through a metered proxy, twice the bandwidth. A re-run of a five minute browser call can easily cost more than the checkpointing itself. Winner: depends on your volume. For low volume, Cloud’s convenience often costs less than your hours.
session persistence
This is where the two differ in a way that is easy to miss. The graph checkpoint persists graph state: messages, variables, which node is next. It does not persist your browser process, your open socket or your half-finished HTTP upload. Those die with the worker on both sides.
So persistence has two layers. The graph layer is handled by the checkpointer, self-hosted or managed. The tool layer is yours: save cookies and storage state to disk or a store, and key the work by a stable id so a retry can pick up or detect “already done”. If your agent loops instead of resuming, see how to stop a LangGraph agent from looping until it hits the recursion limit. Winner: tie on graph state, self-hosted on tool state because the files live on a machine you control.
concurrent connections
Self-hosted concurrency is whatever your worker pool and Postgres connection limit allow. Long calls hold a worker for minutes, so size the pool for the slowest call, not the average one. A common failure is exhausting Postgres connections because every waiting run holds one.
On Cloud, concurrency limits come from the plan and the platform, so read them for your tier. Both can handle many threads, but a long call pins capacity either way. Winner: Cloud for not needing to tune it, self-hosted for having no plan ceiling.
use-case verdicts
Four cases that cover most of what I see.
Long call with a side effect (send, pay, book)
Winner: neither, and I’d say that plainly. Whichever you choose, the tool must carry an idempotency key, generated before the call and stored in graph state. If the downstream API accepts the key, a replay is harmless. If it doesn’t, check “did this already happen” before acting. Given issue 7417, I’d treat a duplicate run as something that can happen, not something that can’t. Tie.
Long browser call behind a proxy
Winner: self-hosted. You control the worker, the timeouts and the exit IP, so retries can come from the same sticky session. If the site expects a stable address and you’re doing legitimate work on your own account or with a consenting user’s account, a sticky mobile IP is worth having. I run Singapore Mobile Proxy, which sells real Singapore mobile IPs with sticky sessions. It is Singapore only, so it only helps if your target sees Singapore traffic as normal. For choosing between mobile and residential, my residential vs mobile proxies for browser agents article goes through the trade-offs.
Small team, no ops, mostly short calls with one slow one
Winner: LangGraph Cloud. You get persistence without running Postgres, and a single slow call can be wrapped defensively in your tool. The saving in setup time is real. Just don’t confuse managed persistence with exactly-once execution.
Regulated or data-location-sensitive workload
Winner: self-hosted. You decide where state is stored and which region the worker runs in. This is not legal advice, and you should check your own obligations, but it is simpler to answer “where does this data live” when you run the database.
who should pick Vendor A
Pick self-hosted checkpointing if:
- you already run Postgres and have someone on call
- your long calls are browser or network heavy and you need a pinned egress IP
- you need to set timeouts at every layer yourself
- you want to inspect checkpoints directly with SQL when something odd happens, which is useful when chasing a duplicate run
- you have data-location requirements
The honest downside is that every crash is your crash. You’ll own migrations, connection limits and backups.
who should pick Vendor B
Pick LangGraph Cloud if:
- you want the graph running and persisted this week, not next month
- your team has no appetite for operating Postgres and workers
- most tool calls are short and the long ones can be made idempotent
- you’re fine with the platform’s limits on region, egress and run time, once you have read them as of the date you sign up
The honest downside is less visibility. If a run behaves oddly, you debug through the vendor’s tooling and your support tier rather than your own logs. For a bug like the one in issue 7417, that can mean waiting on a fix rather than patching a worker yourself.
verdict overall
Winner: it depends, and the dependence is on one question. Can your long tool call run twice safely?
If yes, either product works, and I’d pick Cloud for a small team and self-hosted for anyone who needs control of the network path. If no, fix the tool first. Write an idempotency key into state before the call, record completion after it, and check that record at the start of any retry. Wrap the side effect in its own step, as the durable execution docs advise, so a resume doesn’t replay it. Do that and the question in issue 7417 matters much less, because a silent re-run becomes a no-op.
What I would not do is trust either product to guarantee exactly-once execution of an arbitrary tool call that takes minutes. Process crashes, timeouts, and retries make that guarantee a property of your code. For anything that drives a real browser, also read how to sandbox a computer use agent and the note on prompt injection for browser agents, since a long call that browses the open web is also a long exposure window. More of my operator write-ups are on the blog.
Written by Xavier Fok
disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-10-04.