← all articles

LangGraph durability sync vs async: what actually survives a crash

i run agents that drive browsers and phones for hours at a time, and the question i get most after a worker dies is “where did it resume from, and did it do the last thing twice?” LangGraph has a durability argument that decides the answer, and the two modes people argue about are sync and async. the question in langgraph issue 8039 is, as i read it, about exactly this: which one you can trust after a crash. i’ll answer it directly here.

short version. sync writes the checkpoint for a step before the next step starts, so after a hard kill you lose at most the step that was in flight. async writes the checkpoint in the background while the next step is already running, so after a hard kill you can lose the last completed step and re-run it. neither mode gives you exactly-once side effects. that part is on you, and i spend a whole section on it below.

my verdicts by use case: pick sync for agents that send messages, spend money, or write to systems you do not own. pick async for read-heavy research agents where a replayed step costs a few cents and nothing else. for a long browser session with logins, i lean sync and pair it with real session persistence, covered in how to persist browser sessions for ai agents. one honest note on method: i did not run a benchmark suite for this article and i will not invent numbers. what follows is the documented behaviour as of September 2026, plus reasoning about the mechanism, plus a small kill test you can run yourself.

TL;DR comparison table

durability=sync durability=async
what is written checkpoint for step N is persisted before step N+1 begins checkpoint for step N is persisted while step N+1 runs
worst case after a hard kill the in-flight step is lost and re-runs the last completed step may also be lost and re-run
latency per step adds the checkpointer write to the critical path write is off the critical path
pricing free, it is a parameter in the open source library. you pay for your database same. slightly less database wait, same storage
support github issues and the langchain community channels, no sla on the oss library same
target user agents with side effects, human-in-the-loop, long unattended runs fast, mostly read-only graphs where replay is cheap
default no yes, as of the current docs

there is a third mode, exit, which only persists when the graph finishes or errors out cleanly. i leave it out of the head-to-head because a crash means you have nothing from that run to resume from.

sync mode at a glance

sync is the conservative option. after each superstep finishes, the runtime writes the checkpoint and waits for the write to complete before it schedules the next step. you set it per call:

graph.invoke(inputs, config, durability="sync")

the LangGraph durable execution docs describe it as the mode with the highest durability, at some performance cost. that matches what you would expect from the mechanism. the checkpointer write sits between two steps, so every step pays for a round trip to sqlite or postgres.

what you get in return is a clean rule for recovery. if the process is killed, the newest checkpoint in your store corresponds to the last step that fully completed. resume from the same thread_id and the graph picks up at the step that was in flight. that one step runs again from the start. anything it did before the kill happened twice from the outside world’s point of view unless you made it idempotent.

async mode at a glance

async is the default in current releases per the docs. the runtime hands the checkpoint write to a background task and starts the next step immediately. the LangGraph persistence docs cover the checkpoint model itself, and the durable execution page states the trade: better performance, small risk that the process dies before the checkpoint lands.

that risk is the whole story. under a graceful shutdown, or an exception that the runtime catches, the pending write gets flushed and you never notice. under kill -9, an oom kill, a spot instance reclaim, or a pod eviction with no grace period, the write can be sitting in memory when the process vanishes. the store then holds the checkpoint from step N-1 while the world has already seen step N and maybe some of step N+1.

head-to-head

the template for these comparisons normally scores proxy vendors on pool size and rotation. that does not fit here, so i score the two modes on the axes that decide what survives a crash. same eight-axis idea, different axes.

checkpoint guarantee

with sync, the guarantee is one sentence: every step that returned has a durable checkpoint before anything else happens. with async, the guarantee is weaker: every step that returned will have a durable checkpoint eventually, provided the process lives long enough to flush it.

so the direct answer to “what survives a crash” is this. in sync, all completed steps survive. in async, all completed steps survive except possibly the most recent one, and possibly more than one if your checkpointer is slow enough that writes queue up behind each other. i have not measured how deep that queue gets on any particular database, and it will depend on your latency and payload size. so treat “one step” as the typical case, not a bound.

winner: sync, and it is not close on this axis.

replay window

the replay window is how much work re-runs after resume. in sync it is the in-flight step. in async it is the in-flight step plus whatever completed steps had not been flushed.

for a research agent where a step is an llm call and a search, replaying two steps costs tokens and a few seconds. for an agent whose step is “submit the form” or “post the reply”, replaying two steps is a duplicate submission. same code, very different consequences, and the mode choice decides which one you are exposed to.

winner: sync for side-effecting graphs, tie for pure ones.

side effects and duplicates

this is where most people get confused, so i’ll be blunt. sync does not make side effects exactly-once. consider a node that calls an external api and then returns. the checkpoint is written after the node returns. if the process dies after the api call succeeded but before the node returned, the checkpoint does not include that step, resume re-runs it, and the api gets called again. that is true in both modes.

what sync changes is the number of steps exposed to that window. it does not close the window. what closes it is an idempotency key on the outbound call, a check-before-write against the target system, or wrapping the effect in a LangGraph task so its result is saved and replayed rather than re-executed. the durable execution docs recommend that pattern for non-deterministic and side-effecting work, and i agree with them. if you take one thing from this article, take that.

  • put an idempotency key on every outbound write, derived from thread id and step, not from a timestamp
  • check the target system for the effect before doing it, when the target lets you
  • keep the side effect in its own node so the replay boundary is obvious
  • log the key and the outcome, so you can tell a replay from a first run

winner: neither. both need idempotent effects. sync just exposes fewer steps.

latency and throughput

async wins here by construction. the write is off the critical path, so the step time is the step’s own work. sync adds the write time to every step boundary.

how much that costs depends on your checkpointer and where it lives. sqlite on local disk is one story. a postgres instance in another region is another. if you run postgres, the PostgreSQL documentation on asynchronous commit is worth a read, because there are two durability layers stacked here: LangGraph’s mode above, and the database’s own synchronous_commit setting below. a sync LangGraph write to a postgres that has synchronous_commit off can still lose a recently acknowledged commit if the database server crashes. the two settings are independent, and people tune one and forget the other.

for agents where a step takes seconds because it waits on a model or a browser, the checkpoint write is a small fraction of step time and sync is cheap. for graphs with many fast steps, such as a tight tool loop, it adds up. if your graph looks like that, also read how to stop a langgraph agent from looping until it hits the recursion limit, because a runaway loop under sync writes a checkpoint on every lap.

winner: async.

database load and cost

both modes write the same checkpoints, so storage growth is the same. sync holds a connection open for the write while the graph waits. async overlaps the write with the next step, which can mean more concurrent connections at peak, because writes from several threads overlap with other work.

pricing is not a differentiator. the durability mode is a parameter in the open source library, so it costs nothing on its own. your bill is the database. if you use LangChain’s hosted deployment products, check their current pricing page directly, since i do not quote plan prices from memory and they change.

one operational note that bites both modes equally. long-lived postgres connections behind a load balancer or with tight ssl settings will drop, and the checkpointer will raise. i covered the specific failure in how to fix a langgraph postgres checkpointer that dies with an ssl error under load. a dropped connection under async is worse in one specific way: the error surfaces on a background write, later than the step that caused it, so your logs show a failure that seems unrelated to the node that was running.

winner: tie, with a small edge to sync for easier debugging.

resume behaviour

resume works the same way in both modes. you invoke the graph again with the same thread_id and a None input, and the runtime loads the latest checkpoint for that thread and continues. the difference is only which checkpoint is the latest.

one behaviour worth knowing is pending writes. when several nodes run in the same superstep and one fails, the runtime keeps the successful nodes’ writes so that resume does not re-run them. that helps with partial failures inside a step. it does not help with a hard kill of the whole process, because the pending writes are stored through the same checkpointer and can be lost in async mode by the same mechanism.

winner: sync, because “latest checkpoint” is more likely to mean “last completed step”.

concurrent threads

if you run many threads in one process, async gives better aggregate throughput because no thread waits on the store between steps. sync serialises each thread’s own steps against the store but does not block other threads.

the failure mode differs, though. one worker running fifty threads under async that gets oom-killed can leave fifty threads each one step behind their true position. under sync, each thread is at most one in-flight step behind. if fifty threads each replay a step that sent a message, you have fifty duplicates at once. that is the scenario i would design against.

winner: async on throughput, sync on blast radius after a crash.

a kill test you can run

i said i would not invent results, so here is the test instead. it takes ten minutes and tells you what your own setup does.

  • build a three-node graph where each node appends a line to a file and sleeps two seconds
  • run it with your real checkpointer, not an in-memory one, once with durability="sync" and once with durability="async"
  • kill the process with kill -9 (or taskkill /F on windows) during node three
  • resume with the same thread_id and compare the file against the checkpoint history from graph.get_state_history(config)
  • repeat twenty times with the kill at slightly different moments, since the async window is a race and one run proves nothing

what you are looking for is the number of lines written by nodes the checkpoint history does not know about. under sync you should see at most the in-flight node. if you see more, your database is acknowledging writes it has not made durable, and the fix is in postgres, not in LangGraph. if you see the same under async with a fast local sqlite, that is the race behaving as documented. i would rather you run it than trust my summary.

use-case verdicts

unattended agent that writes to external systems

sends emails, files tickets, posts replies, submits forms on properties you own or with the account holder’s consent. winner: sync. the extra write time is noise compared with an llm call, and the smaller replay window means fewer steps to protect with idempotency keys. still add the keys.

research and retrieval agent

reads pages, calls search apis, summarises, writes a report at the end. winner: async. a replayed step costs some tokens. speed matters more than the last step of history, and a crash just means you re-run a step that had no consequences.

long browser or phone session agent

this is my daily case. the agent holds a logged-in browser or a phone for hours, and a crash mid-session is normal. winner: sync, with a caveat that the checkpoint only saves graph state, not the browser. cookies, local storage, and device state need their own persistence, or resume lands you in a logged-out browser at the right node. two posts cover that side: how to make injected storage state cookies actually stick in a browser agent and the session persistence one linked earlier. if your agent needs a phone identity that survives restarts, cloudf.one rents real Android phones in Singapore on dedicated hardware, each with a persistent Singapore mobile IP, controlled from the browser. it is Singapore-only, so it only helps if Singapore is where your agent needs to appear from. cloudf.one

human-in-the-loop approvals

the graph pauses on an interrupt and waits hours or days for a person. winner: sync. an interrupt is a state you must not lose, and you do not want the approval prompt re-issued because a flush was missed. the throughput argument for async does not apply because nothing is running while it waits.

high-volume batch of short graphs

thousands of small, idempotent graphs run in a queue. winner: async, or even exit if you retry whole jobs from scratch and never resume mid-graph.

who should pick sync

  • you run agents that touch things you cannot undo, such as sending, paying, deleting, publishing
  • your workers run on spot or preemptible hardware, or in containers that get killed without a grace period
  • you use interrupts for human approval and cannot lose the paused state
  • a step takes seconds or longer anyway, so the write is a small share of step time
  • you want the simplest possible answer to “where did it resume from” when something goes wrong at 3am

who should pick async

  • your graph is mostly reads and a replayed step is cheap
  • you have many fast steps and profiling shows the checkpoint write is a real share of wall time
  • your workers shut down gracefully and you handle sigterm by flushing before exit
  • your effects are already idempotent, so a replay is harmless by design
  • you are running high volumes of short graphs and retrying whole jobs is acceptable

verdict overall

the honest verdict is “it depends”, and the dependence is narrower than the docs make it sound. the mode decides how many completed steps you can lose in a hard crash: one in-flight step for sync, that plus the unflushed tail for async. it does not decide whether side effects happen twice. only idempotent design does that.

if you cannot say for certain that a replayed step is harmless, run sync. it costs you a write per step and buys you a smaller, easier to reason about replay window. if you can say it with certainty, async is a sensible default and it is the default for a reason. either way, run the kill test above against your real database once, because the database’s own commit settings can quietly undo the guarantee you thought you picked. this reflects the documented behaviour as of September 2026, and the library changes fast, so check the durable execution page before you rely on any detail here. more posts on running agents in production are on the blog index.

Written by Xavier Fok

disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-09-29.

free download
Why did my agent get blocked? A triage checklist

The checks we run, in order, when a browser or phone agent starts failing: network, fingerprint, behaviour, account. Leave your email and we will also tell you when we publish a new field note, a few times a month at most.

from the team behind this site
A Singapore mobile IP for browser agents

Singapore Mobile Proxy runs real mobile IPs on SingTel, StarHub and M1, with sticky sessions so one task keeps one IP. Singapore only: a fit for SEA or location-agnostic work, the wrong tool if you need a US IP.

see plans →
from the team behind this site
A real Android phone for phone-use agents

cloudf.one hosts real Android phones in Singapore on dedicated hardware, each with a persistent Singapore mobile IP. For agents that need an actual device and a stable carrier identity.

get a phone →
read on
More from The Agent Ops Report

Blocks, sessions, retries, traces, cost per task and phone-use agents. Browse all articles →