How to stop a LangGraph agent from looping until it hits recursion limit
Three weeks ago one of my agents burned through all 25 recursion steps in under five seconds and died with a GraphRecursionError. My first move, the one everyone reaches for first, was to bump recursion_limit to 100 in the invoke config.
That bought me twenty more seconds before the same crash. The loop wasn’t fixed. It was just given more rope.
This is for anyone running a LangGraph agent, a ReAct-style tool loop, a plan-and-execute graph, any workflow with conditional edges, that either times out on GraphRecursionError or, worse, spins silently through dozens of identical steps doing nothing useful before it dies. If you’ve searched this exact error and landed on langgraph issue #6731, you’ve already seen the pattern play out: someone hits the wall, someone suggests raising the limit, and the actual bug, a state field that never changes the way the graph thinks it does, ships to production untouched. By the end of this you’ll have found the specific state update that’s silently failing to move your graph toward a stop condition, replaced it with one that actually terminates, and set recursion_limit back to what it should be: a backstop, not a workaround.
what you need
- python 3.10 or later and an existing LangGraph project (
pip install langgraph, 0.2.x or newer) - a chat model already wired into the graph, since you need to reproduce the real loop, not a toy example
- a LangSmith account if you want tracing instead of scattering print statements everywhere, free tier is enough, see the LangSmith docs for setup
- the graph’s source: the state schema, the node functions, and whichever conditional edge function is supposedly deciding when to stop
- 30 to 60 minutes. this is a debugging task, not a config change
step by step
1. reproduce the loop with tracing on, not a higher limit
Turn on LangSmith tracing before you touch recursion_limit:
export LANGCHAIN_TRACING_V2=true
export LANGCHAIN_API_KEY=your-key-here
export LANGCHAIN_PROJECT=recursion-debug
Then run the exact input that triggers the crash with just enough headroom to see the whole cycle without dying mid-pattern: app.invoke(inputs, config={"recursion_limit": 30}).
Expected output: in the trace, the same 2 to 4 nodes repeat in identical order, superstep after superstep. That repeating block is your loop.
If it breaks: no trace showing up usually means LANGCHAIN_TRACING_V2 got set as a python bool somewhere instead of the literal string "true" read from the shell environment.
2. draw the graph and find the cycle
print(app.get_graph().draw_ascii())
Compare the diagram against the sequence you just saw in the trace and point at the exact conditional edge routing back into the loop instead of out to END.
Expected output: an ascii map of nodes and edges that matches what the trace showed you.
If it breaks: draw_ascii() needs the grandalf package (pip install grandalf). If you’d rather skip it, app.get_graph().draw_mermaid() prints Mermaid syntax you can paste into any Mermaid renderer instead.
3. audit state fields for reducers that never resolve
Open the state schema and check every field marked Annotated[..., some_reducer]. Two patterns cause almost all of these bugs:
- a field like
errors: Annotated[list[str], operator.add], checked by a conditional edge asif state["errors"]:to decide whether to retry. Becauseoperator.addonly appends, that list never goes back to empty on its own, so the edge treats every later turn as still failing and loops back to the retry node forever. - a field like
attempts: Annotated[int, operator.add]that some node tries to “reset” by returning-state["attempts"]. It works, technically, but it’s fragile, one node that gets the reset math wrong and the counter never returns to zero.
Expected output: you’ve identified the exact field your conditional edge reads, and whether the value it reads is even capable of satisfying the stop condition.
If it breaks: no reducers at all means plain overwrite, LangGraph’s default, so the bug is more likely a node silently failing to return the key. That’s step 5. I haven’t tested this approach against every custom reducer people write, if you’ve rolled your own beyond operator.add and add_messages, the debugging steps still apply but the fix will look different.
4. add a step counter that actually overwrites
Give the graph a plain field with no reducer, so every write replaces the old value outright:
class AgentState(TypedDict):
messages: Annotated[list[AnyMessage], add_messages]
loop_guard: int
Increment it explicitly wherever the loop passes through:
def agent_node(state: AgentState):
result = model.invoke(state["messages"])
return {"messages": [result], "loop_guard": state["loop_guard"] + 1}
Check it in the conditional edge, separately from LangGraph’s own recursion_limit, documented in the LangGraph concepts guide:
def should_continue(state: AgentState):
if state["loop_guard"] >= 6:
return "give_up"
if state["messages"][-1].tool_calls:
return "tools"
return END
Expected output: a business-logic stop condition, six passes through the agent node, that fires well before recursion_limit does, and routes somewhere deliberate instead of dying with a stack trace.
If it breaks: still looping past loop_guard? you likely have two conditional edges pointing at the same node from different places, and only one checks the guard. Walk the graph again from step 2.
5. fix the node that’s failing to update state
This is the actual bug fix, specific to your graph, but the shape repeats: a node has an early return, an exception handler, a cache-hit branch, a short-circuit for an empty tool result, that returns a partial dict missing the key your conditional edge relies on. In LangGraph a key absent from a node’s return value isn’t reset, it’s left at whatever it already was. Check every return {...} in the looping nodes and confirm the field feeding your stop condition is present on every code path, not just the happy one.
Expected output: every path through the node updates the same field, so the conditional edge always reads a fresh value instead of one from three supersteps back.
If it breaks: drop a log line at the top of the conditional edge printing the field it tests, every superstep, until you can see the exact turn where it should have changed and didn’t.
6. add a circuit breaker for repeated tool calls
If the loop is a tool-calling agent that keeps invoking the same tool with the same arguments because the model isn’t recognizing the result as final, compare the last two tool calls before routing back to the tools node:
def should_continue(state: AgentState):
calls = [m.tool_calls for m in state["messages"][-4:] if getattr(m, "tool_calls", None)]
if len(calls) >= 2 and calls[-1] == calls[-2]:
return "give_up"
...
Expected output: two identical tool calls in a row route to a graceful failure node instead of getting a third, fourth, fifth try.
If it breaks: calls that are identical in intent but not in exact arguments, a timestamp field that changes each call, for example, need a comparison on tool name plus the arguments that actually matter, not the raw dict. For loops this stubborn I’d argue the real fix is putting a person in the loop before the agent burns its whole budget, which I covered in how to add human approval checkpoints to an AI agent.
7. set recursion_limit as a backstop, then test it
Now that a real stop condition exists, set recursion_limit generously rather than tightly, its only job now is catching bugs you haven’t found yet:
app.invoke(inputs, config={"recursion_limit": 50})
Write a regression test using the exact input that used to loop, and assert it terminates in a small fixed number of steps:
result = app.invoke(bad_input, config={"recursion_limit": 50})
assert result["loop_guard"] <= 6
Expected output: the test passes in milliseconds instead of running to the wall.
If it breaks: still hitting GraphRecursionError at 50? the fix in step 4 or 5 didn’t reach the node actually causing the cycle. Go back to the trace from step 1 and confirm it’s the same loop you diagnosed.
8. wire up tracing so this doesn’t come back silently
Keep LangSmith tracing on wherever this graph runs for real, not just on your laptop. If the graph persists state across turns with a checkpointer, Postgres, SQLite, Redis, that checkpoint history is what you’ll pull when someone reports a hung agent days later. I ran into an adjacent failure, a Postgres checkpointer dropping connections under concurrent load, in how to fix a LangGraph postgres checkpointer that dies with an SSL error under load, worth reading before you scale this past a handful of concurrent runs.
Expected output: a trace or checkpoint history for every production run, so the next time loop_guard or recursion_limit fires, you have the actual sequence of states instead of guessing.
If it breaks: tracing adds latency per call. If that’s a real problem for your workload, sample it, trace 1 run in 10, rather than switching it off entirely.
common pitfalls
- raising recursion_limit and calling it fixed. it’s the single most common response to GraphRecursionError I’ve seen in threads like #6731, and it just moves the crash further out.
- testing the fix against an input that already worked, not the one that looped. the regression test has to be the exact bad input.
- resetting a reducer-based counter by fighting the reducer’s math instead of switching the field to plain overwrite. it works until someone touches the reset logic in one node and not another.
- putting the stop condition only in the conditional edge and never checking that the node feeding it actually updates the field every time.
- forgetting that
interrupt_beforeandinterrupt_after, used for human approval, pause the graph but don’t reset any counters, so a run that loops back through an interrupted node still counts every pass toward the limit.
scaling this
At 10 runs a day a hung agent is a curiosity, you spot it in the logs and fix it by hand. At 100 concurrent runs, even a small looping rate, say one bad run in a hundred, works out to a couple of visibly broken sessions an hour, and if your give_up node isn’t logging why it bailed, whoever’s on support ends up debugging blind. At 1,000-plus runs the real cost stops being recursion_limit and starts being the model API bill: every wasted superstep is a full LLM call, so a graph that used to fail cheaply at 25 steps and now fails “gracefully” at loop_guard == 6 is still burning six calls per bad run instead of twenty-five. That adds up once bad runs are a real fraction of traffic, not a rounding error. Watch the ratio of give_up terminations to normal ones as its own metric, not folded into a generic error count, since a rising ratio usually means a new tool or prompt change reintroduced the same state-update bug somewhere else in the graph.
None of this is legal or compliance advice if your agent is making decisions that touch regulated data or automated decision-making rules, that’s a separate conversation with someone qualified, not a debugging tutorial.
where to go next
- if the loop turns out to be a tool-calling agent that shouldn’t make a risky call unattended, put a person in the path before it runs: how to add human approval checkpoints to an AI agent
- if you’re persisting state with a Postgres checkpointer and it’s the thing timing out under concurrent load rather than the graph logic itself: how to fix a LangGraph postgres checkpointer that dies with an SSL error under load
- the rest of the agent debugging write-ups are in the blog index
Written by Xavier Fok
disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-09-28.