Idempotency key — so a retry doesn't duplicate side effects.
Permanent-failure path — what happens after retries are exhausted.
This is essentially how durable-execution systems such as Temporal work: external operations are placed in Activities, whose executions have persisted state, timeouts, and retry policies.
The key is to separate orchestration from AI work. Your agent should be a set of durable steps with explicit state, rather than one long function that happens to call an LLM.
Systems such as Temporal are specifically designed for this: workflow execution can resume after crashes, network failures, or infrastructure outages rather than losing the workflow's progress.
and have the downstream system deduplicate that key.
Durable execution systems generally assume activities can execute more than once; Temporal specifically recommends making Activities idempotent because a worker can succeed and crash before reporting completion.
Put AI-specific recovery above mechanical retries
There's an important distinction:
Transient failure
↓
automatic retry
↓
still failing
↓
agent-level recovery
↓
change strategy / ask human / abort
For example, if a tool repeatedly rejects the LLM's arguments, blindly retrying the same call is wasteful.
Instead:
try:
result = await tool_call(args)
except ToolValidationError as e:
return await agent_replan(
original_plan=plan,
failure=str(e),
)
That is different from:
except TimeoutError:
retry()
LangGraph follows a similar distinction: node-level retry policies handle transient failures, while an error handler can take over after retries are exhausted.
The current LangGraph RetryPolicy supports initial delay, exponential backoff, maximum interval, attempt count, jitter, and custom exception filtering.
For an agent, I'd typically have different policies for different nodes:
Temporal's current AI guidance specifically describes this model for long-running agents that call LLMs/tools, wait for humans, and resume after crashes.
A practical recipe
If you're adding this to an existing flaky agent, I'd do it in this order:
Give every agent run a stable ID.
Break the agent into discrete steps.
Persist state after every successful step.
Move external calls behind retryable operations.
Add exponential backoff + jitter.
Classify errors as retryable vs. permanent.
Give every side-effecting operation an idempotency key.
Add per-step timeouts.
After retry exhaustion, either replan, request human intervention, or fail explicitly.
Only then consider retrying the entire workflow.
The result is much more reliable than "retry the agent": completed work stays completed, transient failures get retried locally, and genuinely broken decisions get routed into recovery instead of endlessly replaying the same action.
plan = llm("Decide what to do")
result = call_customer_api(plan)
answer = llm("Summarize the result")
Instead, make the boundaries explicit:
async def workflow(request):
state = await load_state(request.id)
if "plan" not in state:
state["plan"] = await run_llm(
request,
idempotency_key=f"{request.id}:plan",
)
await save_state(request.id, state)
if "api_result" not in state:
state["api_result"] = await call_api(
state["plan"],
idempotency_key=f"{request.id}:api",
)
await save_state(request.id, state)
if "answer" not in state:
state["answer"] = await run_llm(
state["api_result"],
idempotency_key=f"{request.id}:answer",
)
await save_state(request.id, state)
return state["answer"]
Now if the worker dies after api_result is persisted, recovery starts from that checkpoint rather than repeating the whole agent.
Don't retry everything
A useful policy is:
Failure
Retry?
Typical policy
Network timeout
Yes
Exponential backoff + jitter
HTTP 429
Yes
Respect Retry-After
HTTP 500/503
Yes
Exponential backoff
LLM transient error
Yes
2–5 attempts
Invalid tool arguments
Usually no
Fix/replan
Authentication failure
No
Refresh credentials/fail
Bad user input
No
Ask for correction
Deterministic application bug
No
Alert/deploy fix
Temporal explicitly recommends retrying failure-prone Activities rather than blindly retrying the entire Workflow.
Make side effects idempotent
This is probably the most important detail.
Imagine:
charge_card(amount)
The request succeeds, but your worker crashes before receiving the response.
and have the downstream system deduplicate that key.
Durable execution systems generally assume activities can execute more than once; Temporal specifically recommends making Activities idempotent because a worker can succeed and crash before reporting completion.
Put AI-specific recovery above mechanical retries
There's an important distinction:
Transient failure
↓
automatic retry
↓
still failing
↓
agent-level recovery
↓
change strategy / ask human / abort
For example, if a tool repeatedly rejects the LLM's arguments, blindly retrying the same call is wasteful.
Instead:
try:
result = await tool_call(args)
except ToolValidationError as e:
return await agent_replan(
original_plan=plan,
failure=str(e),
)
That is different from:
except TimeoutError:
retry()
LangGraph follows a similar distinction: node-level retry policies handle transient failures, while an error handler can take over after retries are exhausted.
The current LangGraph RetryPolicy supports initial delay, exponential backoff, maximum interval, attempt count, jitter, and custom exception filtering.
For an agent, I'd typically have different policies for different nodes:
Temporal's current AI guidance specifically describes this model for long-running agents that call LLMs/tools, wait for humans, and resume after crashes.
A practical recipe
If you're adding this to an existing flaky agent, I'd do it in this order:
Adding durable execution and retries to a flaky AI agent workflow requires shifting from in-memory state management to an orchestration engine or checkpointer pattern . Because LLM calls time out, external APIs rate-limit, and tools fail mid-execution, you need a system that records state transitions and replays or retries cleanly.
Here are the primary ways to build durability and retries into your agent stack, broken down by architectural approach:
Use a Dedicated Durable Orchestration Engine (e.g., Temporal)
How it works: Engines like Temporal treat your agent's steps as durable workflows and individual tool calls/LLM requests as isolated activities. State and event histories are persisted externally. If a worker crashes or an API returns a 500 error, the platform automatically retries the specific activity using a backoff policy and resumes execution from the exact point of failure.
Implementation path: You can wrap your agent steps using native integrations, such as the Temporal LangGraph Plugin Documentation , which runs graph nodes as retryable activities.
How it works: Frameworks like LangGraph support state persistence through checkpointers . By swapping an in-memory checkpointer for a persistent database (PostgreSQL, Redis , or SQLite ), every super-step of the agent's reasoning loop is saved.
Implementation path: If a node fails, you catch the exception, log the state, and reload the graph from the last committed checkpoint in the database to resume or route to a fallback prompt.
Implement Application-Level Retries with Exponential Backoff and Jitter
How it works: LLM providers (OpenAI, Anthropic ) and vector databases frequently throw rate limits (429) or temporary gateway errors (502, 504). Wrapping individual LLM calls or tool executions in a robust retry decorator (using libraries like tenacity in Python) isolates failures locally.
Configuration tip: Always apply exponential backoff combined with random jitter to prevent "thundering herd" issues when hitting recovering downstream APIs.
Design Idempotent Tools and Safe Side-Effects
How it works: Retries only work safely if repeating an action doesn't duplicate side effects (e.g., sending the same email three times or double-charging a credit card).
Best practice: Design your agent tools to accept idempotency keys, or separate your agent graph into "reasoning/planning" phases (which are pure and re-playable) and "action/mutation" phases (which require explicit transactional guarantees or external state checks before firing).
If you want, tell me:
What framework or stack are you currently using to build your agent (e.g., LangGraph, CrewAI, custom Python loops)?
Are you worried more about LLM rate limits/timeouts or infrastructure/server crashes?
I can dive into community discussions on the r/LocalLLM or LangChain Community subreddits regarding how other developers handle gotchas like prompt-drift during workflow replays, or provide a tailored code snippet for your setup.
You can add durable execution and retries to a flaky AI agent workflow by integrating a dedicated orchestration framework like Temporal, Inngest , or Pydantic Logfire/AI that automatically checkpoints state, replays events, and wraps LLM and tool calls in resilient retry policies.
AI agent loops fail due to transient network drops, rate limits, flaky APIs, and unpredictable model outputs. Moving from basic scripts to production-grade durability requires structural design changes.
Core Strategies for Durable AI Workflows
Use a Durable Workflow Engine : Offload state management to platforms like Temporal or Inngest. These frameworks record every step (LLM generation, tool execution) in an append-only event log so that if a node crashes mid-flight, execution resumes from the last successful checkpoint instead of restarting from scratch.
Isolate Tool Calls as Activities : Treat every external API call, database write, or deterministic calculation as an isolated "activity" or step with its own custom retry and exponential backoff configuration.
Implement Hard Guardrails and Budgets : Never rely solely on the LLM to decide when to stop. Enforce strict orchestrator-level limits on maximum iterations, execution time, and total token expenditure to prevent infinite loops.
Leverage Suspend/Resume for Human-in-the-Loop : Use native workflow pause features when the agent requires user validation or external data. Durable engines can sleep indefinitely without wasting compute resources until a webhook or human response arrives.
Comparison of Popular Durable Execution Options
Tool / Framework
Best For
Key Mechanism
Temporal
Complex, long-running, multi-service agent graphs
Code-as-workflow with deterministic event replay and robust workers
Inngest
Serverless and event-driven architectures
Step functions using standard queues and HTTP triggers
Pydantic AI
Pydantic-native Python agents
Step-by-Step Implementation Approach
Define State Boundaries : Map your agent workflow into distinct, discrete steps (e.g., Parse Prompt → Call LLM → Execute Tool → Evaluate Progress).
Wrap External Calls : Apply explicit retry policies (e.g., max 3 attempts with exponential backoff) specifically targeting HTTP 429 rate limits or 5xx server errors from LLM providers.
Persist Context : Ensure conversation history and intermediate tool outputs are saved externally to a persistent database or state store before moving to the next action.
Add Progress Detectors : Track whether sequential tool iterations are returning identical outputs; if the agent gets stuck in a loop, programmatically terminate or escalate the task.
If you share which language or framework (e.g., Python with LangChain/Pydantic AI, or TypeScript) and infrastructure (serverless vs. dedicated workers) you are using, I can provide a code example for setting up retries.
A flaky AI agent workflow usually fails for the same reasons distributed systems fail: processes die, APIs time out, tools return transient errors, and partial side effects happen. Durable execution adds a workflow layer that persists progress, retries failed steps, and resumes from checkpoints instead of restarting the whole agent.
Have I already processed message_id?
yes -> return previous result
no -> send
Idempotency keys are one of the main techniques for making retries safe.
5. Handle different failure classes differently
Not every error deserves a retry.
Retry
timeout
connection reset
HTTP 429 rate limit
temporary provider outage
Do not retry
invalid prompt
permission denied
malformed tool arguments
policy rejection
Example:
try:
result = call_llm()
except RateLimitError:
retry_later()
except InvalidRequestError:
fail_workflow()
For LLM APIs specifically, rate limits often need special handling (for example, respecting retry-after headers), while invalid requests usually should not consume retry budget.
6. Add heartbeats for long-running tools
Some agent tools take minutes:
browser automation
document processing
code execution
batch jobs
Have them report progress:
activity started
|
├── downloaded files
├── processed 500/1000 records
└── completed
If the worker dies, the workflow knows whether to retry or resume.
Workflow:
generate report
wait for approval
continue
The workflow can sleep for days without holding compute.
Practical stack options
Common approaches:
Temporal Technologies — workflow engine with replay, activities, retries, and long-running execution patterns.
Durable Task frameworks — similar patterns for checkpointing and retries in cloud environments.
Agent frameworks with durable execution integrations — for example, OpenAI Agents SDK integrations with Temporal.
A production-grade agent usually ends up looking like:
User request
|
Durable workflow
|
+-- Plan step (LLM)
|
+-- Tool A (retry + idempotency)
|
+-- Tool B (retry + timeout)
|
+-- Human approval signal
|
+-- Final response
The key design shift is: your agent is no longer a running program; it is a recoverable process whose state and progress are stored outside the worker.
1. Make each meaningful operation a durable activity
Instead of:
def run_agent(prompt):
data = fetch_data()
answer = llm(data)
result = call_api(answer)
return result
structure it conceptually as:
def workflow(input):
data = activity(fetch_data)
answer = activity(call_llm, data)
result = activity(call_api, answer)
return result
The workflow engine persists the fact that fetch_data completed. If the worker crashes while calling the LLM, you don't need to redo the entire workflow.
2. Put retries around activities, not the whole agent
You can implement the orchestration yourself with a database + queue, but once workflows become long-running or have many failure modes, a durable execution system such as Temporal can remove a substantial amount of infrastructure you would otherwise have to build.
The most important design rule is:
The workflow owns state and recovery; activities own external side effects; retries belong to activities; side effects must be idempotent.
If you tell me whether your agent is Python/TypeScript and what framework you're using (LangGraph, OpenAI Agents SDK, custom, etc.), I can show the concrete implementation pattern.
Adding durable execution and robust retries to an AI agent workflow transforms it from a fragile demo into a production-ready system . Because AI agents are stateful, long-running, and prone to external hiccups (rate limits, flaky APIs, or transient LLM provider outages), standard try/catch blocks and fire-and-forget queues fall short.
To achieve durability, you need state checkpointing, automatic state reconstruction , and isolated step-level retries.
1. Choose Your Architecture Approach
You can implement durable execution via dedicated general-purpose engines, agent-native frameworks, or serverless developer primitives.
Dedicated Durable Orchestrators: Engines like Temporal AI Documentation or Inngest treat your entire agent loop as code, automatically persisting every step, variable, and tool call to an event history database. If a server dies mid-execution, another worker picks up precisely where it left off.
Agent-Native Frameworks with Checkpointing: Frameworks like LangGraph use built-in memory savers (like Postgres or Redis checkpointers) to save state after every node transition, allowing you to replay or fork execution history.
Cloud & Serverless Workflows: Platforms like Cloudflare Workflows or AWS Durable Functions allow you to write step-based orchestration directly into serverless functions with native state persistence.
2. Implement Step-by-Step Design Patterns
Separate Reasoning from Side Effects (Activities vs. Workflows)
Never execute non-deterministic LLM calls or external API side-effects (like sending an email or writing to a database) directly inside a raw retry wrapper without recording state. Divide your agent into:
The Orchestrator (The Brain): Decides what to do next.
Activities / Tools (The Hands): Isolated functions that perform the actual web requests, DB writes, or LLM completions. These should be individually retryable.
Configure Intelligent, Budget-Aware Retries
LLM providers and external APIs fail differently. Standard linear retries will drain your token budget or trigger stricter rate limits. Configure your retries with:
Exponential Backoff with Jitter: Increase wait time exponentially (𝑡=b a s e×2 a t t e m p t+j i t t e r ) to prevent thundering herds on rate-limited APIs.
Exception Filtering: Retry transient errors (429 Rate Limit, 500 Internal Server Error, timeout exceptions) but fail fast on deterministic errors (400 Bad Request, invalid JSON tool arguments, authentication failure).
Max Retry Caps: Limit LLM generation retries to 3-5 attempts before falling back to a smaller model or routing to a human-in-the-loop review queue.
Persist State Checkpoints Incrementally
If an agent takes 10 sequential steps and fails on step 9, you don't want to re-run steps 1 through 8 (and re-pay for those input tokens). Ensure your framework checkpoints the conversation history and tool outputs immediately after every discrete node execution.
Handle Human-in-the-Loop (HITL) as a Paused State
If your agent requires user approval before executing a high-risk tool (e.g., executing code or making a financial transaction), use a durable primitive. The workflow should pause indefinitely without consuming active compute resources, waiting safely for an external webhook or UI interaction to wake it back up.
If you want to tailor this further, tell me:
What framework or stack are you currently using to build the agent (e.g., LangGraph, raw Python/Asyncio, Temporal, LlamaIndex)?
Are your failures mostly related to LLM rate limits / timeouts or downstream tool/API failures?
I can provide a concrete code snippet or architecture blueprint for your exact setup.
Adding durable execution and robust retries to a flaky AI agent workflow transforms your system from a fragile script that breaks on a single rate limit (429) or API timeout into a resilient engine that survives infrastructure crashes and transient third-party failures.
Two primary architectural patterns achieve this: using a native framework checkpointer (like LangGraph with PostgreSQL) or pairing your agent logic with an enterprise durable execution engine (like Temporal).
1. Choose Your Orchestration Approach
Framework-Native Checkpointing (e.g., LangGraph + Postgres): Best if your state and agent control loops are self-contained. It saves state transitions to a database after every node execution. If a step fails, you can resume from the last known good state.
Durable Execution Engine (e.g., Temporal Integration): Best for production systems requiring guaranteed execution across distributed workers, complex timeouts, and advanced exponential backoff policies for flaky external APIs (OpenAI, Anthropic, or custom tools).
2. Implementation Steps
Isolate LLM Calls and Tool Executions into Activities/Nodes
Separate your pure cognitive/routing logic from side-effect-heavy or flaky operations (network requests, API calls to LLMs, database writes). In a setup like the Temporal LangGraph Integration , flaky API calls are wrapped as activities so they can fail and retry independently without corrupting the broader workflow memory.
Configure Exponential Backoff Retry Policies
LLM providers frequently throw rate limits (429) and gateway timeouts (504). Configure explicit retry policies with jitter and backoff multipliers rather than standard linear loops:
Initial interval: 1 second
Backoff coefficient: 2.0 (doubling the wait time each retry)
Maximum interval: 60 seconds
Maximum attempts: 5 to 10 for transient network/rate-limit errors; fail fast on 400-level client validation errors.
If you want to dive deeper, let me know:
Which AI orchestration framework (LangGraph, CrewAI, LlamaIndex, or custom Python code) are you currently using?
What kind of flakiness or errors are hitting you the hardest (rate limits, timeouts, or state corruption mid-run)?
I can provide a tailored code snippet or architectural layout for your specific stack.
A system such as Temporal is designed specifically for this: workflows can resume after crashes, network failures, or infrastructure outages rather than disappearing.
retry same operation
vs.
ask the agent to reconsider
vs.
terminate workflow
For example, a 429 from an API might warrant an infrastructure retry, while "tool argument 'customer_id' is invalid" should probably send a structured error back to the agent so it can correct its plan.
7. Add deterministic workflow logic around nondeterministic AI
A useful rule is:
Durable workflow code should be deterministic; nondeterministic/external work should happen in activities/tasks.
So instead of putting arbitrary LLM calls, network requests, and current-time lookups directly into the durable workflow, isolate them behind activity boundaries.
That is one of the reasons workflow engines like Temporal work well for AI agents: the orchestration layer can reliably replay/recover while the actual AI/tool calls happen as externally managed activities.
A practical production recipe
If I were hardening a flaky agent today, I'd implement these in roughly this order:
Give every run a durable workflow ID.
Break the agent into discrete activities such as retrieval, LLM calls, tool calls, and persistence.
Checkpoint after every meaningful state transition.
Add exponential backoff + jitter to transient failures.
Make side-effecting tools idempotent.
Classify errors into retryable, agent-correctable, and fatal.
Put hard limits on steps, retries, tokens, and wall-clock time.
Record every attempt with workflow ID, step ID, tool, model, latency, error, and attempt number.
Add human escalation for workflows that repeatedly fail.
Test recovery explicitly by killing workers halfway through workflows.
For a greenfield implementation, I'd seriously consider Temporal for the durable execution layer; its core promise is crash-resumable execution, including for AI-agent workloads.
If you're already using LangGraph, Python/asyncio, Node, or a queue like Celery, the implementation details differ substantially. I can show a concrete durable/retryable architecture for whichever stack you're using.
Yes. The key is to treat your agent as a durable state machine, not as one long function call.
A robust architecture looks like:
Request
│
▼
Durable workflow
├── Load state
├── Plan
├── Call tool/API ── retryable
├── Run LLM step ── retryable
├── Validate result
├── Persist checkpoint
└── Continue / compensate / fail
Platforms such as Temporal are designed specifically to resume workflows after crashes, network failures, and infrastructure outages rather than restarting from scratch.
1. Separate orchestration from side effects
Make your workflow deterministic and put unreliable operations into activities/tasks:
def agent_workflow(request):
state = load_state(request)
plan = plan_task(request, state)
for step in plan:
result = execute_task(step) # retryable activity
state = update_state(state, result)
save_checkpoint(state)
return state
You generally don't need to persist every transient token-level detail, but you should be able to answer:
"What did the agent believe, what did it attempt, and why did it move to the next state?"
7. Consider a workflow engine once this gets serious
If your workflow needs to survive process/container/VM crashes and resume days later, don't reinvent durable execution with a database plus a bunch of polling workers.
A system such as Temporal provides durable workflow execution and lets activities be retried independently. Its documentation explicitly positions this model for AI-agent workflows as well as traditional business processes.
Retry operations, checkpoint state, make side effects idempotent, and put a bounded state machine around the LLM.
If you're building this in Python/TypeScript, I can also show a concrete implementation using Temporal, LangGraph, or a lightweight Postgres-based approach.
To add durable execution and retries to a flaky AI agent workflow, adopt a durable workflow orchestration engine (such as Temporal, Inngest , or cloud native durable functions) to checkpoint state and manage retries automatically.
AI agents fail often due to rate limits, flaky third-party tool APIs, server timeouts, and unpredictable LLM outputs. Traditional try/catch blocks and ad-hoc loops fail to scale or recover state when infrastructure crashes.
Step-by-Step Implementation Guide
Separate Deterministic Orchestration from Non-Deterministic AI Calls
Keep your workflow definition deterministic (pure code logic for routing, state, and loops).
Isolate LLM generations and tool executions into distinct, isolated "activities" or "steps" that can fail, timeout, and be retried independently.
Integrate a Durable Execution Engine
Choose an orchestration layer designed for state persistence:
Temporal for self-hosted or managed robust code-as-workflow infrastructure.
Inngest for serverless and event-driven architectures.
Deep control over state replays and robust multi-language SDKs
Inngest
Serverless, fast setup, and event-driven apps
Frictionless integration with modern web frameworks and serverless queues
AWS Durable Functions
Serverless architectures natively on AWS
Zero compute cost while execution is suspended for human approval
If you can share which programming language or AI framework (e.g., LangChain, OpenAI Agents SDK, custom scripts) you are using, I can provide a concrete code snippet demonstrating a retry and checkpoint loop.
Modular capabilities like TemporalDurability attached directly to agents
suspend and resume
Set Strict Timeouts
AI models can hang indefinitely waiting for a token stream or response. Apply granular timeouts:
Start-to-close timeout on individual LLM generation steps (e.g., max 30 seconds per prompt turn).
Workflow-level timeout to prevent runaway agent loops from consuming infinite budget if a hallucination loop occurs.
Persist State Incrementally
Ensure message histories, tool outputs, and agent scratchpads are serialized and checkpointed immediately after each step. If using a framework like LangGraph, configure a persistent checkpointer (PostgreSQL or Redis) rather than an in-memory saver so that state survives process restarts.
Handle Non-Deterministic Outputs Safely
LLM completions are non-deterministic, but your workflow orchestrator expects deterministic state tracking. Save the actual returned text/response of the LLM call into your durable event history or database during the first successful attempt. Do not re-execute the LLM call on a workflow replay; use the recorded output instead.
Invalid tool arguments
Usually no
Fix/replan
Authentication error
No
Escalate
Permission denied
No
Escalate
Business-rule rejection
No
Replan/compensate
Configure Granular Retry Policies per Activity
Attach exponential backoff retry policies directly to tool calls and LLM API requests to gracefully handle HTTP 429s (rate limits) or 503s (service unavailability).
Define maximum attempts, initial interval, maximum interval, and backoff coefficients.
Leverage Checkpointing and Event Replay
Rely on the engine's built-in state machine to record every completed step.
If a worker crashes mid-workflow, the runtime replays history up to the failure point and resumes instantly from the exact next step without re-running expensive prior actions.
Implement Suspend/Resume for Human-in-the-Loop Interventions
Use asynchronous callback or sleep primitives provided by the durable runtime to pause execution indefinitely while waiting for user approvals or external webhook inputs without bleeding compute resources.