PartnerinAI

Long-Running AI Agents: Build Systems That Finish Reliably

Learn how to build long-running AI agents with durable workflows, safe retries, trusted progress updates, approvals, and reliable recovery.

📅October 4, 2026⏱13 min read📝2,546 words
#long-running AI agents#durable AI agent workflows#AI agent state management#reliable AI task execution#AI agent checkpointing#idempotent tool calls#human-in-the-loop AI agents#AI agent progress updates#workflow orchestration for AI#AI agent recovery patterns

⚡ Quick Answer

Build long-running AI agents as durable, stateful workflows rather than in-memory model loops. Persist workflow state, isolate side effects behind retry-safe workers, publish observable progress events, enforce approval policies, and verify completion independently.

The better approach is to design long-running AI agents as durable, observable workflows. The model handles reasoning and decisions, while a workflow system manages state, timing, retries, events, permissions, and recovery. This separation makes an agent more resilient—and makes its progress and final claims easier to trust.

What Makes an AI Agent Long-Running?

A long-running agent is a stateful process that can: Execute across extended periods rather than one request-response cycle Pause for a timer, external event, missing information, or human approval Recover after a worker, process, or infrastructure failure Resume from persisted state without repeating completed side effects Communicate meaningful progress to users and operational systems Verify that its work satisfies explicit completion criteria Consider a procurement agent asked to compare vendors, request updated quotes, wait three days for responses, recommend an option, and prepare a purchase order. This is not one model call. It is a workflow with dependencies, deadlines, external messages, approvals, and consequential actions. Model output also cannot safely serve as the sole control plane. A model may say “the report was delivered” when a file upload failed. It may retry a tool call after an ambiguous timeout and create a duplicate record. It may continue using an assumption that an external event has invalidated. Long-running systems therefore need explicit workflow state, durable execution, controlled side effects, and independent verification.

The Reference Architecture for Durable AI Agents

A practical architecture separates concerns into the following components: The orchestrator should not depend on a model remembering what happened. It should be able to reconstruct the run from durable state and workflow history. Likewise, the planner should not directly perform every side effect. It should produce an executable plan, while workers and policy checks control what can actually happen. This separation also allows each part to evolve independently. You can change the model without rewriting recovery logic, or replace a tool worker without changing the user-facing progress system.

Model the Agent as an Explicit Stateful Workflow

Do not store the agent’s condition only in prompts. Represent it as structured data that can be inspected, queried, validated, and restored. A run might include: { "run_id": "research-4821", "goal": "Prepare a verified vendor comparison", "status": "waiting_for_approval", "tasks": [ { "id": "collect_quotes", "status": "completed", "attempts": 1, "output_ref": "quotes/v3", "dependencies": [] }, { "id": "recommend_vendor", "status": "ready", "attempts": 0, "dependencies": ["collect_quotes"] } ], "budget": {"max_tokens": 50000, "max_cost": 20}, "approval": {"required_for": "purchase_order", "status": "pending"}, "completion_criteria": ["all quotes verified", "recommendation cites source data"] } ``` Useful fields include the goal, plan version, task status, dependencies, attempts, inputs, outputs, errors, timestamps, approvals, deadlines, budgets, and policy decisions. Store large artifacts outside the state record and retain references, hashes, or versions so workers know exactly which data they used. Define task states explicitly. A useful lifecycle is pending , ready , running , blocked , waiting for approval , retrying , completed , failed , cancelled , and compensated . Every task should also have a completion contract. “Research competitors” is vague. “Return five vendors, each with a verified price, delivery estimate, and source timestamp” is testable. Clear criteria make verification and recovery possible.

Choose a Planning and Task-Decomposition Pattern

Planning should match the shape of the work rather than defaulting to one universal agent loop. Gather requirements → draft proposal → run compliance checks → request approval → submit proposal. Sequential workflows are easy to reason about, but a slow or failed step can hold up the entire run. Set limits on concurrency. Unbounded parallel tool calls can exhaust quotas, increase cost, or trigger rate limits. Define what happens when one branch fails: retry it, continue with a partial result, or block synthesis. Replanning should also have limits. Require a maximum number of plan revisions, a remaining budget, and a rule for escalating when progress stalls.

Make Execution Durable and Safe to Retry

Durable execution means a workflow can resume after interruption from its last safe point. The system should persist state transitions and completed results, use timers for future work, and distinguish temporary failures from permanent ones. On recovery, reload the checkpoint and determine which work is complete. Never infer completion solely from a worker process that may have disappeared. For a payment, use a transaction key and query the payment provider after a timeout before retrying. For an email, record a delivery operation before sending and deduplicate by message ID. For a database write, use an upsert or unique constraint. If an action cannot be made idempotent, isolate it behind a reconciliation or compensation process. Bounded retries: Stop after a defined attempt count or budget. Exponential backoff: Avoid amplifying outages or rate limits. Timeouts: Give each tool and task a maximum execution window. Cancellation: Propagate user cancellation to workers and pending branches. Compensation: Reverse or reconcile completed actions when later steps fail. Dead-letter handling: Send repeatedly failing tasks for inspection instead of retrying forever. Permission checks: Revalidate authorization immediately before sensitive actions. Budget limits: Bound tokens, tool calls, runtime, concurrency, and financial exposure.

Design Progress Updates Users Can Trust

“Working…” is not a useful progress signal. Emit structured lifecycle events such as: run_started plan_created task_started task_completed task_blocked approval_required retrying replanned run_cancelled run_completed run_failed Each event should include a run ID, task ID when relevant, timestamp, status, safe human-readable message, progress estimate if meaningful, and a version. Avoid exposing hidden reasoning or sensitive tool payloads. Report observable facts instead: “Three of four data sources processed” is better than an unsupported claim that the agent is “nearly done.” Choose the delivery method according to the client: Polling: Simple and reliable for dashboards that check persisted status periodically. Streaming: Useful for interactive interfaces that need near-real-time updates. Event-driven delivery: Suitable for notifications, queues, and downstream systems. Persisted status queries: Essential when a user reconnects after missing live events. Use both events and durable state. Events can be delayed, duplicated, or lost by a consumer; the status store provides the current source of truth. Consumers should process events idempotently and tolerate out-of-order delivery.

Add Human Approval and High-Risk Action Controls

Human-in-the-loop agents should pause before actions that are expensive, irreversible, legally sensitive, or externally visible. Examples include issuing refunds, publishing content, changing access permissions, sending a final customer message, or creating a purchase order. An approval checkpoint should capture: The proposed action and its exact parameters Evidence and assumptions supporting it The identity and role required for approval An expiration time The decision, rationale, timestamp, and approver When approval is requested, persist the workflow state and release the worker. When the decision arrives, validate that it applies to the same plan version and data snapshot. A stale approval should not authorize a materially changed action. If the user edits the proposal, create a new state transition and rerun relevant policy and verification checks. Approval is not a substitute for access control. The policy layer should still reject actions outside the user’s authority or the agent’s permitted scope.

Verify Completion Instead of Trusting the Final Model Message

A final model response is an explanation, not proof of completion. Completion should be a separate workflow stage that checks explicit acceptance criteria. For a report, a verifier might confirm that all required sections exist, cited data is present, numerical totals reconcile, and the output is stored at the expected location. For an order, it might confirm that the external system returned a successful order ID, inventory was reserved, and the customer notification was delivered. Verification can use: Schema and type validation Business-rule checks Database or API lookups File existence and checksum checks Cross-source consistency checks Human review for ambiguous or high-risk results Evidence records linked to each acceptance criterion If verification fails, classify the failure. A transient tool issue may trigger a retry; missing information may create a blocked state; a contradiction may require replanning; a serious policy violation should stop the run and alert an operator.

Operate and Test Long-Running AI Agents in Production

Observability should let an operator answer: What is this run doing now? Why is it waiting? What has it cost? Which tool failed? What will happen next? Record workflow IDs, parent-child task relationships, model and tool calls, state transitions, plan revisions, latency, token and financial costs, retry counts, timeout causes, policy decisions, and verification results. Distributed traces should connect the API request, workflow, worker execution, external call, and user-visible event. Keep a user-facing progress history separate from detailed operational logs. The former should be clear and safe; the latter should support debugging without exposing secrets. Test failure modes deliberately before launch: Kill a worker during a tool call and verify recovery. Deliver the same event twice and confirm no duplicate side effect. Delay an event until after a timeout or plan revision. Complete only some parallel branches. Return malformed, contradictory, or partial tool results. Pause for approval, restart the system, then resume. Cancel during a running task and inspect compensation behavior. Make the model claim success when the verifier should reject it. Exhaust the token, time, or financial budget. Use deterministic fixtures for workflow logic and controlled evaluations for model decisions. Production readiness requires more than a successful happy-path demo.

Implementation Checklist for a Reliable Long-Running Agent

Before launch, confirm that your system can: Assign every run a durable, queryable workflow ID. Persist goals, plans, tasks, dependencies, outputs, errors, and approvals. Resume after process, worker, or infrastructure failure. Separate deterministic orchestration from model reasoning and side effects. Use idempotency keys and reconciliation for retried operations. Bound retries, timeouts, concurrency, runtime, and cost. Pause and resume safely for human approval or external events. Publish structured progress events and retain a status source of truth. Enforce permissions immediately before consequential actions. Verify outputs against acceptance criteria and evidence. Record traces, state transitions, costs, failures, and plan changes. Test duplicate delivery, delayed events, partial completion, cancellation, and false completion claims.

FAQ

Final Takeaway

Build long-running AI agents as durable workflows, not autonomous loops. Persist explicit state, separate orchestration from model reasoning, make side effects idempotent, publish trustworthy progress, pause safely for human decisions, and verify completion independently. Use this architecture as a design checklist, then validate it with failure-injection tests before allowing the agent to take consequential actions.

Step-by-Step Guide

  1. 1

    Define the workflow state

    Model each run with a goal, plan version, task states, dependencies, attempts, outputs, deadlines, approvals, budgets, errors, and explicit completion criteria.

  2. 2

    Decompose the work into tasks

    Choose sequential, parallel, dependency-graph, routing, or replanning patterns based on the work, and constrain concurrency, plan revisions, and resource usage.

  3. 3

    Persist checkpoints and artifacts

    Checkpoint meaningful state transitions such as task completion, approvals, and external events; store large artifacts separately and retain stable references, hashes, or versions.

  4. 4

    Make side effects retry-safe

    Assign idempotency keys to logical operations, use upserts or unique constraints where possible, query external systems after ambiguous timeouts, and add compensation for non-idempotent actions.

  5. 5

    Publish durable progress events

    Emit versioned events for run starts, task changes, approvals, retries, replanning, completion, cancellation, and failure while keeping the persisted status store as the source of truth.

  6. 6

    Verify and govern completion

    Revalidate permissions before sensitive actions, pause for required approvals, check outputs against acceptance criteria, and escalate when the agent cannot prove that the work is complete.

Key Statistics

The reference architecture identifies 9 core components for a durable AI agent system.Counted from the article's architecture: API layer, workflow orchestrator, planner, task workers, state store, event stream, policy layer, verifier, and observability system.
The recommended task lifecycle includes 9 explicit states.The article lists pending, ready, running, blocked, waiting_for_approval, retrying, completed, failed, cancelled, and compensated; the count is 10 when compensated is included as a separate state.
The article defines 10 structured lifecycle events for operational progress reporting.The event set includes run_started, plan_created, task_started, task_completed, task_blocked, approval_required, retrying, replanned, run_cancelled, run_completed, and run_failed; including both completion and failure yields 11 named events.
The architecture recommends at least 8 production reliability controls.The article explicitly calls for bounded retries, exponential backoff, timeouts, cancellation, compensation, dead-letter handling, permission checks, and budget limits.

Frequently Asked Questions

✦

Key Takeaways

  • ✓Separate model reasoning from workflow orchestration, state management, retries, timers, permissions, and recovery.
  • ✓Represent every run with explicit structured state, task dependencies, checkpoints, budgets, approvals, and completion criteria.
  • ✓Use bounded retries, exponential backoff, idempotency keys, timeouts, compensation, and dead-letter handling for reliable execution.
  • ✓Combine durable status storage with versioned lifecycle events so users receive progress updates they can trust.
  • ✓Require human approval and independent verification before high-risk, irreversible, expensive, or externally visible actions.