Boundlayer
Durable ExecutionWorkflow OrchestrationTemporalAWS Step FunctionsBusiness Automation

Durable Execution for Business Workflows: Temporal, AWS Step Functions, and Reliable Automation

A practical guide to durable workflow orchestration with Temporal and AWS Step Functions: retries, idempotency, timers, human approvals, compensation, versioning, and AI agents.

21 min read

By BoundLayer Engineering Team

BoundLayer is a senior engineering partner for SaaS, fintech, AI automation, cloud infrastructure, legacy modernization, Web3, IoT, GPU computing, and data systems.

Most business automations are easy to demonstrate and surprisingly hard to operate.

A service receives an order, calls a payment provider, creates a shipment, updates a CRM, waits for an approval, and sends a notification. The happy path may fit in one function. Production adds timeouts, duplicate webhooks, partial failures, rate limits, human delays, deployments, and processes that remain active for days or months.

At that point, the problem is no longer simply how to call several APIs. The problem is how to preserve progress and business intent while every component involved can fail.

Durable execution is an architectural approach for solving that problem. A durable workflow records enough execution history to recover after a process, worker, container, or region-level interruption. The application can continue from a known point instead of restarting the entire process or relying on operators to reconstruct state from logs.

This guide explains where durable execution belongs, how Temporal and AWS Step Functions differ, why queues alone are not a workflow engine, and which design decisions determine whether an automation is genuinely reliable.

Why ordinary application code breaks down

Consider a customer onboarding workflow:

  1. Validate the application.
  2. Run identity and compliance checks.
  3. Ask an analyst to review ambiguous results.
  4. Provision an account in several systems.
  5. Generate agreements.
  6. Wait for an electronic signature.
  7. Activate billing and notify the customer.

This process crosses several reliability boundaries. External calls can time out after completing successfully. A review can take three days. A webhook may arrive twice or before the waiting worker starts. A new application version may deploy while thousands of workflows are in flight.

Teams commonly begin with a request handler, add a background queue, introduce status columns, schedule a retry job, and create an operations dashboard. Each addition is reasonable, but together they form an implicit workflow engine whose semantics are spread across application code, database rows, cron jobs, and tribal knowledge.

Typical symptoms include:

  • records stuck in processing with no clear owner;
  • retries that repeat payments, emails, or provisioning calls;
  • workflows that cannot resume after a deployment;
  • timers represented by frequent database scans;
  • manual approvals handled through one-off scripts;
  • no reliable answer to “what is this case waiting for?”;
  • compensating actions that exist only in an incident runbook.

A queue improves delivery between components. It does not, by itself, model the complete lifecycle of a multi-step business process.

What durable execution means

A durable execution platform persists workflow progress independently of the worker currently executing code. When work stops because of a crash, deployment, network interruption, or long timer, another worker can reconstruct the workflow and continue.

The exact implementation varies by platform. Temporal records workflow events and reconstructs workflow state by replaying deterministic workflow code. AWS Step Functions persists state transitions in a managed state machine. BPMN engines persist process tokens and variables against an explicit process model.

The common properties are more important than the mechanism:

  • Durable state: process progress is not held only in memory.
  • Reliable timers: a workflow can sleep for hours or months without occupying a thread.
  • Controlled retries: backoff, attempt limits, and non-retryable failures are explicit.
  • External interaction: signals, callbacks, or task tokens can resume waiting work.
  • Operational visibility: operators can see current state, history, and failure details.
  • Recovery: workers can restart without abandoning the business process.

Durable execution does not make external side effects exactly once. A workflow engine cannot atomically control a bank API, an email provider, and your database. Activities still need idempotency, reconciliation, and clear ownership of business state.

A production reference architecture

A robust workflow system normally separates six responsibilities.

1. Entry and validation

An API, event consumer, scheduler, or administrator starts the workflow. The entry point validates identity, tenant context, authorization, and an idempotency key before starting a new execution.

The workflow ID should usually derive from a stable business identifier, such as tenant/order/version. Starting the same operation twice can then return the existing execution or fail predictably instead of creating duplicate work.

2. Workflow orchestration

The orchestrator owns control flow: sequence, parallel branches, retries, deadlines, timers, signals, and compensation. It should express business decisions clearly without embedding network clients or large data transformations.

3. Activities or tasks

Activities perform side effects: call an API, write to a database, generate a document, run a model, or publish a message. They are the boundary where timeouts, retry behavior, credentials, rate limiting, and idempotency must be deliberate.

4. Business systems of record

Workflow history is not a replacement for the domain database. Orders, applications, invoices, and customer accounts still need authoritative business records with constraints and audit data.

The workflow tracks how work progresses. The domain model tracks what the business currently believes to be true.

5. Event and integration layer

Queues, streams, webhooks, and a transactional outbox connect the workflow to other systems. For reliable database-to-event publication, use the outbox pattern or change data capture instead of trying to commit a database transaction and publish a message independently.

6. Operations and observability

Engineers need technical telemetry, while operations teams need business-level search and intervention. The platform should support both without forcing people to inspect raw worker logs.

Orchestration versus choreography

Event-driven choreography works well when services react independently to facts such as OrderPaid or CustomerVerified. It reduces central coordination and can keep bounded contexts decoupled.

It becomes difficult when the process has a clear end-to-end owner, long waits, conditional branches, compensation, or strict deadlines. Understanding the current state then requires reconstructing events across multiple systems.

Durable orchestration is usually the better choice when:

  • a single business outcome spans several services;
  • progress must be queryable from one place;
  • steps must execute in a controlled order;
  • the process waits for people or external callbacks;
  • failures require coordinated compensation;
  • deadlines or service-level commitments matter.

This is not an all-or-nothing decision. A workflow can orchestrate critical steps while publishing domain events for loosely coupled downstream consumers. The important point is to make ownership explicit.

Keep workflow logic and activities separate

In code-first engines such as Temporal, workflow code is replayed to rebuild state. It therefore must be deterministic: given the same recorded history, it must make the same decisions.

Do not perform these operations directly in deterministic workflow code:

  • read the current clock without a workflow-safe API;
  • generate unrecorded random values;
  • make network or database calls;
  • read mutable configuration directly;
  • run an LLM and assume it will return the same output;
  • depend on unordered iteration whose order may change.

Put non-deterministic work in activities. The engine records the activity result, and replay consumes that recorded result rather than repeating the side effect.

This separation also improves testing. Workflow tests can cover state transitions with mocked activity results, while activity tests can focus on API contracts, persistence, and failure behavior.

Retries, timeouts, and heartbeats

“Retry three times” is not a reliability policy. A useful policy reflects the failure mode of each operation.

Distinguish timeout types

Define the boundaries that matter:

  • Schedule-to-start: how long a task may wait for a worker.
  • Start-to-close: how long one attempt may run.
  • Schedule-to-close: the total time allowed across retries.
  • Heartbeat timeout: how long a long-running task may remain silent.

A five-second payment authorization and a four-hour data export should not share the same settings.

Classify failures

Retry transient network errors, throttling, and temporary dependency failures with exponential backoff and jitter. Do not retry invalid input, failed authorization, unsupported state transitions, or a definitive business rejection.

Persist a machine-readable failure category. String matching against error messages is too fragile for production policy.

Heartbeat long activities

Long activities should report progress. A heartbeat tells the engine the worker is alive and can carry resumable checkpoint data, such as the last processed page or object key.

Without checkpoints, retrying a four-hour job may discard four hours of successful work. With checkpoints, a replacement worker can continue near the failure point.

Idempotency is still mandatory

AWS Step Functions Standard provides exactly-once workflow execution semantics, and Temporal provides durable activity scheduling. Neither guarantee means an external side effect happens exactly once.

Suppose a payment provider processes a charge but the response is lost. From the workflow's perspective, the activity failed. Retrying without an idempotency key may charge the customer again.

Every side-effecting activity should answer four questions:

  1. What stable key identifies this business operation?
  2. Where is the result of the first successful attempt stored?
  3. What happens when the dependency receives the same key again?
  4. How do we reconcile an ambiguous result?

Good patterns include:

  • provider-supported idempotency keys;
  • a database table with a unique operation key and stored result;
  • compare-and-set state transitions;
  • deduplication at the consumer boundary;
  • read-before-write reconciliation for dependencies without idempotency support.

The activity should return the same logical result for repeated execution. “The engine retries safely” is only true when the activity contract makes retries safe.

Human approvals and external callbacks

Business workflows often pause for an analyst, customer, supplier, or regulator. Do not keep a worker process open or poll every few seconds.

Instead, persist a durable wait and expose a controlled signal or callback path:

  1. The workflow creates an approval task with a stable correlation ID.
  2. The application displays the task to an authorized user.
  3. The user decision is validated and recorded in the domain audit log.
  4. The application signals the workflow.
  5. The workflow validates the current phase and continues.

Add a durable timer for escalation or expiry. If the decision does not arrive within the business deadline, the workflow can notify an owner, request more information, cancel the case, or move it to an exception queue.

Signals can arrive more than once or at an unexpected phase. Treat them like public API calls: authenticate them, enforce tenant boundaries, validate payloads, and make handlers idempotent. This is especially important in multi-tenant SaaS systems.

Compensation instead of distributed rollback

Long-running workflows cannot rely on one database transaction. Once a shipment is booked or a third-party account is created, rollback means performing a new business action.

A saga records forward actions and their compensations:

Forward actionPossible compensation
Reserve inventoryRelease reservation
Authorize paymentVoid authorization
Create cloud resourceDelete or quarantine resource
Book shipmentCancel shipment
Activate subscriptionSchedule termination or credit

Compensation is not always a perfect inverse. An email cannot be unsent, a settled transfer may require a refund, and deletion may violate retention rules. Model the real business remediation rather than pretending the system can return to its original state.

Register compensation before or atomically with the side effect wherever possible. If a worker crashes immediately after a successful action, the workflow must still know how to clean it up.

Versioning long-running workflows

Normal application code assumes the new deployment replaces the old one. Durable workflows may run across dozens of deployments, so code evolution becomes part of the data model.

Unsafe changes can cause replay to make a different decision from the recorded history. Examples include reordering branches, changing a timer path, or introducing a new activity unconditionally into an old execution.

Use the platform's supported versioning strategy:

  • patch or version gates for compatible code paths;
  • worker versioning or build IDs for routing executions;
  • immutable state-machine versions and aliases;
  • continue-as-new for workflows with very long histories;
  • migration only when its operational value justifies the risk.

Before deployment, replay representative production histories against the new workflow code. This catches non-determinism and compatibility mistakes before workers encounter them in production.

Observability must expose business state

Infrastructure metrics alone do not answer the questions operators ask:

  • How many applications are waiting for review?
  • Which orders will miss their fulfillment deadline?
  • Why is this customer onboarding blocked?
  • How many activities are retrying against a particular provider?
  • Which compensation steps require manual intervention?

Capture workflow identifiers, business identifiers, tenant IDs, current phase, attempt count, deadlines, and failure categories as searchable attributes or projections. Keep sensitive payloads out of unprotected metadata.

Instrument workers and activities with traces, metrics, and structured logs. Propagate correlation context across the API, workflow, activity, queue, and external-call boundaries. Our OpenTelemetry production guide covers sampling, cardinality, Collector design, and cost control for this layer.

Alert on symptoms that imply customer impact:

  • workflow age beyond its expected duration;
  • retry storms or exhausted retries;
  • task queue backlog and worker saturation;
  • timer or callback deadline violations;
  • compensation failure;
  • unexpected growth in workflow history;
  • divergence between workflow state and the domain database.

Security and tenant isolation

A workflow engine contains business metadata and often coordinates privileged actions. Treat it as production control-plane infrastructure.

Important controls include:

  • separate namespaces, accounts, or environments for production and non-production;
  • least-privilege worker identities per task queue or integration domain;
  • encrypted transport and storage;
  • secrets supplied to activities at runtime, not stored in workflow input;
  • payload encryption or codecs for sensitive fields;
  • authenticated and authorized signals, callbacks, and operator actions;
  • immutable audit records for human decisions;
  • tenant context included and validated at every activity boundary;
  • retention and deletion policies aligned with regulatory obligations.

Avoid placing large documents, tokens, or regulated data directly into workflow history. Store large payloads in an appropriate encrypted system and pass a versioned reference with integrity metadata.

Temporal, Step Functions, queues, or application code?

The right choice depends on workflow duration, complexity, hosting model, language needs, operational maturity, and how deeply the workflow belongs to the product.

OptionBest fitStrengthsWatch for
Application code and databaseShort, simple flows inside one serviceMinimal platform overhead, familiar debuggingCustom retries, timers, recovery, and operations accumulate quickly
Queue-based workersIndependent asynchronous jobs and bufferingThroughput, backpressure, decouplingNo native end-to-end process model; teams often reinvent state machines
AWS Step Functions StandardAWS-centric integrations and auditable workflows up to one yearFully managed, visual state, service integrations, durable executionState-transition cost, service-specific definitions, payload and history limits
AWS Step Functions ExpressHigh-volume, short workflows up to five minutesHigh event rate, lower per-execution overheadAt-least-once asynchronous execution; unsuitable for long waits
TemporalProduct workflows with complex code, long waits, signals, and frequent evolutionCode-first model, rich durable semantics, testing, self-hosted or cloudRequires deterministic design and a maintained worker/platform model
BPMN engine such as CamundaProcesses shared between business and engineering teamsExplicit process diagrams, human tasks, governanceModel and runtime complexity; code-heavy logic can become awkward
Data orchestratorScheduled analytical and data dependenciesBackfills, datasets, lineage, batch operationsUsually not the right owner for transactional customer workflows

Use AWS Step Functions when the workflow is mostly AWS service integration, the managed operational model is valuable, and the state-machine constraints fit. Standard Workflows provide durable, auditable executions for long-running flows; Express Workflows suit short, high-volume processing with different delivery semantics.

Use Temporal when workflows are core product logic, need rich signals and long-lived state, or benefit from normal application-language abstractions and strong local testing. Temporal's event history and replay model are powerful, but the team must understand determinism and workflow versioning.

Use a queue when the unit of work is genuinely independent. Do not adopt a workflow engine for every background job.

For official implementation details, consult the current Temporal documentation and AWS Step Functions workflow type guide, because quotas and platform capabilities evolve.

Durable execution for AI agents

AI agents make workflow reliability more important, not less. A useful agent often combines model calls, retrieval, tool execution, approvals, and long waits. Without durable orchestration, a worker crash can lose the plan, repeat an expensive model call, or execute a tool twice.

Model inference and tool calls should be treated as activities because they are non-deterministic side effects. Persist the selected model, prompt or prompt version, tool arguments, result, token usage, and policy decision needed for audit and replay. Do not call a model directly from deterministic workflow logic.

A safe agent workflow often looks like this:

  1. Load a bounded, authorized context snapshot.
  2. Ask the model for a structured proposed action.
  3. Validate the proposal against deterministic policy.
  4. Request human approval for high-impact actions.
  5. Execute the tool with an idempotency key.
  6. Verify the resulting system state.
  7. Record the outcome and continue or compensate.

This architecture separates probabilistic reasoning from deterministic control. It complements the evaluation, audit, and intervention practices in our AgentOps guide and the broader comparison of AI agents and RPA.

Common failure patterns

One giant workflow

A workflow that coordinates every detail of a customer account becomes difficult to evolve and may accumulate unbounded history. Use child workflows or explicit process boundaries around coherent business outcomes.

Tiny activities for every line of code

Each activity adds persistence, scheduling, and operational overhead. Group operations into meaningful retry and ownership boundaries rather than remotely invoking trivial calculations.

Retrying permanent failures

Aggressive retries can amplify an outage and consume dependency quotas. Classify errors, cap attempts, use backoff and jitter, and route exhausted work to a visible remediation path.

Storing all business data in workflow history

Workflow history is an execution record, not a general-purpose database. Store authoritative records in domain systems and pass compact references.

Assuming exactly once removes duplicates

The workflow may be scheduled exactly once while an HTTP request is processed twice. Idempotency belongs at every side-effect boundary.

No operational intervention model

Production processes need supported ways to retry, cancel, pause, signal, compensate, and correct data. Ad hoc database edits create a second, undocumented control plane.

Ignoring in-flight executions during deployment

Workflow code is coupled to persisted history. Test replay compatibility and define how old executions reach completion before changing control flow.

A practical implementation plan

Phase 1: Select one valuable workflow

Choose a process with visible operational pain: manual recovery, duplicate effects, long waits, or unclear ownership. Map every step, dependency, deadline, and human decision before choosing technology.

Phase 2: Define the reliability contract

For each activity, document:

  • idempotency key and deduplication behavior;
  • timeout and retry policy;
  • permanent versus transient failures;
  • authoritative system of record;
  • compensation or reconciliation action;
  • required credentials and tenant scope;
  • observability and audit fields.

Phase 3: Build a thin vertical slice

Implement one complete path with production-grade entry validation, a real dependency, a durable wait, telemetry, and operator visibility. A shallow demo across twenty integrations proves less than one path that survives failure.

Phase 4: Test failure, not only success

Kill workers mid-activity. Lose responses after successful side effects. Deliver callbacks twice and out of order. Deploy a new workflow version with old executions active. Disable a dependency and observe retry pressure. Verify compensation and manual recovery.

Phase 5: Establish platform ownership

Define who owns namespaces, task queues, worker capacity, deployment compatibility, retention, encryption, quotas, upgrades, and incident response. A workflow engine removes custom reliability code, but it does not remove operational responsibility.

Phase 6: Expand from measured results

Track stuck-case count, manual interventions, completion time, duplicate effects, recovery time, and engineering effort. Expand the platform where those metrics justify it, not because every asynchronous function looks like a workflow.

The business case

Durable execution is valuable when failed processes create support work, revenue leakage, compliance risk, or customer-visible delays. It converts hidden recovery logic into explicit, testable control flow.

The goal is not to add another infrastructure product. The goal is to make important business operations resumable, observable, and safe under real failure conditions.

BoundLayer designs and implements production workflow systems across backend services, AWS infrastructure, integrations, data pipelines, human approvals, and AI automation. We can audit an existing collection of queues and cron jobs, identify the right orchestration boundaries, choose an appropriate platform, and carry the workflow through implementation and operations.

For broader modernization work, see how a forward-deployed engineering partner connects architecture decisions to the actual operating process rather than stopping at a diagram.

Building a business workflow that must survive failures?

We design and implement durable workflow systems with Temporal, AWS Step Functions, queues, integrations, human approvals, observability, and recovery controls.

Free consultation

Get a Free 30-Minute Technical Consultation

Share a few details about your project and we'll get back to you within 48 hours with a clear next step.

  • No sales pressure — a senior engineer, not a sales rep
  • Clear next step within 48 hours
  • We can sign an NDA before we talk

By submitting, you agree to be contacted about your request. We respect your privacy and can sign an NDA on request.