Failure Is Not an Option: Durable Execution + Dapr = π - Marc Duiker, Diagrid
Marc Duiker, Diagrid
KubeCon + CloudNativeCon Europe 2025 Β· Session
Overview
In the realm of distributed systems, failure is not merely an option; it's an inevitability. This insightful talk by Marc Duiker, a Developer Advocate at Diagrid, delves into the critical challenge of building resilient applications that can not only withstand failures but also recover from them automatically and gracefully. Titled "Failure Is Not an Option: Durable Execution + Dapr = π," the presentation introduces Durable Execution as a powerful paradigm for achieving this resilience, particularly when combined with the Dapr (Distributed Application Runtime) framework. The session highlights how Dapr Workflow, a relatively new and stable API within Dapr, provides developers with the tools to implement robust, stateful workflows that persist their state across failures, ensuring business processes complete successfully even in the face of outages.

Key moments
- 0:00 Introduction: Failure is inevitable, let's fail successfully
- 2:00 Monetary cost of IT failures: hundreds of billions
- 3:20 Introducing Dapr: The Distributed Application Runtime and CNCF graduation
- 4:40 Understanding durable execution for automatic recovery from failures
- 6:00 Visualizing how durable execution and state persistence work
Failure Is Not an Option: Durable Execution + Dapr = π
Speakers: Marc Duiker, Developer Advocate, Diagrid
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=KPNuLwXNkNQ
Overview
In the realm of distributed systems, failure is not merely an option; it's an inevitability. This insightful talk by Marc Duiker, a Developer Advocate at Diagrid, delves into the critical challenge of building resilient applications that can not only withstand failures but also recover from them automatically and gracefully. Titled "Failure Is Not an Option: Durable Execution + Dapr = π," the presentation introduces Durable Execution as a powerful paradigm for achieving this resilience, particularly when combined with the Dapr (Distributed Application Runtime) framework. The session highlights how Dapr Workflow, a relatively new and stable API within Dapr, provides developers with the tools to implement robust, stateful workflows that persist their state across failures, ensuring business processes complete successfully even in the face of outages.
The core premise of Duiker's talk is that while IT failures are costly and disruptive, modern architectural patterns and tools can transform these failures into "successful failures" β situations where systems automatically self-heal or minimize user impact. This is achieved through the concept of durable execution, where the state of long-running operations is reliably persisted, allowing them to resume from their last known good point after an interruption. For developers grappling with the complexities of microservices, asynchronous communication, and distributed state management, Dapr Workflow emerges as a compelling solution, simplifying the creation of highly available and fault-tolerant applications.
The talk is particularly relevant for architects and developers building or migrating to cloud-native, microservice-based applications, especially those already using or considering Dapr. It offers a practical, code-centric exploration of how to design and implement workflows that are inherently resilient, shifting the burden of failure recovery from complex custom logic to a robust, open-source runtime. By leveraging Dapr's sidecar model and its integration with various state stores, Dapr Workflow aims to abstract away much of the distributed systems complexity, allowing developers to focus on business logic rather than intricate error handling and retry mechanisms.
Background
βΆ Watch: Introduction: Failure is inevitable, let's fail successfully (0:00)
The pervasive nature and substantial financial impact of IT failures underscore the necessity of resilient system design. Marc Duiker referenced a 2015 IEEE Spectrum blog post, "Lessons Learned from a Decade of IT Failure," which estimated the monetary cost of IT failures to be in the "hundreds of billions" a decade ago, a figure likely even higher today. This escalating cost is largely attributable to the increasing complexity of modern applications, which are predominantly distributed, relying on intricate webs of synchronously and asynchronously communicating services, message brokers, and diverse state stores. With numerous moving parts, the probability of something going wrong inherently increases.
This challenge is not new; the Fallacies of Distributed Computing, a list of false assumptions compiled in 1994, has long highlighted the inherent difficulties of building distributed systems. These fallacies, such as "the network is reliable" or "latency is zero," often lead developers to underestimate the complexities of distributed environments. Consequently, designing applications that can automatically recover from failures and limit their impact on users remains a paramount concern.
Approximately five years ago, a group of "smart people" initiated the Dapr (Distributed Application Runtime) open-source project to address these challenges. Dapr has since grown significantly, becoming a graduated project within the CNCF in November, used by hundreds of companies across various verticals. Dapr functions as a sidecar process running alongside any application, offering a rich set of APIs that abstract away common distributed systems patterns. These APIs include state management, pub/sub, service invocation, and more, enabling developers to build and run microservices at scale with greater ease, regardless of their chosen language or framework.
The specific problem Dapr Workflow aims to solve is the need for durable execution. Durable execution refers to the ability to run code in a stateful manner, meaning that if the process executing the code fails, another process can seamlessly resume execution from the point of failure. This is achieved by persistently storing the application's state to disk. While the term "durable execution" might be relatively new to some, the underlying concept has been implemented for many years by workflow engines. These engines orchestrate business processes composed of multiple tasks or activities, which are small units of work like calling an API, saving state, or publishing a message.
The mechanism behind durable execution and workflow engines involves meticulously persisting all state changes to a state store. As Duiker illustrated with an animation, when a workflow starts, its input and a unique workflow ID are stored. Each activity's input and output are also persisted. Crucially, after each activity completes, the workflow doesn't simply proceed to the next step; instead, it effectively "replays" from the beginning, reading the stored state to determine which activities have already completed successfully. This replay mechanism, combined with state persistence, ensures that even if the workflow process crashes, it can restart and continue from its last recorded state without re-executing completed idempotent activities or losing context.
Workflows can be authored in different ways: visually via click-and-drag designers, declaratively via JSON-based step functions, or through workflow as code (WAC) solutions. Marc Duiker expressed a strong preference for WAC, citing benefits such as version control integration, peer review via pull requests, and the ability to unit test business logic, which is critical for correctness. Examples of WAC solutions include Temporal, Azure Durable Functions, and the focus of this talk, Dapr Workflow. The Dapr Workflow API has been stable since the Dapr 1.15 release in February, following over two years of development, marking a significant milestone for the project.
Key Findings
βΆ Watch: Monetary cost of IT failures: hundreds of billions (2:00)
The central revelation of this presentation is the maturity and capability of Dapr Workflow as a robust solution for durable execution in distributed systems. Its stabilization with the Dapr 1.15 release in February signifies its readiness for production use, offering developers a powerful tool to build resilient, self-recovering applications.
A key finding is that Dapr Workflow effectively tackles the inherent unreliability of distributed computing by providing an automatic failure recovery mechanism. The demonstration vividly illustrated that even if an application process crashes mid-workflow, Dapr's sidecar, coupled with a persistent state store, enables the workflow to resume precisely from its last known good state upon restart, without requiring manual intervention or re-triggering the initial request. This capability is a significant contribution to simplifying the development of highly available systems.
Dapr Workflow supports several common and powerful workflow patterns, allowing developers to express complex business logic clearly and concisely in code. These patterns include:
- Task Chaining: Executing activities sequentially, where the output of one activity serves as the input for the next.
- Fan-out/Fan-in: Running multiple independent activities in parallel and then aggregating their results once all have completed.
- Monitor Pattern: Implementing reoccurring tasks, such as nightly cleanup jobs, by creating timers and continuing the workflow as a new instance after a specified delay.
- External System Interaction: Enabling workflows to pause and wait for external events, such as human approvals, before proceeding.
Architecturally, Dapr Workflow integrates seamlessly into the Dapr ecosystem. The workflow engine resides within the Dapr sidecar, establishing a gRPC stream with the application. While the application's code defines the workflow logic, the sidecar is responsible for scheduling activities, managing state persistence to the configured state store (e.g., Redis), and orchestrating the replay mechanism. This design offloads crucial durability concerns from the application developer to the Dapr runtime.
Furthermore, the talk highlighted critical considerations for designing durable workflows effectively. These include:
- Deterministic Code: Workflows must be deterministic to ensure consistent replay behavior. Non-deterministic operations (like generating random GUIDs or current timestamps) need special handling or encapsulation within activities.
- Idempotent Activities: Activities should be designed to be idempotent, meaning they can be executed multiple times without producing unintended side effects, due to Dapr Workflow's "at least once" execution guarantee.
- Versioning Strategies: Significant changes to workflow logic (e.g., altering activity order or input/output types) constitute breaking changes that necessitate proper versioning, typically by creating new workflow definitions.
- Efficient Argument Passing: Minimizing the size of data passed between activities is crucial, as all such arguments are persisted to the state store, impacting performance and storage costs.
These findings collectively demonstrate Dapr Workflow's potential to significantly enhance the resilience and maintainability of distributed applications, providing a robust, opinionated framework for durable execution.
Technical Deep Dive
βΆ Watch: Introducing Dapr: The Distributed Application Runtime and CNCF graduation (3:20)
The technical foundation of Dapr Workflow is built upon the established principles of durable execution and Dapr's sidecar architecture. The Dapr Workflow engine is an integral part of the Dapr sidecar, which runs as a companion process to the application. Communication between the application, where the workflow logic resides, and the Dapr sidecar occurs via a gRPC stream. This stream facilitates the scheduling of workflow activities and the crucial persistence of workflow state. The sidecar orchestrates the workflow's lifecycle, including initiating replays and ensuring state consistency, while the application code provides the business logic.
A fundamental aspect of durable execution is state persistence. Dapr Workflow leverages Dapr's existing state store components. Developers define their chosen state store (e.g., Redis, Cosmos DB, SQL Server) using a YAML component file, such as components/statestore.yaml. A critical configuration detail for Dapr Workflow is setting actorStateStore: true within the state store component, as Dapr Workflow is built on top of Dapr's actor model, though developers do not directly interact with the actor API. This state store is where all workflow inputs, outputs, and intermediate states are durably saved, enabling recovery.
Dapr Workflow promotes workflow as code, supporting multiple languages including C#, Java, JavaScript, Python, and Go. In the C# example provided, a workflow is defined as a class inheriting from a Dapr SDK Workflow class, implementing a RunAsync method. This method receives a WorkflowContext, which is the primary interface for interacting with the workflow engine.
Key workflow patterns are implemented using the WorkflowContext:
- Task Chaining: Sequential execution is achieved by calling
await context.CallActivityAsync<TOutput>("ActivityName", inputPayload);. WhenCallActivityAsyncis invoked, the workflow engine schedules the activity. The application code then stops executing and is unloaded from memory. Upon completion of the activity, the workflow engine replays the workflow from the beginning, using the persisted state to skip already completed activities, and resumes execution of theRunAsyncmethod from the point whereawaitwas called.
- Fan-out/Fan-in: For parallel execution of independent activities,
context.CallActivityAsynccalls are made withoutawaitwithin a loop, storing the returnedTaskobjects in a list. The workflow then waits for all these tasks to complete usingawait Task.WhenAll(listOfTasks);. This allows for efficient parallel processing, followed by aggregation of results.
- Monitor Pattern: To implement reoccurring tasks or timed delays,
await context.CreateTimer(TimeSpan.FromHours(24));can be used. After the timer expires, the workflow becomes active again. For creating a new, fresh instance of the workflow (e.g., for a nightly cleanup job),context.ContinueAsNew(newInput);is invoked. This differs from standard replay as it discards the previous history, starting a new workflow instance with fresh state.
- External System Interaction: Workflows can pause and wait for external stimuli using
await context.WaitForExternalEventAsync<TEventPayload>("EventName");. This is ideal for scenarios requiring human approval or integration with external systems, where the workflow needs to halt until a specific event payload arrives.
Retry policies can be applied to activities via WorkflowTaskOptions. For instance, new WorkflowTaskOptions { RetryPolicy = new ConstantRetryPolicy(TimeSpan.FromSeconds(5), 3) } allows for retrying an activity a specified number of times with a constant delay. Exponential backoff policies are also supported.
Compensation actions are a crucial aspect of transactional integrity in distributed systems. While not a built-in Dapr Workflow primitive, they are implemented by writing specific activities to undo previous work. The example showed a try-catch block around a RegisterShipment activity. If an exception occurs, the catch block invokes an UndoUpdateInventory activity, effectively rolling back a prior state change. Exceptions within activities are wrapped in a WorkflowTaskFailedException by the Dapr SDK.
An activity itself is a simple class inheriting from WorkflowActivity, implementing a RunAsync method. Activities interact with the Dapr client to perform actual work, such as await daprClient.GetStateAsync<ProductInventory>("statestore", productId); to retrieve state or await daprClient.SaveStateAsync("statestore", productId, updatedInventory); to persist changes. Crucially, while workflows must be deterministic, activities can be non-deterministic, as their execution results are stored and replayed.
For Dapr to recognize workflows and activities, they must be registered during application startup using services.AddDaprWorkflow(workflowOptions => { workflowOptions.RegisterWorkflow<ValidateOrderWorkflow>(); workflowOptions.RegisterActivity<UpdateInventoryActivity>(); / ... / });.
Initiating a workflow is done via the DaprWorkflowClient, typically from an endpoint: var instanceId = await daprWorkflowClient.ScheduleNewWorkflowAsync("ValidateOrderWorkflow", orderId.ToString(), order);. This method is asynchronous and returns the workflow's instance ID, not its final result, because workflows can be long-running processes (hours, days, or months). The instance ID is essential for correlating subsequent operations or querying the workflow's status.
Demo / Proof of Concept
βΆ Watch: Understanding durable execution for automatic recovery from failures (4:40)
The core of Marc Duiker's presentation culminated in a compelling demonstration of Dapr Workflow's durable execution capabilities, specifically its automatic failure recovery. The scenario simulated a common e-commerce order processing workflow involving multiple distributed services.
The demo architecture comprised two main applications:
- A Workflow Application: This hosts the Dapr Workflow logic for validating and processing orders.
- A Shipping Application: A separate microservice responsible for retrieving shipping providers, calculating costs, and registering shipments.
- An Inventory Service: Implicitly managed by the Workflow Application, interacting with a Dapr state store to check and update product stock.
The workflow itself was structured as follows:
- Activity 1:
UpdateInventory: Checks if sufficient stock is available. If yes, it mutates the inventory; otherwise, the workflow exits. This activity uses the Dapr client to interact with the configured state store. - Activity 2:
GetShippingProviders: Retrieves a list of available shipping providers for the product. - Activity 3 (Fan-out/Fan-in):
GetShippingCost: For each shipping provider, this activity is called in parallel to obtain shipping cost estimates. The workflow then waits for all these calls to complete. - Business Logic: The workflow code then selects the cheapest shipping provider from the aggregated results.
- Activity 4:
RegisterShipment: Attempts to register the shipment with the chosen provider. This step was wrapped in atry-catchblock. - Compensation Action (
UndoUpdateInventory): IfRegisterShipmentfails, a compensation activity is invoked to reverse the initialUpdateInventoryaction, ensuring transactional consistency.
The demonstration setup involved:
- A Redis state store configured via
components/statestore.yaml, withactorStateStore: trueenabled. - A
dapr.yamlfile to orchestrate the local startup of both the Workflow and Shipping applications, along with their respective Dapr sidecars, using thedapr run -f dapr.yamlcommand.
The critical steps of the demo were:
- Initial Setup: Marc started the applications using
dapr run. He then initialized the inventory with "5 rubber ducks" via a dedicated endpoint. - Workflow Initiation: He sent a request to the
validate orderendpoint of the Workflow Application, attempting to order "2 rubber duckies." This initiated a new Dapr Workflow instance. - Simulated Failure: Crucially, while the workflow was in progress (specifically, after the shipping service was contacted to retrieve costs, but before it could return the information), Marc manually stopped all running processes (Workflow App, Shipping App, and Dapr sidecars). This simulated a severe application crash or infrastructure failure. The logs showed the last activity was the shipping service retrieving costs.
- Automatic Recovery: Marc then restarted the applications using the same
dapr runcommand. - Observation: Upon restart, the Dapr Workflow engine automatically detected the incomplete workflow. It seamlessly replayed the workflow from its last persisted state. The logs showed that the shipping services were contacted again, the cheapest one was selected, and finally, the
orchestration status completedlog message appeared. This demonstrated that the workflow continued from where it left off, successfully completing the order despite the intermediate crash, without any manual re-triggering of the order request. - Verification: The final state of the workflow (input and output) was retrieved via a
GET /status/{instanceId}endpoint, confirming its successful completion and the correct output.
This demo powerfully illustrated Dapr Workflow's ability to provide automatic, durable execution, making applications resilient to transient or even catastrophic failures by persisting state and enabling seamless recovery.
Defensive Implications
βΆ Watch: Visualizing how durable execution and state persistence work (6:00)
Implementing durable workflows with Dapr brings significant resilience benefits, but it also introduces specific design considerations that developers must address to ensure correctness and stability. Marc Duiker outlined four key defensive implications:
- Deterministic Workflows:
- Problem: Workflows are designed to be replayed multiple times from their persisted state. If the workflow code's behavior changes during a replay (i.e., it's non-deterministic), it can lead to a mismatch between the stored state and the current runtime execution, resulting in errors or inconsistent outcomes.
- Examples of Non-Deterministic Code: Directly using
Guid.NewGuid()to generate unique identifiers orDateTime.Now/DateTime.UtcNowto get current timestamps within the workflow logic. - Solution: Workflow code must be deterministic. Dapr Workflow provides helper methods on the
WorkflowContextto handle common non-deterministic scenarios, such ascontext.NewGuid()for generating stable GUIDs andcontext.CurrentUtcDateTimefor getting a consistent UTC timestamp during replay. Any other non-deterministic logic should be encapsulated within an activity. Activity code can be non-deterministic because its execution results are captured and stored, not replayed in the same manner as the workflow orchestrator logic.
- Idempotent Activities:
- Problem: Dapr Workflow guarantees "at least once" execution for activities. This means an activity might be executed more than once, especially during failure recovery and replay scenarios. If an activity is not idempotent (meaning it produces the same result and has no additional side effects when executed multiple times with the same input), it can lead to duplicate data, incorrect state, or unintended operations.
- Example: An activity that performs a
SQL INSERTwith a primary key might fail on a second attempt if the record already exists. - Solution: Design activities to be idempotent. This often involves:
- Using upsert operations (insert if not exists, update if exists) instead of simple inserts.
- Performing a read-before-write check to ensure the state is as expected before making changes.
- Ensuring that external APIs or services called by activities are themselves idempotent or that the activity logic handles repeated calls gracefully.
- Versioning:
- Problem: Making significant changes to a workflow, such as altering the order of activities, adding or removing activities, or changing the input/output types of activities, constitutes a breaking change. When new code with these changes is deployed while older workflow instances are still "in flight" (not yet completed), the new code's expectations about the workflow's state history will clash with the state persisted by the old version, leading to runtime errors.
- Example: Changing an activity's input from a simple
orderItemto theresultAof a previous activity. - Solution: The easiest and most recommended approach is to version your workflows. This means that whenever a breaking change is made, a new workflow definition should be created (e.g.,
ValidateOrderWorkflowV1,ValidateOrderWorkflowV2). The old version continues to process existing "in flight" instances, while new client requests are directed to the new version. This requires updating client applications to call the appropriate workflow version. Never change an existing workflow definition in a breaking way; always add a new one.
- Efficient Argument Passing:
- Problem: All arguments passed between activities (inputs and outputs) are serialized and persisted to the underlying state store. Passing large objects or documents between activities can lead to significant overhead in terms of storage consumption, serialization/deserialization latency, and overall performance.
- Example: An activity retrieves a large document, and this entire document is then passed as input to a subsequent activity. This results in the document being saved twice in the state store (once as the output of the first, once as the input of the second).
- Solution: Design activities to be more cohesive and less granular. Instead of breaking down operations into very small "micro-activities" that pass large amounts of data, consolidate related operations into a single activity. For instance, an activity could perform both a "get" and an "update" operation, passing only a small identifier (like an ID) as input, rather than retrieving a large object in one activity and then passing that entire object to another for an update. This minimizes the amount of data stored and retrieved from the state store, improving efficiency.
By diligently adhering to these defensive programming practices, developers can harness the full power of Dapr Workflow to build robust, scalable, and truly resilient distributed applications.
Key Takeaways
- Durable Execution for Resilience: Dapr Workflow provides a robust framework for durable execution, enabling distributed applications to automatically recover from failures by persisting workflow state, ensuring long-running processes complete successfully.
- Leverages Dapr Sidecar Architecture: The Dapr Workflow engine runs within the Dapr sidecar, abstracting away complex distributed systems concerns like state management and failure recovery, and communicating with application code via a gRPC stream.
- Supports Common Workflow Patterns: Developers can implement complex business logic using well-established workflow patterns such as task chaining, fan-out/fan-in, monitor patterns for recurring tasks, and external system interaction for human approvals or external events.
- Workflow as Code (WAC) Benefits: Dapr Workflow embraces the "workflow as code" paradigm, supporting multiple languages (C#, Java, JavaScript, Python, Go), which facilitates source control, peer review, and robust unit testing of business logic.
- Critical Design Considerations: To ensure reliability, workflows must adhere to principles of determinism (especially during replay), activities must be idempotent (due to "at least once" execution guarantee), versioning is crucial for managing breaking changes, and efficient argument passing is necessary to optimize state store usage.
- Production Readiness: The Dapr Workflow API achieved stability with the Dapr 1.15 release in February, building on Dapr's status as a graduated CNCF project, making it a mature option for building resilient cloud-native applications.
About the Speaker(s)
Marc Duiker is a Developer Advocate at Diagrid, a company co-founded by the creators of Dapr. He is also an active Dapr community manager, contributing significantly to the Dapr ecosystem and its adoption. Known for his engaging presentations, Marc is a proponent of pixel art and uses VS Code for his entire presentation and demo setup, showcasing his practical, developer-centric approach. His work at Diagrid focuses on helping developers understand and leverage Dapr to build resilient and scalable distributed applications.
Reviews
Dr. Zero (Offensive Security Researcher) β STRONG ACCEPT
This talk by Marc Duiker provides a solid, practical deep dive into Dapr Workflow, a critical component for building resilient distributed systems. It clearly articulates the challenge of inevitable failures in distributed environments and presents durable execution via Dapr as a robust solution. The presentation effectively covers architectural details, common workflow patterns, and crucial defensive design considerations, all reinforced by a compelling live demo of automatic failure recovery. While the core concept of durable execution isn't novel, its mature implementation within Dapr is highly relevant and actionable for developers.
Heather Calloway (CISO) β STRONG ACCEPT
Marc Duiker's presentation on Durable Execution with Dapr Workflow provides a highly relevant and actionable framework for building resilient distributed systems. It directly addresses the substantial business risk posed by IT failures, offering concrete technical mechanisms to ensure business processes complete successfully even in the face of outages. For any organization grappling with the complexities of cloud-native architectures, this talk offers a clear path to improving foundational system availability and reducing the downstream impact of inevitable failures.