You have built an agent workflow in .NET. Three executors, a fan-out for parallel research, a barrier to collect results, and a final summariser. It works beautifully in your console app. Then a deployment goes wrong, the process recycles, and everything — the intermediate state, the LLM responses you already paid for, the human approval that someone clicked two hours ago — evaporates. You start over from scratch.

This is the gap between a demo and a production system. The Microsoft Agent Framework (MAF) gives you a clean, graph-based model for orchestrating agents and deterministic logic. But its in-memory execution model treats process boundaries as hard walls. The Microsoft.Agents.AI.DurableTask package removes those walls, giving your workflows automatic checkpointing, distributed execution, and the ability to survive restarts — all without changing a single executor.

A quick refresher on MAF workflows

If you have read the earlier coverage of the Microsoft Agent Framework, you already know the building blocks. A workflow is a directed graph of executors connected by edges. Each executor is a strongly-typed processing unit — it takes TInput, does work, and produces TOutput. The WorkflowBuilder wires them together:

Workflows/OrderCancellationWorkflow.cs
var workflow = new WorkflowBuilder(validateRequest)
    .WithName("CancelOrder")
    .WithDescription("Handles order cancellation with refund processing")
    .AddEdge(validateRequest, checkEligibility)
    .AddEdge(checkEligibility, processRefund)
    .AddEdge(processRefund, notifyCustomer)
    .Build();

Edge types control the flow. AddEdge is sequential. AddFanOutEdge branches into parallel executors. AddFanInBarrierEdge waits for all parallel branches to complete before continuing. AddSwitch routes conditionally based on the output of the previous executor.

The critical detail: the workflow definition is just a graph — it says nothing about how it runs. That separation is what makes the durable extension possible.

Stage one: in-memory execution

During development, you run workflows in-process with zero infrastructure:

Program.cs
using Microsoft.Agents.AI;
using Microsoft.Agents.AI.Workflows;

var workflow = BuildCancelOrderWorkflow();
var request = new CancellationRequest("ORD-98234", "Customer changed mind");

await foreach (var evt in InProcessExecution.RunStreamingAsync(workflow, request))
{
    if (evt is ExecutorCompletedEvent completed)
    {
        Console.WriteLine($"[{completed.ExecutorId}] finished");
    }
}

InProcessExecution.RunStreamingAsync takes the graph, walks it, and yields WorkflowEvent objects as each executor completes. No Docker, no connection strings, no external dependencies. The trade-off is obvious: kill the process and everything is lost.

This is perfectly fine for development. You can iterate on executor logic, test edge conditions, and debug the graph structure without waiting for infrastructure. The mistake is shipping this to production.

Stage two: adding durability with the Durable Task Scheduler

The Durable Task Scheduler (DTS) is the persistence backbone. It stores workflow state, coordinates distributed execution, and provides an observability dashboard. For local development, it runs as a Docker container:

terminal
docker run -d --name dts-emulator \
  -p 8080:8080 \
  -p 8082:8082 \
  mcr.microsoft.com/dts/dts-emulator:latest

Port 8080 is the gRPC endpoint your application connects to. Port 8082 serves the dashboard where you can inspect running orchestrations, view execution history, and diagnose failures.

To make your workflow durable, add the Microsoft.Agents.AI.DurableTask package and swap the hosting. The workflow definition stays identical:

Program.cs
using Microsoft.Agents.AI;
using Microsoft.Agents.AI.DurableTask;
using Microsoft.DurableTask.Client.AzureManaged;
using Microsoft.DurableTask.Worker.AzureManaged;
using Microsoft.Extensions.Hosting;

var workflow = BuildCancelOrderWorkflow();

var host = Host.CreateDefaultBuilder(args)
    .ConfigureDurableWorkflows(workflows =>
    {
        workflows.AddWorkflow(workflow);
        workflows.UseDurableTaskScheduler(
            "Endpoint=http://localhost:8080;TaskHub=default;Authentication=None");
    })
    .Build();

await host.RunAsync();

ConfigureDurableWorkflows does the heavy lifting. It registers each executor as a durable activity, wires up the orchestration, and configures the DTS connection. Your executor code is untouched — the framework wraps it transparently.

What changes at runtime is significant. After each executor completes, its output is checkpointed to the DTS. If the process crashes halfway through the workflow, the next instance picks up from the last checkpoint rather than starting over. If an executor calls an LLM and gets a response, that response is persisted — a replay will not call the LLM again.

// TIP

The DTS emulator stores state in memory, so it is not suitable for production. It exists purely to let you develop and test durable workflows locally without an Azure subscription.

Stage three: deploying to Azure Functions

The same workflow definition deploys to Azure Functions with a different hosting call:

Program.cs
using Microsoft.Agents.AI.Hosting.AzureFunctions;

var workflow = BuildCancelOrderWorkflow();

FunctionsApplication
    .CreateBuilder(args)
    .ConfigureFunctionsWebApplication()
    .ConfigureDurableWorkflows(workflows =>
        workflows.AddWorkflow(workflow))
    .Build()
    .Run();

The Azure Functions host automatically provisions HTTP endpoints for your workflow:

Endpoint Purpose
POST /api/workflows/CancelOrder/run Starts a new workflow run
GET /api/workflows/CancelOrder/status/{runId} Checks run status (opt-in)
POST /api/workflows/CancelOrder/respond/{runId} Submits human approval responses

Start a workflow by posting the input as the request body. The response returns 202 Accepted with a run ID you can use to poll for status or submit responses to human-in-the-loop gates.

If you need synchronous execution — waiting for the result before returning — add the x-ms-wait-for-response: true header. The function will hold the connection open until the workflow completes. This is useful for short workflows where the caller expects an immediate result.

// WARNING

The synchronous wait has a timeout governed by your Azure Functions host configuration. For workflows that run longer than a few minutes, use the asynchronous pattern with status polling instead.

Durable agents: the simpler path for single-agent scenarios

Not every use case needs a full workflow graph. If you have a single agent (or a small set of independent agents) that needs durable conversation threads, the ConfigureDurableAgents API is a more direct fit:

Program.cs
using Microsoft.Agents.AI;
using Microsoft.Agents.AI.Hosting.AzureFunctions;

AIAgent triageAgent = chatClient.AsAIAgent(
    instructions: "You are a customer support triage agent.",
    name: "TriageAgent");

FunctionsApplication
    .CreateBuilder(args)
    .ConfigureFunctionsWebApplication()
    .ConfigureDurableAgents(options =>
        options.AddAIAgent(triageAgent))
    .Build()
    .Run();

This creates HTTP endpoints for the agent with automatic conversation persistence. Each conversation gets a thread ID (returned in the x-ms-thread-id response header) that you can pass back on subsequent requests to maintain context. The conversation history survives process restarts, scale-out events, and instance migrations.

The distinction matters: ConfigureDurableWorkflows is for multi-step orchestrations with explicit graph topologies. ConfigureDurableAgents is for conversational agents that need persistent threads. Choose based on whether your problem is an orchestration or a conversation.

Shared state across executors

Workflows frequently need to pass data between executors beyond the direct input-output chain. The shared state API handles this:

Executors/ValidateRequestExecutor.cs
public class ValidateRequestExecutor : Executor<CancellationRequest, ValidationResult>
{
    protected override async Task<ValidationResult> HandleAsync(
        CancellationRequest input,
        IWorkflowContext context,
        CancellationToken cancellationToken)
    {
        var order = await _orderService.GetAsync(input.OrderId, cancellationToken);

        await context.QueueStateUpdateAsync(
            "originalAmount",
            order.TotalAmount,
            "refund",
            cancellationToken);

        return new ValidationResult(IsValid: true, Order: order);
    }
}

A downstream executor reads the state:

Executors/ProcessRefundExecutor.cs
protected override async Task<RefundResult> HandleAsync(
    ValidationResult input,
    IWorkflowContext context,
    CancellationToken cancellationToken)
{
    var amount = await context.ReadStateAsync<decimal>(
        "originalAmount",
        "refund",
        cancellationToken);

    return await _refundService.ProcessAsync(input.Order.Id, amount, cancellationToken);
}

State is scoped by name ("refund" in this example) to avoid collisions between unrelated data. In durable mode, this state is checkpointed alongside executor outputs — it survives restarts just like everything else.

Human-in-the-loop with RequestPort

Some workflows need a human to approve, reject, or provide additional information before continuing. MAF models this with RequestPort<TRequest, TResponse>, a special executor that pauses the workflow and waits for an external response:

Workflows/RefundApprovalWorkflow.cs
var approvalGate = new RequestPort<RefundApproval, ApprovalDecision>("ManagerApproval");

var workflow = new WorkflowBuilder(calculateRefund)
    .WithName("RefundWithApproval")
    .AddEdge(calculateRefund, approvalGate)
    .AddEdge(approvalGate, processRefund)
    .Build();

When the workflow reaches the RequestPort, it emits an event and suspends. In Azure Functions, the /respond/{runId} endpoint accepts the approval response. When deployed with DTS, the workflow pauses without consuming compute — it can wait for hours, days, or weeks. The serverless host scales to zero during the wait, eliminating cost.

This is fundamentally different from polling or holding a connection open. The orchestration state is persisted, the process can terminate entirely, and when the human response arrives, DTS rehydrates the workflow and resumes from exactly where it paused.

Sub-workflows and composition

Complex systems decompose into smaller, reusable workflows. MAF supports this with BindAsExecutor, which turns a workflow into an executor that can be used as a node in a parent workflow:

Workflows/OrderLifecycleWorkflow.cs
var refundWorkflow = BuildRefundWorkflow();
var refundExecutor = refundWorkflow.BindAsExecutor("RefundSubWorkflow");

var lifecycle = new WorkflowBuilder(receiveOrder)
    .WithName("OrderLifecycle")
    .AddSwitch(routeDecision, new Dictionary<string, IExecutor>
    {
        ["cancel"] = refundExecutor,
        ["modify"] = modifyExecutor,
        ["complete"] = shipExecutor
    })
    .Build();

The sub-workflow runs as a nested orchestration with its own checkpointing, state scope, and failure boundaries. A failure in the refund sub-workflow does not corrupt the parent orchestration's state.

Observability out of the box

The DTS dashboard (port 8082 in the emulator, or the Azure portal in production) shows every running and completed orchestration. You can see which executor is currently active, inspect the inputs and outputs of completed steps, and view the full replay history.

In Azure Functions, Application Insights integration gives you distributed tracing across workflows. Each executor emits trace spans, so you can follow a workflow execution through your existing monitoring pipeline without custom instrumentation.

// NOTE

The DTS dashboard is a diagnostic tool, not a workflow management UI. To programmatically interact with running workflows — checking status, sending events, or cancelling runs — use the IWorkflowClient interface.

Common pitfalls

Non-deterministic executor logic. Durable orchestrations replay executor history to reconstruct state. If an executor produces different output on replay (because it reads the current time, generates a random number, or calls an external service), the replay diverges. Use IWorkflowContext for time-dependent logic and ensure side effects are idempotent.

Large payloads in executor outputs. Every executor's output is serialised and stored in the DTS. If an executor returns a 50 MB document, that 50 MB gets checkpointed, replayed, and passed to the next executor. Keep outputs lean — store large data externally and pass references.

Confusing workflows with agents. A workflow is a predetermined graph; an agent decides its own path via an LLM. If you wrap an agent as an executor, the agent's internal tool calls are not individually checkpointed — only the final response is. A crash mid-agent-execution replays the entire agent call. Design for this by keeping agent executors focused on single-turn interactions.

Skipping the emulator. Deploying directly to Azure without local testing means every debugging cycle involves deployment, log streaming, and guesswork. The DTS emulator is lightweight and fast. Use it.

Ignoring the connection string in production. The local emulator connection string (Authentication=None) must never reach production. Use Azure Managed Identity with the DTS resource in Azure, and keep connection strings in environment variables or Azure Key Vault.

Summary