Connection Resiliency in EF Core: Surviving Transient Failures

Database connections fail. Network blips, Azure SQL elastic pool rebalancing, brief maintenance windows — in cloud environments, transient failures are a fact of life. EF Core's connection resiliency feature automatically retries failed operations, turning what would be an unhandled exception into a brief, invisible pause.

Enabling Retry Logic

The SQL Server provider includes a built-in execution strategy. Enable it with a single call:

Example.cs
services.AddDbContext<AppDbContext>(options =>
    options.UseSqlServer(connectionString,
        sqlOptions => sqlOptions.EnableRetryOnFailure()));

With default settings, this retries up to 6 times with an exponential backoff. It only retries errors that SQL Server identifies as transient — timeouts, connection drops, and specific error codes.

Customising the Retry Strategy

Control the retry count, delay, and which errors to retry:

Example.cs
services.AddDbContext<AppDbContext>(options =>
    options.UseSqlServer(connectionString,
        sqlOptions => sqlOptions.EnableRetryOnFailure(
            maxRetryCount: 3,
            maxRetryDelay: TimeSpan.FromSeconds(10),
            errorNumbersToAdd: [4060, 40197, 40501, 40613])));

The errorNumbersToAdd parameter extends the default list of transient error codes. The defaults already cover the most common transient SQL Server errors, but you can add application-specific ones.

How It Works Under the Hood

When EF Core detects a transient failure, it:

  1. Catches the SqlException
  2. Checks whether the error code is in the transient list
  3. Waits for an exponentially increasing delay (with random jitter)
  4. Retries the entire operation

The retry wraps the complete SaveChanges or query operation. If the operation partially succeeded before the failure, EF Core will re-execute the entire thing.

Transactions and Retry Logic

This is where resiliency gets tricky. If you're using explicit transactions, the retry strategy needs to replay the entire transaction — not just the failing statement:

Example.cs
// This WON'T work with retry logic:
using var transaction = await context.Database.BeginTransactionAsync();
context.Products.Add(product);
await context.SaveChangesAsync();
context.Orders.Add(order);
await context.SaveChangesAsync(); // If this fails and retries, the Add above won't replay
await transaction.CommitAsync();

Instead, use CreateExecutionStrategy to wrap the entire transaction:

Example.cs
var strategy = context.Database.CreateExecutionStrategy();

await strategy.ExecuteAsync(async () =>
{
    await using var transaction = await context.Database.BeginTransactionAsync();

    context.Products.Add(product);
    await context.SaveChangesAsync();

    context.Orders.Add(order);
    await context.SaveChangesAsync();

    await transaction.CommitAsync();
});

Now if any part fails transiently, the entire block — including creating the transaction — is retried from scratch.

Custom Execution Strategies

For more control, implement your own ExecutionStrategy:

CustomRetryStrategy.cs
public class CustomRetryStrategy : SqlServerRetryingExecutionStrategy
{
    private readonly ILogger<CustomRetryStrategy> _logger;

    public CustomRetryStrategy(
        DbContext context,
        int maxRetryCount,
        TimeSpan maxRetryDelay,
        ILogger<CustomRetryStrategy> logger)
        : base(context, maxRetryCount, maxRetryDelay, null)
    {
        _logger = logger;
    }

    protected override bool ShouldRetryOn(Exception exception)
    {
        var shouldRetry = base.ShouldRetryOn(exception);

        if (shouldRetry)
        {
            _logger.LogWarning(exception,
                "Transient database error detected. Retrying...");
        }

        return shouldRetry;
    }

    protected override TimeSpan? GetNextDelay(Exception lastException)
    {
        var delay = base.GetNextDelay(lastException);

        if (delay.HasValue)
        {
            _logger.LogInformation(
                "Next retry in {Delay}ms", delay.Value.TotalMilliseconds);
        }

        return delay;
    }
}

Register it:

Example.cs
services.AddDbContext<AppDbContext>(options =>
    options.UseSqlServer(connectionString,
        sqlOptions => sqlOptions.ExecutionStrategy(
            deps => new CustomRetryStrategy(
                deps.CurrentContext.Context,
                maxRetryCount: 5,
                maxRetryDelay: TimeSpan.FromSeconds(30),
                deps.CurrentContext.Context
                    .GetService<ILogger<CustomRetryStrategy>>()))));

Idempotency Matters

Because the strategy retries the entire operation, your code must be idempotent. If SaveChanges sent the INSERT but the connection dropped before the acknowledgement arrived, the retry will attempt the INSERT again.

Guard against this:

Example.cs
var strategy = context.Database.CreateExecutionStrategy();

await strategy.ExecuteAsync(async () =>
{
    // Check if the operation already succeeded
    var exists = await context.Orders.AnyAsync(o => o.ExternalId == externalId);
    if (exists) return;

    context.Orders.Add(new Order { ExternalId = externalId, /* ... */ });
    await context.SaveChangesAsync();
});

Using unique constraints or natural keys helps the database reject duplicate inserts even if your code doesn't catch them.

PostgreSQL and Other Providers

Npgsql (PostgreSQL) has its own retry strategy:

Example.cs
services.AddDbContext<AppDbContext>(options =>
    options.UseNpgsql(connectionString,
        npgsqlOptions => npgsqlOptions.EnableRetryOnFailure()));

The concept is identical; only the transient error codes differ.

Monitoring Retries

Combine retry strategies with EF Core's diagnostic events to track retry frequency:

Example.cs
services.AddDbContext<AppDbContext>(options =>
    options.UseSqlServer(connectionString,
            sqlOptions => sqlOptions.EnableRetryOnFailure())
           .LogTo(Console.WriteLine,
               [DbLoggerCategory.Infrastructure.Name],
               LogLevel.Warning));

If you see frequent retries, investigate the root cause — it might be a connection pool exhaustion issue, a database under heavy load, or a network problem worth addressing directly.

Connection resiliency is essential for any cloud-hosted application. Enable it early, wrap your transactions properly, and keep your operations idempotent. The few minutes of setup will save you from countless transient failures in production.