API Rate Limiting in ASP.NET Core

Rate limiting protects your API from abuse, prevents resource exhaustion, and ensures fair usage across clients. ASP.NET Core 7+ includes built-in rate limiting middleware with support for fixed window, sliding window, token bucket, and concurrency limiter algorithms. No third-party packages needed.

Setting Up Rate Limiting

Add the rate limiting services and middleware:

Program.cs
var builder = WebApplication.CreateBuilder(args);

builder.Services.AddRateLimiter(options =>
{
    options.RejectionStatusCode = StatusCodes.Status429TooManyRequests;

    options.AddFixedWindowLimiter("fixed", limiterOptions =>
    {
        limiterOptions.PermitLimit = 100;
        limiterOptions.Window = TimeSpan.FromMinutes(1);
        limiterOptions.QueueProcessingOrder = QueueProcessingOrder.OldestFirst;
        limiterOptions.QueueLimit = 0; // No queuing — reject immediately
    });
});

var app = builder.Build();

app.UseRateLimiter();

app.MapGet("/api/products", GetProducts)
    .RequireRateLimiting("fixed");

app.Run();

This allows 100 requests per minute per client. When exceeded, the client receives a 429 Too Many Requests response.

Rate Limiting Algorithms

Fixed Window

Divides time into fixed intervals. Each interval has a request quota that resets at the boundary.

Example.cs
options.AddFixedWindowLimiter("fixed", o =>
{
    o.PermitLimit = 100;
    o.Window = TimeSpan.FromMinutes(1);
});

The weakness: a client can make 100 requests at 12:00:59 and another 100 at 12:01:01 — 200 requests in two seconds that happen to straddle a window boundary.

Sliding Window

Addresses the boundary problem by dividing each window into segments and sliding across them:

Example.cs
options.AddSlidingWindowLimiter("sliding", o =>
{
    o.PermitLimit = 100;
    o.Window = TimeSpan.FromMinutes(1);
    o.SegmentsPerWindow = 6; // 10-second segments
});

With 6 segments, unused permits from older segments roll off gradually rather than all at once. This produces smoother rate enforcement.

Token Bucket

Allows short bursts while enforcing a long-term average rate. Tokens are added to a bucket at a steady rate. Each request consumes a token.

Example.cs
options.AddTokenBucketLimiter("token-bucket", o =>
{
    o.TokenLimit = 100;            // Maximum burst size
    o.ReplenishmentPeriod = TimeSpan.FromSeconds(10);
    o.TokensPerPeriod = 20;       // 20 tokens every 10 seconds
    o.AutoReplenishment = true;
});

This allows a burst of up to 100 requests but sustains a maximum of 120 requests per minute (20 tokens every 10 seconds).

Concurrency Limiter

Limits the number of concurrent requests rather than requests per time window:

Example.cs
options.AddConcurrencyLimiter("concurrency", o =>
{
    o.PermitLimit = 10;
    o.QueueProcessingOrder = QueueProcessingOrder.OldestFirst;
    o.QueueLimit = 5;
});

This is useful for endpoints that perform expensive operations — database-heavy reports, file processing, or external API calls.

Per-Client Rate Limiting

The examples above apply a single limit across all clients. In practice, you want per-client limits based on API key, user identity, or IP address:

Example.cs
builder.Services.AddRateLimiter(options =>
{
    options.RejectionStatusCode = StatusCodes.Status429TooManyRequests;

    options.AddPolicy("per-user", httpContext =>
    {
        var userId = httpContext.User.FindFirstValue(ClaimTypes.NameIdentifier)
            ?? httpContext.Connection.RemoteIpAddress?.ToString()
            ?? "anonymous";

        return RateLimitPartition.GetFixedWindowLimiter(userId, _ =>
            new FixedWindowRateLimiterOptions
            {
                PermitLimit = 100,
                Window = TimeSpan.FromMinutes(1)
            });
    });
});

Each authenticated user gets their own 100 requests/minute quota. Unauthenticated requests fall back to IP-based limiting.

Tiered Rate Limits

Different API tiers can have different limits:

Example.cs
options.AddPolicy("tiered", httpContext =>
{
    var tier = httpContext.User.FindFirstValue("subscription_tier") ?? "free";

    return RateLimitPartition.GetTokenBucketLimiter(
        $"{tier}:{httpContext.User.Identity?.Name}",
        _ => tier switch
        {
            "enterprise" => new TokenBucketRateLimiterOptions
            {
                TokenLimit = 1000,
                ReplenishmentPeriod = TimeSpan.FromSeconds(10),
                TokensPerPeriod = 200,
                AutoReplenishment = true
            },
            "pro" => new TokenBucketRateLimiterOptions
            {
                TokenLimit = 500,
                ReplenishmentPeriod = TimeSpan.FromSeconds(10),
                TokensPerPeriod = 100,
                AutoReplenishment = true
            },
            _ => new TokenBucketRateLimiterOptions
            {
                TokenLimit = 50,
                ReplenishmentPeriod = TimeSpan.FromSeconds(10),
                TokensPerPeriod = 10,
                AutoReplenishment = true
            }
        });
});

Response Headers

Communicate rate limit status to clients using standard headers. Customise the rejection response:

Example.cs
options.OnRejected = async (context, ct) =>
{
    context.HttpContext.Response.StatusCode = StatusCodes.Status429TooManyRequests;

    if (context.Lease.TryGetMetadata(MetadataName.RetryAfter, out var retryAfter))
    {
        context.HttpContext.Response.Headers.RetryAfter =
            ((int)retryAfter.TotalSeconds).ToString();
    }

    await context.HttpContext.Response.WriteAsJsonAsync(new ProblemDetails
    {
        Title = "Too Many Requests",
        Detail = "Rate limit exceeded. Please retry after the specified delay.",
        Status = StatusCodes.Status429TooManyRequests
    }, ct);
};

The Retry-After header tells clients exactly when they can retry, preventing unnecessary immediate retries.

Applying Rate Limits Selectively

Apply different policies to different endpoints:

Example.cs
app.MapGet("/api/products", GetProducts)
    .RequireRateLimiting("per-user");

app.MapPost("/api/orders", CreateOrder)
    .RequireRateLimiting("tiered");

// Health checks and documentation should not be rate limited
app.MapGet("/health", () => Results.Ok())
    .DisableRateLimiting();

Distributed Rate Limiting

The built-in middleware stores state in memory, which means it does not work across multiple server instances. For distributed scenarios, you have two options:

  1. Sticky sessions — route each client to the same server instance. Simple but reduces load balancer flexibility.
  2. Redis-backed rate limiting — use a package like RedisRateLimiting that stores counters in Redis, providing shared state across instances.

For most applications, the built-in middleware is sufficient. When you scale horizontally, plan for distributed state from the start.