API Rate Limiting in ASP.NET Core
Rate limiting protects your API from abuse, prevents resource exhaustion, and ensures fair usage across clients. ASP.NET Core 7+ includes built-in rate limiting middleware with support for fixed window, sliding window, token bucket, and concurrency limiter algorithms. No third-party packages needed.
Setting Up Rate Limiting
Add the rate limiting services and middleware:
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddRateLimiter(options =>
{
options.RejectionStatusCode = StatusCodes.Status429TooManyRequests;
options.AddFixedWindowLimiter("fixed", limiterOptions =>
{
limiterOptions.PermitLimit = 100;
limiterOptions.Window = TimeSpan.FromMinutes(1);
limiterOptions.QueueProcessingOrder = QueueProcessingOrder.OldestFirst;
limiterOptions.QueueLimit = 0; // No queuing — reject immediately
});
});
var app = builder.Build();
app.UseRateLimiter();
app.MapGet("/api/products", GetProducts)
.RequireRateLimiting("fixed");
app.Run();
This allows 100 requests per minute per client. When exceeded, the client receives a 429 Too Many Requests response.
Rate Limiting Algorithms
Fixed Window
Divides time into fixed intervals. Each interval has a request quota that resets at the boundary.
options.AddFixedWindowLimiter("fixed", o =>
{
o.PermitLimit = 100;
o.Window = TimeSpan.FromMinutes(1);
});
The weakness: a client can make 100 requests at 12:00:59 and another 100 at 12:01:01 — 200 requests in two seconds that happen to straddle a window boundary.
Sliding Window
Addresses the boundary problem by dividing each window into segments and sliding across them:
options.AddSlidingWindowLimiter("sliding", o =>
{
o.PermitLimit = 100;
o.Window = TimeSpan.FromMinutes(1);
o.SegmentsPerWindow = 6; // 10-second segments
});
With 6 segments, unused permits from older segments roll off gradually rather than all at once. This produces smoother rate enforcement.
Token Bucket
Allows short bursts while enforcing a long-term average rate. Tokens are added to a bucket at a steady rate. Each request consumes a token.
options.AddTokenBucketLimiter("token-bucket", o =>
{
o.TokenLimit = 100; // Maximum burst size
o.ReplenishmentPeriod = TimeSpan.FromSeconds(10);
o.TokensPerPeriod = 20; // 20 tokens every 10 seconds
o.AutoReplenishment = true;
});
This allows a burst of up to 100 requests but sustains a maximum of 120 requests per minute (20 tokens every 10 seconds).
Concurrency Limiter
Limits the number of concurrent requests rather than requests per time window:
options.AddConcurrencyLimiter("concurrency", o =>
{
o.PermitLimit = 10;
o.QueueProcessingOrder = QueueProcessingOrder.OldestFirst;
o.QueueLimit = 5;
});
This is useful for endpoints that perform expensive operations — database-heavy reports, file processing, or external API calls.
Per-Client Rate Limiting
The examples above apply a single limit across all clients. In practice, you want per-client limits based on API key, user identity, or IP address:
builder.Services.AddRateLimiter(options =>
{
options.RejectionStatusCode = StatusCodes.Status429TooManyRequests;
options.AddPolicy("per-user", httpContext =>
{
var userId = httpContext.User.FindFirstValue(ClaimTypes.NameIdentifier)
?? httpContext.Connection.RemoteIpAddress?.ToString()
?? "anonymous";
return RateLimitPartition.GetFixedWindowLimiter(userId, _ =>
new FixedWindowRateLimiterOptions
{
PermitLimit = 100,
Window = TimeSpan.FromMinutes(1)
});
});
});
Each authenticated user gets their own 100 requests/minute quota. Unauthenticated requests fall back to IP-based limiting.
Tiered Rate Limits
Different API tiers can have different limits:
options.AddPolicy("tiered", httpContext =>
{
var tier = httpContext.User.FindFirstValue("subscription_tier") ?? "free";
return RateLimitPartition.GetTokenBucketLimiter(
$"{tier}:{httpContext.User.Identity?.Name}",
_ => tier switch
{
"enterprise" => new TokenBucketRateLimiterOptions
{
TokenLimit = 1000,
ReplenishmentPeriod = TimeSpan.FromSeconds(10),
TokensPerPeriod = 200,
AutoReplenishment = true
},
"pro" => new TokenBucketRateLimiterOptions
{
TokenLimit = 500,
ReplenishmentPeriod = TimeSpan.FromSeconds(10),
TokensPerPeriod = 100,
AutoReplenishment = true
},
_ => new TokenBucketRateLimiterOptions
{
TokenLimit = 50,
ReplenishmentPeriod = TimeSpan.FromSeconds(10),
TokensPerPeriod = 10,
AutoReplenishment = true
}
});
});
Response Headers
Communicate rate limit status to clients using standard headers. Customise the rejection response:
options.OnRejected = async (context, ct) =>
{
context.HttpContext.Response.StatusCode = StatusCodes.Status429TooManyRequests;
if (context.Lease.TryGetMetadata(MetadataName.RetryAfter, out var retryAfter))
{
context.HttpContext.Response.Headers.RetryAfter =
((int)retryAfter.TotalSeconds).ToString();
}
await context.HttpContext.Response.WriteAsJsonAsync(new ProblemDetails
{
Title = "Too Many Requests",
Detail = "Rate limit exceeded. Please retry after the specified delay.",
Status = StatusCodes.Status429TooManyRequests
}, ct);
};
The Retry-After header tells clients exactly when they can retry, preventing unnecessary immediate retries.
Applying Rate Limits Selectively
Apply different policies to different endpoints:
app.MapGet("/api/products", GetProducts)
.RequireRateLimiting("per-user");
app.MapPost("/api/orders", CreateOrder)
.RequireRateLimiting("tiered");
// Health checks and documentation should not be rate limited
app.MapGet("/health", () => Results.Ok())
.DisableRateLimiting();
Distributed Rate Limiting
The built-in middleware stores state in memory, which means it does not work across multiple server instances. For distributed scenarios, you have two options:
- Sticky sessions — route each client to the same server instance. Simple but reduces load balancer flexibility.
- Redis-backed rate limiting — use a package like
RedisRateLimitingthat stores counters in Redis, providing shared state across instances.
For most applications, the built-in middleware is sufficient. When you scale horizontally, plan for distributed state from the start.