Every .NET project that touches AI eventually hits the same wall. You wire up the OpenAI SDK in one service, Ollama in another for local development, and maybe Azure OpenAI in production. Each provider has its own client types, its own way of handling streaming, its own function-calling conventions. Swapping providers means rewriting calling code, and testing means either hitting a real API or building a bespoke mock. It is the ILogger problem all over again — but for large language models.
Microsoft.Extensions.AI is Microsoft's answer. It provides a set of abstractions — IChatClient, IEmbeddingGenerator<TInput, TEmbedding>, and VectorStoreCollection<TKey, TRecord> — that sit between your application code and whichever AI service you happen to be using. The libraries shipped as GA with .NET 10 and have already surpassed three million NuGet downloads. If you are writing .NET code that calls an LLM, these are the interfaces you should be coding against.
The packages
The library ships as three NuGet packages, each with a clear purpose:
- Microsoft.Extensions.AI.Abstractions — the core exchange types. This is what library authors reference when they implement a provider. It contains
IChatClient,IEmbeddingGenerator<TInput, TEmbedding>, and the content types (TextContent,UriContent,ImageContent). - Microsoft.Extensions.AI — the package most applications should reference. It pulls in the abstractions and adds middleware components: distributed caching, OpenTelemetry, logging, and automatic function invocation.
- Microsoft.Extensions.VectorData.Abstractions — exchange types for vector stores, including
VectorStoreCollection<TKey, TRecord>and search primitives.
All three target netstandard2.0, so they work on .NET 8 LTS, .NET 9, .NET 10, and even .NET Framework 4.6.2. The current stable version is 10.4.1.
// TIP
Most applications should reference Microsoft.Extensions.AI directly. It implicitly depends on Microsoft.Extensions.AI.Abstractions, so you get everything in one package reference.
IChatClient: one interface, every provider
The core abstraction is IChatClient. It defines two methods — GetResponseAsync for complete responses and GetStreamingResponseAsync for streamed tokens — and that is it. Every provider implements the same interface, so your application code never sees provider-specific types.
Here is what swapping between Ollama for development and Azure OpenAI for production looks like:
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddChatClient(sp =>
{
if (builder.Environment.IsDevelopment())
{
return new OllamaApiClient("http://localhost:11434", "qwen3");
}
return new AzureOpenAIClient(
new Uri(builder.Configuration["AzureOpenAI:Endpoint"]!),
new DefaultAzureCredential())
.GetChatClient("gpt-4.1")
.AsIChatClient();
});
The calling code never changes:
public class SummaryService(IChatClient chatClient)
{
public async Task<string> SummariseAsync(string document)
{
var response = await chatClient.GetResponseAsync(
$"Summarise the following document in three bullet points:\n\n{document}");
return response.Text;
}
}
When you need streaming — for a chat interface or real-time output — swap to GetStreamingResponseAsync:
app.MapPost("/chat", async (IChatClient chatClient, ChatRequest request) =>
{
var stream = chatClient.GetStreamingResponseAsync(request.Message);
return Results.Stream(async responseStream =>
{
var writer = new StreamWriter(responseStream);
await foreach (var update in stream)
{
await writer.WriteAsync(update.Text);
await writer.FlushAsync();
}
}, "text/plain");
});
That is the same code whether the backing model is GPT-4.1, Claude, Gemini, or a local Ollama instance. The provider is a deployment concern, not an application concern.
The middleware pipeline
The real power of Microsoft.Extensions.AI is not just the abstraction — it is the middleware model. If you have used ASP.NET Core middleware or HttpClient delegating handlers, this will feel instantly familiar. You compose a pipeline of behaviours around your chat client using a builder:
builder.Services.AddChatClient(sp =>
new AzureOpenAIClient(
new Uri(config["AzureOpenAI:Endpoint"]!),
new DefaultAzureCredential())
.GetChatClient("gpt-4.1")
.AsIChatClient())
.UseLogging()
.UseDistributedCache()
.UseOpenTelemetry();
Each .Use*() call wraps the inner client in a delegating layer. Requests flow inward through logging, then caching, then telemetry, then hit the actual provider. Responses flow back out through the same layers in reverse.
The built-in middleware covers the most common cross-cutting concerns:
| Middleware | What it does |
|---|---|
UseLogging() |
Logs requests and responses via ILogger |
UseDistributedCache() |
Caches responses using IDistributedCache |
UseOpenTelemetry() |
Emits traces and metrics for AI operations |
UseFunctionInvocation() |
Automatically invokes tool functions when the model requests them |
Writing custom middleware
You are not limited to the built-in components. The .Use() method accepts a delegate, so you can inject any behaviour:
builder.Services.AddChatClient(sp => CreateInnerClient(sp))
.Use(async (messages, options, nextAsync, cancellationToken) =>
{
using var lease = await rateLimiter.AcquireAsync(1, cancellationToken);
if (!lease.IsAcquired)
throw new InvalidOperationException("Rate limit exceeded.");
return await nextAsync(messages, options, cancellationToken);
})
.UseOpenTelemetry();
This slots a concurrency limiter into the pipeline without touching the provider or the consuming code. The pattern is identical to ASP.NET Core middleware — same mental model, same composability.
Structured output
One of the most useful features is built-in structured output. Instead of parsing JSON strings from LLM responses, you can deserialise directly to a C# type:
public record ReceiptItem(string Name, float Price);
public record Receipt(string Merchant, List<ReceiptItem> Items, float Total);
public class ReceiptProcessor(IChatClient chatClient)
{
public async Task<Receipt?> ExtractAsync(string receiptText)
{
var response = await chatClient.GetResponseAsync<Receipt>(
$"Extract structured data from this receipt:\n\n{receiptText}");
return response.TryGetResult(out var receipt) ? receipt : null;
}
}
The SDK handles the JSON schema generation, system prompt injection, and response parsing. You get a strongly-typed C# object back, or a clear indication that the model's output did not match the expected shape.
Function calling with AIFunctionFactory
When you need the model to invoke your code — to query a database, call an API, or perform a calculation — Microsoft.Extensions.AI provides AIFunctionFactory and the UseFunctionInvocation() middleware:
public class OrderAssistant(IChatClient chatClient)
{
[Description("Look up the status of an order by its ID")]
static async Task<string> GetOrderStatus(int orderId, OrderRepository repo)
{
var order = await repo.FindAsync(orderId);
return order is not null
? $"Order {orderId}: {order.Status}, shipped {order.ShippedDate:d}"
: $"Order {orderId} not found.";
}
public async Task<string> AskAsync(string question, OrderRepository repo)
{
var client = chatClient
.AsBuilder()
.UseFunctionInvocation()
.Build();
var response = await client.GetResponseAsync(
question,
new ChatOptions
{
Tools = [AIFunctionFactory.Create(GetOrderStatus, [repo])]
});
return response.Text;
}
}
The flow works like this: the model receives the function schema, decides it needs to call GetOrderStatus, the SDK invokes the method locally, feeds the result back to the model, and the model produces a natural language response. The UseFunctionInvocation() middleware handles the entire loop, including multi-step chains where the model calls several functions before producing a final answer.
// WARNING
Function invocation executes real code. Always validate inputs and scope the functions the model can call to the minimum required for the task. Do not expose administrative or destructive operations as AI functions without explicit safeguards.
Embeddings and vector search
The embedding side of the library follows the same pattern. IEmbeddingGenerator<TInput, TEmbedding> provides a provider-agnostic interface for generating embeddings, and VectorStoreCollection<TKey, TRecord> handles storage and search:
builder.Services.AddEmbeddingGenerator(sp =>
new AzureOpenAIClient(
new Uri(config["AzureOpenAI:Endpoint"]!),
new DefaultAzureCredential())
.GetEmbeddingClient("text-embedding-3-small")
.AsIEmbeddingGenerator());
builder.Services.AddSqliteCollection<int, Product>(
"products",
"Data Source=products.db");
Define your record type with attributes that describe the schema:
public class Product
{
[VectorStoreKey]
public int Id { get; set; }
[VectorStoreData]
public required string Name { get; set; }
[VectorStoreData]
public required string Category { get; set; }
[VectorStoreVector(Dimensions: 1536)]
public string? Description { get; set; }
}
Then search with LINQ-style filter expressions:
public class ProductSearch(VectorStoreCollection<int, Product> collection)
{
public async IAsyncEnumerable<Product> SearchAsync(
string query, string category, int limit = 5)
{
await foreach (var result in collection.SearchAsync(
query,
top: limit,
new() { Filter = r => r.Category == category }))
{
yield return result.Record;
}
}
}
The embedding generation happens automatically — the VectorStoreCollection uses the registered IEmbeddingGenerator to convert the search query into a vector, then performs similarity search against the stored embeddings. Swap SQLite for Qdrant, Azure AI Search, or Cosmos DB by changing the DI registration, not the search code.
The provider ecosystem
Around 100 NuGet packages now implement these abstractions. The major players include:
- Microsoft.Extensions.AI.OpenAI — OpenAI and Azure OpenAI
- Microsoft.Extensions.AI.Ollama (via OllamaSharp) — local models
- Anthropic, Google, and HuggingFace community packages
- Vector stores — SQLite, Qdrant, Cosmos DB, Azure SQL, Azure AI Search
Agent frameworks like Semantic Kernel and AutoGen are built on top of the same interfaces, so anything you compose at the IChatClient level — caching, telemetry, rate limiting — applies transparently to agents built with those frameworks.
// NOTE
The MCP C# SDK also builds on Microsoft.Extensions.AI types for format translation between MCP and LLM message formats, so the abstractions serve as a common foundation across the .NET AI ecosystem.
Common pitfalls
Referencing provider packages directly in application code. If your service classes take OpenAIClient instead of IChatClient, you have coupled your application to a specific provider. Depend on the abstraction, configure the concrete client in your composition root.
Forgetting the middleware order matters. Middleware executes in the order you register it. If you put caching before logging, cached responses will not appear in your logs. If you put telemetry after function invocation, tool calls will not generate spans. Think about the order the same way you think about ASP.NET Core middleware.
Not using UseFunctionInvocation() with tool-calling models. If you pass Tools in ChatOptions but do not have the function invocation middleware in the pipeline, the model will return tool call requests as raw response content instead of actually invoking them. The middleware is what closes the loop.
Ignoring structured output failures. GetResponseAsync<T> returns a response that may not successfully deserialise. Always check TryGetResult rather than assuming the output will match your type. Models can return malformed JSON, especially with complex nested schemas.
Registering IEmbeddingGenerator as scoped or transient. The embedding generator typically wraps an HTTP client. Register it as a singleton (or use the AddEmbeddingGenerator extension which handles lifetime correctly) to avoid socket exhaustion.
Summary
Microsoft.Extensions.AIprovidesIChatClient,IEmbeddingGenerator, andVectorStoreCollection— three abstractions that decouple your AI code from specific providers.- The middleware pipeline (
UseLogging,UseDistributedCache,UseOpenTelemetry,UseFunctionInvocation) follows the same pattern as ASP.NET Core middleware, making it immediately familiar. - Structured output with
GetResponseAsync<T>eliminates manual JSON parsing from LLM responses. - Function calling via
AIFunctionFactoryandUseFunctionInvocation()handles multi-turn tool invocation automatically. - Vector search with
VectorStoreCollectionsupports LINQ-style filters and automatic embedding generation. - Around 100 provider packages implement the abstractions, from OpenAI and Ollama to Semantic Kernel and AutoGen.
- The packages target
netstandard2.0, so they work across .NET 8, .NET 9, .NET 10, and .NET Framework 4.6.2.