Every AI feature you ship today comes with invisible baggage. There's the API key your ops team rotates quarterly, the per-token costs that scale unpredictably with usage, the latency spike when your user's connection drops, and the privacy policy you had to rewrite because customer data now leaves the device. For many applications — think document summarisation in a desktop tool, voice transcription in a meeting app, or semantic search in a local knowledge base — the cloud dependency is the problem, not the solution.

Microsoft's Foundry Local is an attempt to remove that baggage entirely. It ships AI models as part of your application, running inference on the user's own hardware with zero cloud calls, zero per-token costs, and zero network latency. Version 1.1, announced on 12 May 2026, adds text embeddings, live audio transcription, and — critically for enterprise shops — retargets the C# SDK to netstandard2.0, making it compatible with everything from .NET Framework 4.6.1 to .NET 10.

What Foundry Local actually is

Foundry Local is not a wrapper around a REST API. The core is a platform-specific native library (.dll on Windows, .so on Linux, .dylib on macOS) that loads directly into your application's process. That library calls ONNX Runtime for model execution, with pluggable execution providers that automatically select the best available hardware — CUDA for NVIDIA GPUs, WebGPU via Dawn for cross-platform GPU access, QNN for Qualcomm NPUs, OpenVINO for Intel, and CPU as a fallback.

The C# SDK is a thin managed wrapper around this native library. You get a FoundryLocalManager singleton, a model catalogue with download/load/unload lifecycle management, and client objects for chat completions, embeddings, and audio transcription. The entire runtime adds roughly 20 MB to your application — models are downloaded separately and cached on disk.

There's also an optional embedded web server that exposes an OpenAI-compatible REST API within your process, which lets you use the official OpenAI SDK or Microsoft.Extensions.AI adapters if you prefer that programming model.

Getting started

Install the CLI first:

terminal
# Windows
winget install Microsoft.FoundryLocal

# macOS
brew tap microsoft/foundrylocal
brew install foundrylocal

The CLI is your main tool for exploring the model catalogue and testing models before you write code:

terminal
# List all available models
foundry model list

# Filter to GPU-optimised chat models
foundry model list --filter device=GPU --filter task=chat-completion

# Download and interactively chat with a model
foundry model run phi-4-mini

Models come in hardware-optimised variants. When you download a model by alias (like phi-4-mini), Foundry Local selects the variant best suited to your hardware. You can override this in code if needed.

For the SDK, add the NuGet package:

terminal
dotnet add package Microsoft.AI.Foundry.Local

On Windows, you can optionally add Microsoft.AI.Foundry.Local.WinML for WinML-based hardware acceleration and driver management via Windows Update.

Chat completions with the native SDK

The native API is the most direct way to use Foundry Local. No web server, no HTTP overhead — just in-process function calls:

Program.cs
using Microsoft.AI.Foundry.Local;
using Betalgo.Ranul.OpenAI.ObjectModels.RequestModels;

var config = new Configuration
{
    AppName = "my-assistant",
    LogLevel = Microsoft.AI.Foundry.Local.LogLevel.Information
};

await FoundryLocalManager.CreateAsync(config);
var mgr = FoundryLocalManager.Instance;

await mgr.DownloadAndRegisterEpsAsync((epName, percent) =>
    Console.WriteLine($"Registering {epName}: {percent:F0}%"));

var catalogue = await mgr.GetCatalogAsync();
var model = await catalogue.GetModelAsync("phi-4-mini");

await model.DownloadAsync(progress =>
    Console.WriteLine($"Downloading: {progress:F0}%"));
await model.LoadAsync();

Once the model is loaded, you get a chat client and stream responses:

Program.cs
var chatClient = await model.GetChatClientAsync();
chatClient.Settings.Temperature = 0.7f;
chatClient.Settings.MaxTokens = 1024;

var messages = new List<ChatMessage>
{
    new ChatMessage { Role = "system", Content = "You are a helpful coding assistant." },
    new ChatMessage { Role = "user", Content = "Explain the builder pattern in C# with an example." }
};

await foreach (var chunk in chatClient.CompleteChatStreamingAsync(messages, CancellationToken.None))
{
    Console.Write(chunk.Choices[0].Message.Content);
}

The message types come from the Betalgo.Ranul.OpenAI package, which is a transitive dependency of the SDK. This is a deliberate design choice — it means the request/response shapes are already familiar if you've used any OpenAI-compatible API.

// NOTE

The native C# SDK currently exposes only a streaming API (CompleteChatStreamingAsync). If you need a non-streaming call, use the web server integration path with the official OpenAI SDK instead.

Tool calling

Foundry Local supports function calling through the native SDK, which is essential for building agentic applications that run entirely offline:

Services/WeatherAgent.cs
var tools = new List<ToolDefinition>
{
    new ToolDefinition
    {
        Type = "function",
        Function = new FunctionDefinition
        {
            Name = "get_current_weather",
            Description = "Get the current weather for a location",
            Parameters = new PropertyDefinition
            {
                Type = "object",
                Properties = new Dictionary<string, PropertyDefinition>
                {
                    ["location"] = new PropertyDefinition
                    {
                        Type = "string",
                        Description = "City name, e.g. 'London'"
                    }
                },
                Required = ["location"]
            }
        }
    }
};

chatClient.Settings.ToolChoice = ToolChoice.Auto;

await foreach (var chunk in chatClient.CompleteChatStreamingAsync(messages, tools, CancellationToken.None))
{
    if (chunk.Choices[0].FinishReason == "tool_calls")
    {
        var toolCall = chunk.Choices[0].Message.ToolCalls[0];
        // Execute the function locally and feed the result back
    }
    else
    {
        Console.Write(chunk.Choices[0].Message.Content);
    }
}

This is where on-device AI gets genuinely interesting. Your application can reason about user requests, call local functions (database queries, file system operations, device APIs), and synthesise results — all without any data leaving the machine.

Text embeddings

Version 1.1 introduces embedding support, which unlocks semantic search, RAG pipelines, clustering, and similarity matching — all running locally. The default model is qwen3-0.6b-embedding:

Services/SemanticSearch.cs
var embeddingModel = await catalogue.GetModelAsync("qwen3-0.6b-embedding");
await embeddingModel.DownloadAsync();
await embeddingModel.LoadAsync();

var embeddingClient = await embeddingModel.GetEmbeddingClientAsync();

var response = await embeddingClient.GenerateEmbeddingAsync(
    "Dependency injection lifetime scoping in ASP.NET Core");

var vector = response.Data[0].Embedding;
Console.WriteLine($"Dimensions: {vector.Count}");

Batch embeddings work the same way, which matters when you're indexing a document corpus:

Services/SemanticSearch.cs
var documents = new[]
{
    "Transient services are created each time they are requested",
    "Scoped services are created once per request",
    "Singleton services are created once for the application lifetime"
};

var batchResponse = await embeddingClient.GenerateEmbeddingsAsync(documents);

for (var i = 0; i < batchResponse.Data.Count; i++)
{
    var embedding = batchResponse.Data[i].Embedding;
    // Store in a local vector database (SQLite with vector extensions, FAISS, etc.)
}

The embedding format follows the OpenAI embeddings API shape, so if you later decide to swap in a cloud embedding model, the migration is straightforward.

// TIP

For local RAG applications, pair Foundry Local embeddings with a lightweight vector store like SQLite with the sqlite-vec extension. The entire pipeline runs on-device with no external dependencies.

Audio transcription

Foundry Local has supported Whisper-based audio transcription since 1.0, and 1.1 adds a live streaming model (nemotron-speech-streaming-en-0.6b) optimised for real-time speech recognition.

File-based transcription is straightforward:

Services/TranscriptionService.cs
var whisper = await catalogue.GetModelAsync("whisper-small");
await whisper.DownloadAsync();
await whisper.LoadAsync();

var audioClient = await whisper.GetAudioClientAsync();
audioClient.Settings.Language = "en";
audioClient.Settings.Temperature = 0.0f;

await foreach (var chunk in audioClient.TranscribeAudioStreamingAsync(
    "meeting-recording.mp3", CancellationToken.None))
{
    Console.Write(chunk.Text);
}

The streaming transcription returns chunks as they're processed, which is useful for long recordings where you want to show progress rather than waiting for the entire file to complete.

The live streaming model (nemotron-speech-streaming-en-0.6b) is particularly impressive — NVIDIA optimised it from 2.47 GB down to 0.67 GB through quantisation and operator fusion, and it runs faster than real-time on CPU with just 0.56 seconds of algorithmic latency.

// WARNING

The live audio transcription API (LiveAudioTranscriptionSession) is new in 1.1 and the C# documentation is still catching up. The Python SDK has the most complete examples for live mic capture at the time of writing. Check the official docs for the latest C# samples.

Integrating with Microsoft.Extensions.AI

If you're building an ASP.NET Core application or want to use the IChatClient abstraction from Microsoft.Extensions.AI, Foundry Local supports this through its embedded web server:

Program.cs
using Microsoft.AI.Foundry.Local;
using Microsoft.Extensions.AI;
using OpenAI;
using System.ClientModel;

var config = new Configuration
{
    AppName = "my-web-app",
    Web = new Configuration.WebService { Urls = "http://127.0.0.1:0" }
};

await FoundryLocalManager.CreateAsync(config);
var mgr = FoundryLocalManager.Instance;

var catalogue = await mgr.GetCatalogAsync();
var model = await catalogue.GetModelAsync("phi-4-mini");
await model.DownloadAsync();
await model.LoadAsync();

await mgr.StartWebServiceAsync();

var openAiClient = new OpenAIClient(
    new ApiKeyCredential("not-needed"),
    new OpenAIClientOptions
    {
        Endpoint = new Uri(config.Web.Urls + "/v1")
    });

IChatClient chatClient = openAiClient
    .GetChatClient(model.Id)
    .AsIChatClient();

This approach starts a lightweight HTTP server inside your process that speaks the OpenAI protocol. You connect to it with the official OpenAI SDK and then use the AsIChatClient() extension method to get an IChatClient. From there, you can register it in DI, use it with any library that accepts IChatClient, and swap between local and cloud models without changing your application code.

// TIP

Pass http://127.0.0.1:0 as the web service URL to let the OS assign an available port automatically — useful when running multiple instances or in test environments.

The netstandard2.0 story

Version 1.1 retargets the Microsoft.AI.Foundry.Local package from net9.0 to netstandard2.0 (plus net8.0). This might look like a minor infrastructure change, but the implications are significant.

netstandard2.0 compatibility means Foundry Local now works with:

If you're maintaining a WPF application on .NET Framework 4.8 and want to add AI features without migrating to .NET 8, you can now do that. If you have a Unity game that needs on-device inference, that's now possible too. The SDK doesn't force you onto the latest runtime just to use local AI.

Choosing the right model

Foundry Local's catalogue includes models across several size points. Picking the right one depends on your use case and your users' hardware:

Model Size Best for
qwen2.5-0.5b ~500 MB Low-memory devices, simple tasks
phi-4-mini ~2.5 GB General-purpose, good quality/size ratio
phi-4 ~8 GB Higher quality, needs more RAM
qwen2.5-7b ~4.5 GB Strong general performance
deepseek-r1-7b ~4.5 GB Reasoning-heavy tasks
whisper-small ~500 MB Audio transcription
qwen3-0.6b-embedding ~600 MB Text embeddings

Your application should handle the case where a model variant isn't available for the user's hardware. Use model.Variants to inspect what's available and model.SelectVariant() to explicitly choose one:

Services/ModelSelector.cs
foreach (var variant in model.Variants)
{
    var deviceType = variant.Info.Runtime?.DeviceType;
    Console.WriteLine($"Variant: {variant.Info.Runtime?.DeviceType} — {variant.Id}");
}

Common pitfalls

Forgetting to download execution providers. The first call to DownloadAndRegisterEpsAsync downloads GPU-specific runtimes (CUDA, WebGPU, etc.) which can be hundreds of megabytes. Do this during application setup or a first-run experience, not when the user first triggers an AI feature.

Not handling model downloads gracefully. Models range from 500 MB to 8+ GB. Always use the progress callback to show download status, and call IsCachedAsync() to skip re-downloads. Consider downloading models in the background during installation rather than on first use.

Loading too many models simultaneously. Each loaded model consumes significant memory. Use UnloadAsync() to release models you're not actively using. A chat model and an embedding model loaded concurrently is fine; four chat models is probably not.

Assuming GPU availability. Not every user has a compatible GPU. Foundry Local falls back to CPU automatically, but CPU inference is substantially slower for large models. Test your application's responsiveness with CPU-only inference and choose model sizes accordingly.

Ignoring the singleton pattern. FoundryLocalManager is a singleton by design — calling CreateAsync twice throws. Initialise it once during application startup and access it via FoundryLocalManager.Instance everywhere else.

Hardcoding the web service port. If you use the web server integration path, hardcoding a port will fail when another instance is already using it. Use port 0 and let the OS assign one, or read the actual bound address from the configuration after starting the service.

What's next

Foundry Local 1.1 also introduces a Responses API (currently best documented for Python) that provides a higher-level abstraction combining prompt structuring, tool invocation, and output management in a single call. Vision-language models like qwen3-vl-2b-instruct are available for document understanding and visual question answering. The C# surface for these features is still maturing, but the direction is clear — Foundry Local is becoming a full inference platform, not just a chat completion wrapper.

Summary