A traditional JIT compiler makes optimisation decisions without knowing how your code actually behaves at runtime. It does not know which branches are taken most often, which virtual methods are called, or which loops are hot. Profile-Guided Optimisation (PGO) changes this by feeding runtime profiling data back into the JIT, enabling smarter decisions.

What PGO Does

PGO collects data about actual execution patterns and uses it to:

Dynamic PGO

Dynamic PGO is built into the .NET runtime. It works in two phases:

  1. Tier 0 — methods are compiled quickly with minimal optimisation. During execution, the runtime instruments them to collect profiling data (call counts, branch frequencies, type profiles).
  2. Tier 1 — when a method is called enough times, it is recompiled with full optimisations, guided by the profiling data collected in Tier 0.

Enable dynamic PGO in your project:

MyApp.csproj
<PropertyGroup>
    <TieredPGO>true</TieredPGO>
</PropertyGroup>

Or via environment variable:

DOTNET_TieredPGO=1

In .NET 8 and later, dynamic PGO is enabled by default. You do not need to opt in.

What Dynamic PGO Optimises

Guarded Devirtualisation

When profiling shows that a virtual method call consistently targets one concrete type, the JIT inserts a type check and inlines the likely target:

Example.cs
public interface IProcessor { int Process(int x); }

public class FastProcessor : IProcessor
{
    public int Process(int x) => x * 2;
}

public int RunAll(IProcessor[] processors, int value)
{
    int sum = 0;
    foreach (var p in processors)
        sum += p.Process(value); // PGO may devirtualise this
    return sum;
}

Without PGO, p.Process(value) is a virtual call — the JIT must look up the method in the vtable every time. With PGO data showing that p is always FastProcessor, the JIT generates something like:

Example.cs
// Pseudocode of what the JIT emits
if (p is FastProcessor fp)
    sum += fp.Process(value); // inlined: sum += value * 2;
else
    sum += p.Process(value); // fallback virtual call

The hot path is now fully inlined with zero overhead.

Hot/Cold Block Reordering

PGO data tells the JIT which branches are taken most often. The JIT places the hot path in a straight line (no jumps) and moves cold blocks (error handling, rare conditions) to the end of the method. This improves instruction cache utilisation and branch prediction.

Loop Optimisations

Hot loops identified by PGO receive more aggressive optimisation: loop unrolling, bounds check elimination, and better register allocation.

Measuring PGO Impact

Use the DOTNET_JitDisasmSummary environment variable to see which methods were recompiled:

DOTNET_JitDisasmSummary=1 dotnet run -c Release

For more detail, use DOTNET_JitDisasm to see the generated assembly for a specific method:

DOTNET_JitDisasm="RunAll" dotnet run -c Release

Look for [GuardedDevirtualization] annotations in the JIT output to confirm PGO-driven devirtualisation.

Static PGO with NativeAOT

For ahead-of-time compiled applications, you can use static PGO:

  1. Build an instrumented version that records profiling data.
  2. Run the instrumented version with representative workloads.
  3. Rebuild with the profiling data to produce an optimised binary.

This is more complex than dynamic PGO but produces optimal code from the first execution — there is no warm-up period.

Benchmarking PGO

A simple benchmark showing PGO's impact on interface dispatch:

PgoBenchmark.cs
[MemoryDiagnoser]
public class PgoBenchmark
{
    private readonly IProcessor[] _processors;

    public PgoBenchmark()
    {
        _processors = Enumerable.Range(0, 1000)
            .Select(_ => (IProcessor)new FastProcessor())
            .ToArray();
    }

    [Benchmark]
    public int ProcessAll()
    {
        int sum = 0;
        foreach (var p in _processors)
            sum += p.Process(42);
        return sum;
    }
}

With PGO enabled, the inner loop is devirtualised and inlined. Typical speedups range from 20-40% for interface-heavy dispatch patterns.

What PGO Cannot Do

PGO does not help when:

Summary

Dynamic PGO is one of the most significant JIT improvements in recent .NET history. It is enabled by default in .NET 8+ and requires no code changes. It speeds up virtual dispatch, interface calls, and delegate invocations by profiling actual runtime behaviour. If you are running on .NET 8 or later, you are already benefiting from it. If you are on .NET 7, enable it with TieredPGO=true and measure the difference.