A traditional JIT compiler makes optimisation decisions without knowing how your code actually behaves at runtime. It does not know which branches are taken most often, which virtual methods are called, or which loops are hot. Profile-Guided Optimisation (PGO) changes this by feeding runtime profiling data back into the JIT, enabling smarter decisions.
What PGO Does
PGO collects data about actual execution patterns and uses it to:
- Inline hot methods more aggressively.
- Optimise hot loops with better register allocation.
- Devirtualise calls when profiling shows a single concrete type is always used.
- Reorder branches so the common path is predicted correctly.
- Guide tiered compilation to focus optimisation effort where it matters.
Dynamic PGO
Dynamic PGO is built into the .NET runtime. It works in two phases:
- Tier 0 — methods are compiled quickly with minimal optimisation. During execution, the runtime instruments them to collect profiling data (call counts, branch frequencies, type profiles).
- Tier 1 — when a method is called enough times, it is recompiled with full optimisations, guided by the profiling data collected in Tier 0.
Enable dynamic PGO in your project:
<PropertyGroup>
<TieredPGO>true</TieredPGO>
</PropertyGroup>
Or via environment variable:
DOTNET_TieredPGO=1
In .NET 8 and later, dynamic PGO is enabled by default. You do not need to opt in.
What Dynamic PGO Optimises
Guarded Devirtualisation
When profiling shows that a virtual method call consistently targets one concrete type, the JIT inserts a type check and inlines the likely target:
public interface IProcessor { int Process(int x); }
public class FastProcessor : IProcessor
{
public int Process(int x) => x * 2;
}
public int RunAll(IProcessor[] processors, int value)
{
int sum = 0;
foreach (var p in processors)
sum += p.Process(value); // PGO may devirtualise this
return sum;
}
Without PGO, p.Process(value) is a virtual call — the JIT must look up the method in the vtable every time. With PGO data showing that p is always FastProcessor, the JIT generates something like:
// Pseudocode of what the JIT emits
if (p is FastProcessor fp)
sum += fp.Process(value); // inlined: sum += value * 2;
else
sum += p.Process(value); // fallback virtual call
The hot path is now fully inlined with zero overhead.
Hot/Cold Block Reordering
PGO data tells the JIT which branches are taken most often. The JIT places the hot path in a straight line (no jumps) and moves cold blocks (error handling, rare conditions) to the end of the method. This improves instruction cache utilisation and branch prediction.
Loop Optimisations
Hot loops identified by PGO receive more aggressive optimisation: loop unrolling, bounds check elimination, and better register allocation.
Measuring PGO Impact
Use the DOTNET_JitDisasmSummary environment variable to see which methods were recompiled:
DOTNET_JitDisasmSummary=1 dotnet run -c Release
For more detail, use DOTNET_JitDisasm to see the generated assembly for a specific method:
DOTNET_JitDisasm="RunAll" dotnet run -c Release
Look for [GuardedDevirtualization] annotations in the JIT output to confirm PGO-driven devirtualisation.
Static PGO with NativeAOT
For ahead-of-time compiled applications, you can use static PGO:
- Build an instrumented version that records profiling data.
- Run the instrumented version with representative workloads.
- Rebuild with the profiling data to produce an optimised binary.
This is more complex than dynamic PGO but produces optimal code from the first execution — there is no warm-up period.
Benchmarking PGO
A simple benchmark showing PGO's impact on interface dispatch:
[MemoryDiagnoser]
public class PgoBenchmark
{
private readonly IProcessor[] _processors;
public PgoBenchmark()
{
_processors = Enumerable.Range(0, 1000)
.Select(_ => (IProcessor)new FastProcessor())
.ToArray();
}
[Benchmark]
public int ProcessAll()
{
int sum = 0;
foreach (var p in _processors)
sum += p.Process(42);
return sum;
}
}
With PGO enabled, the inner loop is devirtualised and inlined. Typical speedups range from 20-40% for interface-heavy dispatch patterns.
What PGO Cannot Do
PGO does not help when:
- The hot path changes. If different code paths are hot at different times, the profiling data from Tier 0 may not reflect steady-state behaviour.
- Methods are only called once. There is nothing to profile.
- Code is already simple. Scalar arithmetic and direct method calls are already well-optimised.
Summary
Dynamic PGO is one of the most significant JIT improvements in recent .NET history. It is enabled by default in .NET 8+ and requires no code changes. It speeds up virtual dispatch, interface calls, and delegate invocations by profiling actual runtime behaviour. If you are running on .NET 8 or later, you are already benefiting from it. If you are on .NET 7, enable it with TieredPGO=true and measure the difference.