The IIncrementalGenerator Pipeline: Transforms, Filters and Caching

The incremental generator pipeline is what makes modern source generators fast. Instead of receiving the entire compilation and doing all the work yourself, you declare a series of transformations. Roslyn caches the output of each stage and only re-runs stages whose inputs have changed.

Understanding the pipeline operators is essential for writing generators that do not degrade IDE performance.

The Pipeline Model

An incremental generator pipeline starts with a provider — a source of data. The most common providers are:

From a provider, you chain operators to filter, transform, and combine data before finally registering an output.

Where: Filtering

The Where operator removes items from the pipeline. It is the first line of defence against unnecessary work:

Example.cs
var pipeline = context.SyntaxProvider
    .CreateSyntaxProvider(
        predicate: (node, _) => node is ClassDeclarationSyntax,
        transform: (ctx, _) => (ClassDeclarationSyntax)ctx.Node)
    .Where(cls => cls.Modifiers.Any(SyntaxKind.PublicKeyword));

The predicate in CreateSyntaxProvider is a fast syntactic filter — it runs on every node and must be cheap. The transform runs only on nodes that pass the predicate. Then Where applies a further filter.

Select: Transforming

Select transforms each item in the pipeline. Critically, Roslyn compares each output with its previous value using Equals. If the output has not changed, downstream stages are skipped entirely.

Example.cs
var classNames = pipeline
    .Select((cls, _) => cls.Identifier.Text);

This is why it is vital to extract only the data you need into simple, equatable types. If your Select returns the full ClassDeclarationSyntax, Roslyn cannot tell that nothing meaningful changed when trivial whitespace edits occur.

The Golden Rule: Use Value Types or Records

Example.cs
// Bad — SyntaxNode reference changes on every edit
var bad = provider.Select((ctx, _) => ctx.Node);

// Good — extract only the data you need
var good = provider.Select((ctx, _) =>
{
    var symbol = ctx.SemanticModel
        .GetDeclaredSymbol((ClassDeclarationSyntax)ctx.Node);
    return new ClassModel(
        symbol!.Name,
        symbol.ContainingNamespace.ToDisplayString());
});

// Use a record (or struct) with value equality
private record ClassModel(string Name, string Namespace);

Records implement Equals by comparing all fields, so Roslyn's caching works correctly. If you use a class, you must override Equals and GetHashCode yourself.

Collect: Aggregating

Collect gathers all items from an IncrementalValuesProvider<T> (many items) into a single IncrementalValueProvider<ImmutableArray<T>> (one array). Use this when your output depends on the full set of inputs:

Example.cs
var allClasses = classNames.Collect();

context.RegisterSourceOutput(allClasses, (ctx, names) =>
{
    var list = string.Join("\n", names.Select(n => $"    \"{n}\","));
    ctx.AddSource("Registry.g.cs", $$"""
        namespace Generated;

        public static class Registry
        {
            public static readonly string[] AllClasses = {
        {{list}}
            };
        }
        """);
});

Be cautious with Collect. It means the output stage re-runs whenever any item in the collection changes. Use it only when you genuinely need the full set.

Combine: Joining Pipelines

Combine merges two pipelines into a single pipeline of tuples. This is how you access multiple data sources in a single output:

Example.cs
var configAndClasses = context.AnalyzerConfigOptionsProvider
    .Combine(allClasses);

context.RegisterSourceOutput(configAndClasses, (ctx, pair) =>
{
    var (config, classes) = pair;

    config.GlobalOptions.TryGetValue(
        "build_property.RootNamespace", out var ns);

    // Use both config and class list to generate output
});

When combining a single-value provider (like config) with a multi-value provider, Roslyn only re-runs the output when either side changes.

Pipeline Execution Order

The pipeline runs in this order:

  1. Syntax predicates run first (cheapest — no semantic info).
  2. Transforms run on matching nodes (may access semantic model).
  3. Where/Select stages filter and project.
  4. Collect/Combine aggregate and join.
  5. RegisterSourceOutput emits the final code.

Each stage caches its output. If a file is edited but the pipeline output at stage 3 is identical to the previous run (because the edit was in a comment, say), stages 4 and 5 do not execute.

Debugging Pipeline Performance

If your generator is slow, the issue is almost always one of these:

Use Debugger.Launch() or logging to trace which pipeline stages are executing. If a stage runs on every keystroke even when you have not changed the relevant code, your equality implementation is the problem.

Summary

The pipeline is the heart of the incremental generator. Master Where, Select, Collect, and Combine, ensure your intermediate types implement value equality, and your generator will be fast enough to run on every keystroke without the developer noticing.