Modern CPUs can operate on multiple data elements simultaneously using SIMD (Single Instruction, Multiple Data) instructions. A single SIMD instruction can add, compare, or transform 4, 8, 16, or even 32 values at once. .NET exposes this capability through types in System.Numerics and System.Runtime.Intrinsics.

The Three Levels of SIMD in .NET

.NET offers three approaches, from highest to lowest level:

  1. Vector<T> — a portable, hardware-adaptive type that adjusts its width to the CPU's SIMD capabilities. Easy to use, works everywhere.
  2. Vector128<T>, Vector256<T>, Vector512<T> — fixed-width types with cross-platform intrinsics. More control, still portable.
  3. Sse2, Avx2, Arm.AdvSimd — platform-specific intrinsics. Maximum control, but requires separate code paths per architecture.

For most developers, Vector<T> or the fixed-width types are the right choice.

Vector<T>: The Simplest Path

Vector<T> processes Vector<T>.Count elements at once. On a machine with AVX2, Vector<int>.Count is 8 (256 bits / 32 bits per int). On ARM with NEON, it is 4 (128 bits).

Here is a scalar sum:

Example.cs
public static int SumScalar(int[] data)
{
    int sum = 0;
    for (int i = 0; i < data.Length; i++)
        sum += data[i];
    return sum;
}

And the vectorised version:

Example.cs
using System.Numerics;

public static int SumVector(int[] data)
{
    var sumVec = Vector<int>.Zero;
    int i = 0;
    int vectorSize = Vector<int>.Count;

    // Process chunks of Vector<int>.Count elements
    for (; i <= data.Length - vectorSize; i += vectorSize)
    {
        var vec = new Vector<int>(data, i);
        sumVec += vec;
    }

    // Sum the vector lanes
    int sum = 0;
    for (int j = 0; j < vectorSize; j++)
        sum += sumVec[j];

    // Handle remaining elements
    for (; i < data.Length; i++)
        sum += data[i];

    return sum;
}

On a machine with AVX2, the inner loop processes 8 integers per iteration instead of one.

Fixed-Width Vectors: Vector128 and Vector256

For more precise control, use the fixed-width types in System.Runtime.Intrinsics:

Example.cs
using System.Runtime.Intrinsics;

public static int SumVector128(ReadOnlySpan<int> data)
{
    var sum = Vector128<int>.Zero;
    int i = 0;

    for (; i <= data.Length - Vector128<int>.Count; i += Vector128<int>.Count)
    {
        var vec = Vector128.Create(data[i..]);
        sum = Vector128.Add(sum, vec);
    }

    int result = Vector128.Sum(sum);

    for (; i < data.Length; i++)
        result += data[i];

    return result;
}

The Vector128 and Vector256 types provide static methods like Add, Multiply, Equals, LessThan, Shuffle, and many more. These map directly to hardware instructions.

A Practical Example: Counting Characters

Counting occurrences of a character in a string is a common operation. Here is a vectorised version:

Example.cs
public static int CountChar(ReadOnlySpan<char> text, char target)
{
    int count = 0;
    int i = 0;

    if (Vector128.IsHardwareAccelerated && text.Length >= Vector128<ushort>.Count)
    {
        var targetVec = Vector128.Create((ushort)target);
        var countVec = Vector128<ushort>.Zero;

        for (; i <= text.Length - Vector128<ushort>.Count; i += Vector128<ushort>.Count)
        {
            var chunk = Vector128.LoadUnsafe(
                ref Unsafe.As<char, ushort>(ref MemoryMarshal.GetReference(text)),
                (nuint)i);

            var matches = Vector128.Equals(chunk, targetVec);
            countVec -= matches; // -1 for match, 0 for no match
        }

        // Sum all lanes
        for (int j = 0; j < Vector128<ushort>.Count; j++)
            count += countVec[j];
    }

    // Scalar fallback for remaining elements
    for (; i < text.Length; i++)
    {
        if (text[i] == target)
            count++;
    }

    return count;
}

When the JIT Does It for You

The .NET JIT compiler can auto-vectorise some simple loops. In .NET 8+, loops like this may be automatically converted to SIMD:

Example.cs
// The JIT may vectorise this automatically
public static bool Contains(ReadOnlySpan<byte> data, byte value)
{
    foreach (var b in data)
    {
        if (b == value)
            return true;
    }
    return false;
}

However, auto-vectorisation is limited. Complex loops, branches, and method calls prevent it. For guaranteed SIMD performance, write vectorised code explicitly.

Checking Hardware Support

Always check that the hardware supports the instructions you want to use:

Example.cs
if (Vector128.IsHardwareAccelerated)
{
    // Use Vector128 path
}
else
{
    // Scalar fallback
}

For platform-specific intrinsics:

Example.cs
if (Avx2.IsSupported)
{
    // AVX2 path (256-bit)
}
else if (Sse2.IsSupported)
{
    // SSE2 path (128-bit)
}
else
{
    // Scalar fallback
}

Summary

SIMD lets you process multiple data elements per instruction, delivering 2-16x speedups for numeric and data-processing workloads. Start with Vector<T> for simplicity, move to Vector128<T>/Vector256<T> for control, and only drop to platform-specific intrinsics when you need maximum performance on a known architecture. The .NET runtime itself uses SIMD extensively in string.IndexOf, Span.Contains, MemoryExtensions, and many other methods — your code can benefit from the same techniques.