UTF-8 String Literals: Zero-Allocation Text for High-Performance Code

.NET strings are UTF-16 internally, but the modern web runs on UTF-8. Every time you write a string to an HTTP response, serialise JSON, or send data over a socket, there is an encoding conversion happening. C# 11's UTF-8 string literals eliminate that conversion for string constants by letting you express text directly as UTF-8 bytes.

The Basics

Append u8 to any string literal to produce a ReadOnlySpan<byte> containing the UTF-8 encoded bytes:

Example.cs
ReadOnlySpan<byte> greeting = "Hello, world"u8;

This is not a runtime conversion. The compiler embeds the UTF-8 bytes directly in the assembly metadata, so at runtime you get a span pointing at static data — no allocation, no encoding step.

Why This Matters

Consider a typical scenario in an ASP.NET application where you write a content type header:

Example.cs
// Before: UTF-16 string that must be converted to UTF-8 for the response
private const string JsonContentType = "application/json; charset=utf-8";

// After: already in the right encoding
private static ReadOnlySpan<byte> JsonContentType
    => "application/json; charset=utf-8"u8;

In high-throughput services, these encoding conversions add up. Each one involves allocating a byte array and running the encoder. With UTF-8 literals, the bytes are simply there, ready to be written directly to the output buffer.

Practical Examples

UTF-8 literals are particularly useful when working with APIs that accept byte spans. Writing to a PipeWriter or IBufferWriter<byte> becomes more natural:

Example.cs
public void WriteJsonStart(IBufferWriter<byte> writer)
{
    ReadOnlySpan<byte> opening = "{ \"data\": ["u8;
    var span = writer.GetSpan(opening.Length);
    opening.CopyTo(span);
    writer.Advance(opening.Length);
}

They also work well for protocol implementations where you need specific byte sequences:

Example.cs
public static class HttpConstants
{
    public static ReadOnlySpan<byte> CrLf => "\r\n"u8;
    public static ReadOnlySpan<byte> Get => "GET"u8;
    public static ReadOnlySpan<byte> Http11 => "HTTP/1.1"u8;
    public static ReadOnlySpan<byte> ContentLength => "Content-Length"u8;
}

Comparing with Encoding.UTF8.GetBytes

The traditional approach requires runtime work:

Example.cs
// Allocates a new byte[] every time
byte[] bytes = Encoding.UTF8.GetBytes("Hello");

// Better, but still requires a static field and runtime initialisation
private static readonly byte[] HelloBytes = Encoding.UTF8.GetBytes("Hello");

The u8 literal is superior in both cases. It avoids the allocation entirely and the data is available immediately without any initialisation logic:

Example.cs
// No allocation, no conversion, no static field needed
ReadOnlySpan<byte> hello = "Hello"u8;

Type Semantics

A u8 literal has a natural type of ReadOnlySpan<byte>, which means it can be used anywhere a ReadOnlySpan<byte> is expected. It can also be implicitly converted to byte[] when assigned to a byte[] variable, though this does allocate:

Example.cs
ReadOnlySpan<byte> span = "data"u8;  // No allocation
byte[] array = "data"u8;              // Allocates a new byte[]

Prefer the span form unless you specifically need an array.

Raw String Literals and UTF-8

UTF-8 literals combine with raw string literals for embedding multi-line content like JSON templates:

Example.cs
ReadOnlySpan<byte> template = """
    {
        "type": "notification",
        "version": 1
    }
    """u8;

This is especially handy for test fixtures or protocol templates where you want readable multi-line text that is directly usable as bytes.

Limitations

UTF-8 string literals must be constant expressions, so you cannot use string interpolation with them:

Example.cs
// This does not compile
ReadOnlySpan<byte> message = $"Hello, {name}"u8;

For dynamic content, you still need Encoding.UTF8.GetBytes or, better yet, write directly to a buffer using Utf8Formatter or the Utf8.TryWrite method introduced in .NET 8.

Additionally, since ReadOnlySpan<byte> is a ref struct, you cannot store UTF-8 literals as instance fields. Use them as local variables, method returns via expression bodies, or method parameters.

When to Reach for u8

If you are writing performance-sensitive code that outputs text as bytes — HTTP servers, serialisers, protocol handlers, logging pipelines — UTF-8 string literals should be a default choice for any constant text. They are simpler, faster, and more expressive than the alternatives.