UTF-8 String Literals: Zero-Allocation Text for High-Performance Code
.NET strings are UTF-16 internally, but the modern web runs on UTF-8. Every time you write a string to an HTTP response, serialise JSON, or send data over a socket, there is an encoding conversion happening. C# 11's UTF-8 string literals eliminate that conversion for string constants by letting you express text directly as UTF-8 bytes.
The Basics
Append u8 to any string literal to produce a ReadOnlySpan<byte> containing the UTF-8 encoded bytes:
ReadOnlySpan<byte> greeting = "Hello, world"u8;
This is not a runtime conversion. The compiler embeds the UTF-8 bytes directly in the assembly metadata, so at runtime you get a span pointing at static data — no allocation, no encoding step.
Why This Matters
Consider a typical scenario in an ASP.NET application where you write a content type header:
// Before: UTF-16 string that must be converted to UTF-8 for the response
private const string JsonContentType = "application/json; charset=utf-8";
// After: already in the right encoding
private static ReadOnlySpan<byte> JsonContentType
=> "application/json; charset=utf-8"u8;
In high-throughput services, these encoding conversions add up. Each one involves allocating a byte array and running the encoder. With UTF-8 literals, the bytes are simply there, ready to be written directly to the output buffer.
Practical Examples
UTF-8 literals are particularly useful when working with APIs that accept byte spans. Writing to a PipeWriter or IBufferWriter<byte> becomes more natural:
public void WriteJsonStart(IBufferWriter<byte> writer)
{
ReadOnlySpan<byte> opening = "{ \"data\": ["u8;
var span = writer.GetSpan(opening.Length);
opening.CopyTo(span);
writer.Advance(opening.Length);
}
They also work well for protocol implementations where you need specific byte sequences:
public static class HttpConstants
{
public static ReadOnlySpan<byte> CrLf => "\r\n"u8;
public static ReadOnlySpan<byte> Get => "GET"u8;
public static ReadOnlySpan<byte> Http11 => "HTTP/1.1"u8;
public static ReadOnlySpan<byte> ContentLength => "Content-Length"u8;
}
Comparing with Encoding.UTF8.GetBytes
The traditional approach requires runtime work:
// Allocates a new byte[] every time
byte[] bytes = Encoding.UTF8.GetBytes("Hello");
// Better, but still requires a static field and runtime initialisation
private static readonly byte[] HelloBytes = Encoding.UTF8.GetBytes("Hello");
The u8 literal is superior in both cases. It avoids the allocation entirely and the data is available immediately without any initialisation logic:
// No allocation, no conversion, no static field needed
ReadOnlySpan<byte> hello = "Hello"u8;
Type Semantics
A u8 literal has a natural type of ReadOnlySpan<byte>, which means it can be used anywhere a ReadOnlySpan<byte> is expected. It can also be implicitly converted to byte[] when assigned to a byte[] variable, though this does allocate:
ReadOnlySpan<byte> span = "data"u8; // No allocation
byte[] array = "data"u8; // Allocates a new byte[]
Prefer the span form unless you specifically need an array.
Raw String Literals and UTF-8
UTF-8 literals combine with raw string literals for embedding multi-line content like JSON templates:
ReadOnlySpan<byte> template = """
{
"type": "notification",
"version": 1
}
"""u8;
This is especially handy for test fixtures or protocol templates where you want readable multi-line text that is directly usable as bytes.
Limitations
UTF-8 string literals must be constant expressions, so you cannot use string interpolation with them:
// This does not compile
ReadOnlySpan<byte> message = $"Hello, {name}"u8;
For dynamic content, you still need Encoding.UTF8.GetBytes or, better yet, write directly to a buffer using Utf8Formatter or the Utf8.TryWrite method introduced in .NET 8.
Additionally, since ReadOnlySpan<byte> is a ref struct, you cannot store UTF-8 literals as instance fields. Use them as local variables, method returns via expression bodies, or method parameters.
When to Reach for u8
If you are writing performance-sensitive code that outputs text as bytes — HTTP servers, serialisers, protocol handlers, logging pipelines — UTF-8 string literals should be a default choice for any constant text. They are simpler, faster, and more expressive than the alternatives.