You've got a product catalogue with a million rows, users who type vague search queries like "lightweight laptop for travel," and a stakeholder who's just discovered the phrase "semantic search." Until recently, your options were: bolt on a dedicated vector database, pull in a third-party extension package, or resign yourself to writing raw SQL. With EF Core 10 and 11 targeting SQL Server 2025, that's changed. Vector search is now a first-class citizen in Entity Framework Core, complete with LINQ translation, migration support, and a clean integration path alongside traditional full-text search.
This article walks through the full spectrum of vector search capabilities in EF Core — from exact distance calculations to approximate nearest-neighbour queries powered by DiskANN indexes, and finally to hybrid search that merges semantic similarity with keyword matching using Reciprocal Rank Fusion.
Setting up vector columns
SQL Server 2025 introduced the vector data type for storing embeddings — fixed-length arrays of floating-point numbers that represent the "meaning" of a piece of content. EF Core maps these through the SqlVector<float> type, with the number of dimensions specified in the column type.
You can configure vector properties using either data annotations or the fluent API:
public class Product
{
public int Id { get; set; }
public string Name { get; set; } = string.Empty;
public string Description { get; set; } = string.Empty;
[Column(TypeName = "vector(1536)")]
public SqlVector<float> DescriptionEmbedding { get; set; }
}
Or with the fluent API if you prefer keeping your entities clean:
protected override void OnModelCreating(ModelBuilder modelBuilder)
{
modelBuilder.Entity<Product>()
.Property(p => p.DescriptionEmbedding)
.HasColumnType("vector(1536)");
}
The dimension count — 1536 in these examples — must match whatever your embedding model produces. OpenAI's text-embedding-3-small outputs 1536 dimensions, while smaller models like nomic-embed-text produce 768. Get this wrong and you'll hit runtime errors when inserting data.
// TIP
The Microsoft.Extensions.AI library provides IEmbeddingGenerator<string, Embedding<float>>, a provider-agnostic abstraction for generating embeddings. Use it to avoid coupling your data access layer to a specific embedding service.
Generating and storing embeddings
Embedding generation happens outside the database. You produce vectors from your content using a model, then store them alongside the source data:
public class ProductIngestionService(
CatalogueDbContext db,
IEmbeddingGenerator<string, Embedding<float>> embeddingGenerator)
{
public async Task IngestProductAsync(string name, string description)
{
var embedding = await embeddingGenerator.GenerateVectorAsync(description);
db.Products.Add(new Product
{
Name = name,
Description = description,
DescriptionEmbedding = new SqlVector<float>(embedding)
});
await db.SaveChangesAsync();
}
}
For bulk ingestion, batch your embedding calls — most providers support generating multiple embeddings in a single request, which is significantly faster than one-at-a-time.
Exact search with VECTOR_DISTANCE
The simplest form of vector search uses EF.Functions.VectorDistance() to compute the exact distance between two vectors. This scans every row in the table, calculating the distance for each one:
public async Task<List<Product>> SearchExactAsync(string query, int limit = 10)
{
var queryVector = new SqlVector<float>(
await embeddingGenerator.GenerateVectorAsync(query));
return await db.Products
.OrderBy(p => EF.Functions.VectorDistance("cosine", p.DescriptionEmbedding, queryVector))
.Take(limit)
.ToListAsync();
}
This translates to a straightforward SQL query:
SELECT TOP(@limit) [p].[Id], [p].[Name], [p].[Description], [p].[DescriptionEmbedding]
FROM [Products] AS [p]
ORDER BY VECTOR_DISTANCE('cosine', [p].[DescriptionEmbedding], @queryVector)
Three distance metrics are available:
| Metric | Use case |
|---|---|
cosine |
Text embeddings — measures angular distance, ignoring magnitude |
euclidean |
Image embeddings — measures straight-line distance (L2 norm) |
dot |
Normalised embeddings where you want raw inner product similarity |
Exact search is perfectly fine for small datasets — a few thousand rows will return in milliseconds. But it doesn't scale. At a million rows with 1536-dimensional vectors, you're asking SQL Server to perform a billion floating-point operations per query. That's where approximate search comes in.
Approximate search with VECTOR_SEARCH and DiskANN
SQL Server 2025 supports approximate nearest-neighbour (ANN) search through vector indexes built on the DiskANN algorithm. DiskANN constructs a graph structure over your vectors that allows the database to navigate to approximate matches without scanning every row.
// WARNING
VECTOR_SEARCH() and vector indexes are experimental features in SQL Server 2025. The EF Core APIs wrapping them are also subject to change.
Creating a vector index
Configure the index in your OnModelCreating:
protected override void OnModelCreating(ModelBuilder modelBuilder)
{
modelBuilder.Entity<Product>()
.Property(p => p.DescriptionEmbedding)
.HasColumnType("vector(1536)");
modelBuilder.Entity<Product>()
.HasVectorIndex(p => p.DescriptionEmbedding, "cosine");
}
When you add a migration, EF generates:
CREATE VECTOR INDEX [IX_Products_DescriptionEmbedding]
ON [Products] ([DescriptionEmbedding])
WITH (METRIC = COSINE)
The metric you choose for the index must match the metric you use in your search queries. Pick cosine for text embeddings in most cases.
Querying with VectorSearch
Once the index exists, use the VectorSearch() extension method instead of VectorDistance():
public async Task<List<ProductSearchResult>> SearchApproximateAsync(
string query, int topN = 10)
{
var queryVector = new SqlVector<float>(
await embeddingGenerator.GenerateVectorAsync(query));
return await db.Products
.VectorSearch(p => p.DescriptionEmbedding, "cosine", queryVector, topN: topN)
.Select(r => new ProductSearchResult
{
Product = r.Value,
Distance = r.Distance
})
.ToListAsync();
}
This translates to:
SELECT [v].[Id], [v].[Name], [v].[Description], [v].[DescriptionEmbedding]
FROM VECTOR_SEARCH([Products], 'DescriptionEmbedding', @queryVector, 'metric = cosine', @topN)
VectorSearch() returns VectorSearchResult<TEntity>, giving you both the entity through r.Value and the computed distance through r.Distance. You can filter on distance to enforce a similarity threshold:
var results = await db.Products
.VectorSearch(p => p.DescriptionEmbedding, "cosine", queryVector, topN: 50)
.Where(r => r.Distance < 0.3)
.Select(r => new { r.Value.Name, r.Distance })
.ToListAsync();
// NOTE
DiskANN is a disk-friendly algorithm — it uses SSDs efficiently with minimal memory overhead. This means your vector index can handle datasets far larger than available RAM, unlike purely in-memory approaches.
Full-text search with FreeTextTable and ContainsTable
Before diving into hybrid search, it's worth understanding the full-text search improvements in EF Core 11. While EF.Functions.FreeText() and EF.Functions.Contains() have been available for years as Where() predicates, they don't return ranking scores. EF Core 11 introduces their table-valued function counterparts.
First, configure a full-text catalog and index:
protected override void OnModelCreating(ModelBuilder modelBuilder)
{
modelBuilder.HasFullTextCatalog("ftCatalogue");
modelBuilder.Entity<Product>()
.HasFullTextIndex(p => new { p.Name, p.Description })
.HasKeyIndex("PK_Products")
.OnCatalog("ftCatalogue");
}
This generates the full-text infrastructure in a migration — no more hand-writing CREATE FULLTEXT CATALOG in raw SQL blocks. Then query using FreeTextTable():
public async Task<List<ProductSearchResult>> SearchFullTextAsync(string query, int topN = 10)
{
return await db.Products
.FreeTextTable(p => new { p.Name, p.Description }, query, topN: topN)
.Select(r => new ProductSearchResult
{
Product = r.Value,
Rank = r.Rank
})
.OrderByDescending(r => r.Rank)
.ToListAsync();
}
FreeTextTable() returns FullTextSearchResult<TEntity>, providing both the entity and a relevance rank. ContainsTable() works identically but supports the richer CONTAINS query syntax — prefix matching, proximity search, and weighted terms.
Hybrid search: combining vectors and full-text
This is where things get interesting. Vector search excels at finding semantically similar content — "lightweight laptop for travel" will match products described as "portable ultrabook for business trips." Full-text search excels at exact keyword matching — if the user types a specific model number, vector search might miss it entirely.
Hybrid search combines both approaches and uses Reciprocal Rank Fusion (RRF) to merge the results into a single ranked list:
public async Task<List<Product>> SearchHybridAsync(string query, int topN = 20)
{
var queryVector = new SqlVector<float>(
await embeddingGenerator.GenerateVectorAsync(query));
int k = topN;
return await db.Products
.FreeTextTable<Product, int>(query, topN: k)
.LeftJoin(
db.Products.VectorSearch(
p => p.DescriptionEmbedding, queryVector, "cosine", topN: k),
fts => fts.Key,
vs => vs.Value.Id,
(fts, vs) => new
{
Product = vs.Value,
FullTextRank = fts.Rank,
VectorDistance = (double?)vs.Distance
})
.Select(x => new
{
x.Product,
RrfScore = (1.0 / (k + x.FullTextRank))
+ (1.0 / (k + x.VectorDistance) ?? 0.0)
})
.OrderByDescending(x => x.RrfScore)
.Take(10)
.Select(x => x.Product)
.ToListAsync();
}
The RRF formula 1/(k + rank) normalises the scores from both search methods onto a comparable scale. Results that rank highly in both searches bubble to the top; results that are strong in only one method still appear but with a lower combined score.
// IMPORTANT
This query uses a LEFT JOIN, which means results with a high full-text rank but no vector match are included, but results with a high vector score and no full-text match are dropped. A FULL OUTER JOIN would be more appropriate here, but EF Core doesn't support it yet — track dotnet/efcore#37633 if this matters to you.
Common pitfalls
Mismatched dimensions. If your embedding model produces 768-dimensional vectors and your column is defined as vector(1536), inserts will fail at runtime. Always verify the dimension count matches your model's output before running migrations.
Using VectorSearch without an index. Calling VectorSearch() on a table without a vector index will throw a SQL Server error. Use VectorDistance() for exact search when you don't have an index, and VectorSearch() only after creating one with HasVectorIndex().
Wrong distance metric. The metric on your vector index must match the metric in your search query. If you create an index with "cosine" and query with "euclidean", you'll get an error. Decide on a metric early and stick with it.
Forgetting to embed the query. This sounds obvious, but it's easy to accidentally pass raw text to VectorSearch() instead of generating an embedding first. The query must be a SqlVector<float>, not a string.
Mixing exact and approximate search. VectorDistance() and VectorSearch() serve different purposes. VectorDistance() gives perfect results but scales poorly. VectorSearch() scales well but returns approximate results. Don't use VectorDistance() "just to be safe" on a large table — use the appropriate method for your dataset size.
Ignoring the full-text index setup. If you're building hybrid search, the full-text catalog and index must exist before you can use FreeTextTable() or ContainsTable(). EF Core 11 lets you configure these in OnModelCreating, but you still need to apply the migration.
Summary
SqlVector<float>maps .NET properties to SQL Server'svectordata type, with dimensions configured throughHasColumnType("vector(n)")or the[Column]attribute.EF.Functions.VectorDistance()performs exact distance calculations — accurate but slow at scale.HasVectorIndex()creates a DiskANN-backed vector index through EF migrations, enabling approximate nearest-neighbour search.VectorSearch()queries against a vector index usingVECTOR_SEARCH(), returning both entities and distance scores throughVectorSearchResult<TEntity>.FreeTextTable()andContainsTable()are new in EF Core 11, providing ranked full-text search results throughFullTextSearchResult<TEntity>.- Hybrid search combines vector and full-text results using Reciprocal Rank Fusion, giving you the best of semantic and keyword matching in a single LINQ query.
- Vector search requires SQL Server 2025 and EF Core 10+ (with vector indexes and
VECTOR_SEARCH()requiring EF Core 11).