Qdrant with .NET Browse articles

A query can find the right runbook and still miss the condition that makes its instructions safe to follow. Choose chunk boundaries so retrieval can preserve those conditions. Check the returned evidence before you start tuning the generated answer.

The short notes in Indexing Your First Documents fit into one point each. That was enough to verify indexing and build the query path in Querying Qdrant from .NET. A longer runbook needs more deliberate boundaries.

We will extend that console project with a longer support runbook. Each chunk will contain complete paragraphs from one section and have a deterministic point ID. The query still uses the existing embedding setup.

A boundary can separate the answer from its condition

Consider this illustrative runbook passage:

When shutdown starts, stop accepting new work and complete the channel writer. The reader can then drain the items already in the queue.

Drain the remaining items only while the shutdown deadline allows it. If the deadline expires, cancel the active I/O and leave unfinished work for the durable retry path.

A query about draining a queue needs the deadline condition too. If the pipeline retrieves only the first paragraph, the model receives an incomplete procedure.

Making every chunk larger can keep those paragraphs together, but it can also pull unrelated material into the same point. Smaller chunks make narrower passages searchable, at the cost of context. Overlap can preserve a connection across a boundary, but it also creates repeated text in the result set.

I would start with the source structure and inspect the resulting chunks. The conceptual background is covered in RAG Is a Data Problem Before It’s a Prompt Problem. Here, the question is how to make those choices visible in the ingestion code.

Keep the existing embedding contract

Continue with QdrantCollections, .NET 10, and the local Qdrant 1.19.0 setup from Your First Qdrant Collection. To reproduce the earlier samples, keep the packages pinned to their existing baseline:

PackageVersionUsed by
Qdrant.Client1.19.0Both paths
Microsoft.Extensions.AI.OpenAI10.9.0Azure OpenAI
Azure.AI.OpenAI2.1.0Azure OpenAI
Microsoft.Extensions.AI10.9.0Ollama
OllamaSharp5.4.30Ollama

You can preview the chunks without running Qdrant or configuring an embedding provider. When you are ready to index them, start Qdrant in another terminal if its container is stopped:

docker start -a qdrant-local

For indexing and querying, the chosen collection must already exist. The Azure path uses support-documents-v1, support-text-v1, and 1,536-dimensional cosine vectors. The local path uses support-documents-embeddinggemma-v1, support-embeddinggemma-v1, and 768-dimensional cosine vectors. The indexing article includes the provider setup and environment variables if you are starting here. Both provider variants below compile with this package baseline on .NET 10.0.102. OllamaSharp may report CS9057 for its optional source generator on that SDK. This example does not use generated tools.

The embedding profile stays the same: Azure OpenAI receives plain text, and EmbeddingGemma uses its existing document and query prefixes. The chunker chooses the passages and includes a heading in each chunk’s text. Its separate version describes those construction rules. The embedding profile describes how the provider turns that text into a vector. The Azure profile embeds chunk.Text as-is. The EmbeddingGemma profile adds the model-specific document prefix and title format used by the earlier sample.

Define a small, inspectable chunking policy

For this exercise, the source adapter supplies sections with stable IDs, headings, and paragraphs. That keeps parsing separate from chunking. A Markdown or PDF adapter would need to produce that structure first.

The policy groups complete paragraphs within each section. It repeats the section heading in every chunk and allows one paragraph of overlap when the preceding window contains at least two paragraphs. It never carries text into another section.

The example uses an 80-word ceiling, including the heading. The ceiling is deliberately small so this runbook produces several chunks. Choose production chunk sizes for your corpus, and enforce the embedding token limit separately.

If a paragraph cannot fit with its heading, the chunker throws. I prefer that visible failure to silently cutting a procedure in half. A source adapter can supply smaller meaningful blocks, or a later splitter can divide prose at sentence boundaries. Tables and fenced code need their own handling.

The implementation below counts whitespace-separated words. It cannot enforce the model’s token limit. Count the complete formatted embedding input when validating tokens, including headings and provider prefixes.

EmbeddingGemma has a maximum input length of 2,048 tokens. Ollama’s embedding API defaults to truncating oversized inputs. Setting truncate: false makes it reject them instead. This example retains the earlier client’s behavior for the short controlled text below. In a production ingestion path, I would validate token counts or disable truncation before embedding. Silent truncation can remove a condition that the chunker deliberately kept beside an instruction.

Configure one provider

Replace Program.cs with one of these setup blocks, then append the shared program and helper types below. Together they form one program. The provider factory runs only when you proceed to indexing. Preview mode needs neither Azure credentials nor a running Ollama server.

Azure OpenAI

using System.Security.Cryptography;
using System.Text;
using Azure.AI.OpenAI;
using Microsoft.Extensions.AI;
using Qdrant.Client;
using Qdrant.Client.Grpc;
using System.ClientModel;
using static Qdrant.Client.Grpc.Conditions;

const string collectionName = "support-documents-v1";
const string embeddingProfileId = "support-text-v1";
const int vectorSize = 1536;

Func<IEmbeddingGenerator<string, Embedding<float>>> createGenerator = () =>
{
    string endpoint = Environment.GetEnvironmentVariable("AZURE_OPENAI_ENDPOINT")
        ?? throw new InvalidOperationException("Set AZURE_OPENAI_ENDPOINT.");
    string deployment = Environment.GetEnvironmentVariable(
        "AZURE_OPENAI_EMBEDDING_DEPLOYMENT")
        ?? throw new InvalidOperationException("Set AZURE_OPENAI_EMBEDDING_DEPLOYMENT.");
    string key = Environment.GetEnvironmentVariable("AZURE_OPENAI_API_KEY")
        ?? throw new InvalidOperationException("Set AZURE_OPENAI_API_KEY.");

    return new AzureOpenAIClient(new Uri(endpoint), new ApiKeyCredential(key))
        .GetEmbeddingClient(deployment)
        .AsIEmbeddingGenerator();
};

Func<string, string, string> prepareChunk = (title, text) => text;
Func<string, string> prepareQuery = query => query;

Ollama with EmbeddingGemma

Use this block instead if your existing collection contains EmbeddingGemma vectors. Keep the same model artifact used to index and query it.

using System.Security.Cryptography;
using System.Text;
using Microsoft.Extensions.AI;
using OllamaSharp;
using Qdrant.Client;
using Qdrant.Client.Grpc;
using static Qdrant.Client.Grpc.Conditions;

const string collectionName = "support-documents-embeddinggemma-v1";
const string embeddingProfileId = "support-embeddinggemma-v1";
const int vectorSize = 768;

Func<IEmbeddingGenerator<string, Embedding<float>>> createGenerator = () =>
    new OllamaApiClient(
        new Uri("http://localhost:11434"), "embeddinggemma:300m");

Func<string, string, string> prepareChunk = (title, text) =>
    $"title: {title} | text: {text}";
Func<string, string> prepareQuery = query =>
    $"task: search result | query: {query}";

Preview, index, and query the chunks

Append this shared block after the chosen setup. It builds and prints the chunks first. With previewOnly set to true, the program then exits before creating either client or reading Azure credentials. Set it to false to validate the collection and run the indexing and query path.

const string chunkingVersion = "section-paragraphs-80w-overlap1-v1";
bool previewOnly = true;

const string documentId = "worker-shutdown";
const string title = "Stop a background service cleanly";
SourceSection[] sections =
[
    new("shutdown", "Shutdown and the queue", [
        "When shutdown starts, stop accepting new work and complete " +
        "the channel writer. The reader can then drain the items " +
        "already in the queue.",
        "Drain the remaining items only while the shutdown deadline " +
        "allows it. If the deadline expires, cancel the active I/O " +
        "and leave unfinished work for the durable retry path.",
        "Pass the deadline cancellation token to each asynchronous " +
        "operation used during draining. A handler that ignores the " +
        "token can keep running after the shutdown budget expires."
    ]),
    new("capacity", "Queue capacity", [
        "Use a bounded Channel to limit queued work. With FullMode " +
        "set to Wait, producers await available capacity instead " +
        "of letting an in-memory backlog grow without a limit.",
        "Completing the writer prevents further writes. Producers " +
        "must handle channel completion during shutdown rather than " +
        "retrying against the same completed writer."
    ])
];

List<SourceChunk> chunks = Chunker.Build(
    documentId, sections, chunkingVersion, maxWords: 80, overlap: true);

foreach (SourceChunk chunk in chunks)
{
    Console.WriteLine($"{chunk.ChunkId} ({Chunker.WordCount(chunk.Text)} words)");
    Console.WriteLine(chunk.Text);
    Console.WriteLine();
}

if (previewOnly)
{
    return;
}

using var timeout = new CancellationTokenSource(TimeSpan.FromMinutes(2));
CancellationToken token = timeout.Token;
using var client = new QdrantClient(host: "127.0.0.1", port: 6334);

CollectionInfo collection = await client.GetCollectionInfoAsync(
    collectionName, token);
VectorsConfig? config = collection.Config.Params.VectorsConfig;
if (config is null ||
    config.ConfigCase != VectorsConfig.ConfigOneofCase.Params ||
    config.Params.Size != (ulong)vectorSize ||
    config.Params.Distance != Distance.Cosine ||
    !collection.Config.Metadata.TryGetValue(
        "embedding-profile-id", out Value? profile) ||
    profile.StringValue != embeddingProfileId)
{
    throw new InvalidOperationException("Collection contract mismatch.");
}

using IEmbeddingGenerator<string, Embedding<float>> generator = createGenerator();

foreach (SourceChunk chunk in chunks)
{
    ReadOnlyMemory<float> vector = await generator.GenerateVectorAsync(
        prepareChunk(title, chunk.Text), cancellationToken: token);
    if (vector.Length != vectorSize)
    {
        throw new InvalidOperationException("Embedding dimension mismatch.");
    }

    await client.UpsertAsync(
        collectionName: collectionName,
        points: [new PointStruct
        {
            Id = chunk.PointId,
            Vectors = vector.ToArray(),
            Payload =
            {
                ["document_id"] = documentId,
                ["chunk_id"] = chunk.ChunkId,
                ["chunking_version"] = chunkingVersion,
                ["section_id"] = chunk.SectionId,
                ["title"] = title,
                ["heading"] = chunk.Heading,
                ["text"] = chunk.Text
            }
        }],
        wait: true,
        cancellationToken: token);

    var stored = await client.RetrieveAsync(
        collectionName: collectionName,
        id: chunk.PointId,
        withPayload: true,
        withVectors: false,
        cancellationToken: token);
    if (stored.Count != 1 ||
        !stored[0].Payload.TryGetValue("text", out Value? storedText) ||
        storedText.StringValue != chunk.Text)
    {
        throw new InvalidOperationException($"Read-back failed: {chunk.ChunkId}");
    }
}

const string question = "Should I drain all queued work during shutdown?";
ReadOnlyMemory<float> queryVector = await generator.GenerateVectorAsync(
    prepareQuery(question), cancellationToken: token);
if (queryVector.Length != vectorSize)
{
    throw new InvalidOperationException("Query dimension mismatch.");
}

var results = await client.QueryAsync(
    collectionName: collectionName,
    query: queryVector.ToArray(),
    filter: new Filter
    {
        Must =
        {
            Match("document_id", new[] { documentId }),
            Match("chunking_version", new[] { chunkingVersion })
        }
    },
    limit: 3,
    payloadSelector: new[] { "chunk_id", "heading", "text" },
    vectorsSelector: false,
    cancellationToken: token);

Console.WriteLine($"Question: {question}");
foreach (ScoredPoint result in results)
{
    Console.WriteLine($"{result.Payload["chunk_id"].StringValue}: {result.Score:F4}");
    Console.WriteLine(result.Payload["text"].StringValue);
    Console.WriteLine();
}

Append these types at the end of the same file:

public sealed record SourceSection(
    string SectionId, string Heading, string[] Paragraphs);

public sealed record SourceChunk(
    Guid PointId, string ChunkId, string SectionId,
    string Heading, string Text);

public static class Chunker
{
    public static int WordCount(string text) =>
        text.Split((char[]?)null, StringSplitOptions.RemoveEmptyEntries).Length;

    public static List<SourceChunk> Build(
        string documentId, SourceSection[] sections, string version,
        int maxWords, bool overlap)
    {
        if (maxWords <= 0 || string.IsNullOrWhiteSpace(documentId) ||
            string.IsNullOrWhiteSpace(version))
        {
            throw new ArgumentException(
                "Supply a document ID, chunking version, and positive word budget.");
        }

        var chunks = new List<SourceChunk>();
        var sectionIds = new HashSet<string>(StringComparer.Ordinal);
        foreach (SourceSection section in sections)
        {
            if (string.IsNullOrWhiteSpace(section.SectionId) ||
                !sectionIds.Add(section.SectionId) ||
                string.IsNullOrWhiteSpace(section.Heading) ||
                section.Paragraphs.Length == 0 ||
                section.Paragraphs.Any(string.IsNullOrWhiteSpace))
            {
                throw new ArgumentException("Sections need unique IDs and nonempty text.");
            }

            int start = 0;
            while (start < section.Paragraphs.Length)
            {
                int end = start;
                int words = WordCount(section.Heading);
                while (end < section.Paragraphs.Length &&
                    words + WordCount(section.Paragraphs[end]) <= maxWords)
                {
                    words += WordCount(section.Paragraphs[end]);
                    end++;
                }

                if (end == start)
                {
                    throw new InvalidOperationException(
                        $"Paragraph {start} in '{section.SectionId}' exceeds the budget.");
                }

                string chunkId = $"{section.SectionId}/p{start}";
                string text = section.Heading + "\n\n" +
                    string.Join("\n\n", section.Paragraphs[start..end]);
                string identity = $"{version}\n{documentId}\n{chunkId}";
                byte[] hash = SHA256.HashData(Encoding.UTF8.GetBytes(identity));

                chunks.Add(new SourceChunk(
                    new Guid(hash.AsSpan(0, 16), bigEndian: true), chunkId,
                    section.SectionId, section.Heading, text));

                start = overlap && end < section.Paragraphs.Length && end - start > 1
                    ? end - 1
                    : end;
            }
        }

        return chunks;
    }
}

Run dotnet run with previewOnly = true. The controlled source produces these windows:

Chunk IDParagraphs includedBoundary to inspect
shutdown/p0First and second shutdown paragraphsDraining and its deadline condition stay together
shutdown/p1Second and third shutdown paragraphsThe deadline condition repeats beside cancellation guidance
capacity/p0Both capacity paragraphsQueue capacity stays in its own section

Set previewOnly = false and run again to embed, upsert, read back, and query those three chunks. The two-minute timeout is a bound for this small console exercise, including local model loading. Adjust it deliberately for a different job.

The loop writes one point at a time and reads it back so each write is easy to inspect. A larger ingestion job should batch writes and handle failures across the batch explicitly.

Give chunk identity a precise meaning

Each point ID comes from a SHA-256 hash of the chunking version, document ID, and chunk ID, truncated to 16 bytes and interpreted as a big-endian Guid. Using the explicit big-endian constructor gives the hash-to-UUID mapping a defined byte order that can be reproduced in other languages. This is a deterministic sample mapping, not an implementation of the standard UUIDv5 algorithm. The controlled identifiers contain no newlines. A general source adapter still needs an unambiguous identity encoding and a documented namespace.

shutdown/p0 means a window starting at paragraph zero in the shutdown section. Repeating the same input and policy produces the same point IDs, so Qdrant upserts replace those points. The loop still generates embeddings on every run.

These are deterministic positional IDs under one chunking version. Inserting a paragraph changes subsequent paragraph positions, so an unchanged passage can receive a different ID. Conversely, editing text at the same position can replace the point with that ID.

Changing the word budget, overlap policy, or how the chunker includes headings in chunk.Text requires a new chunkingVersion. Include that version in both point identity and payload so a query can select one strategy during a comparison.

Changing provider-specific formatting in prepareChunk or prepareQuery belongs in the embedding profile instead. This changes the embedding contract. Decide whether existing vectors need to be reindexed, even when their dimensions stay the same.

The original whole-document point remains in the collection. The query above excludes it because it has no matching chunking_version. An older query that filters only on document_id can still return it alongside the new chunks. Use the version filter throughout this exercise.

Deterministic IDs do not reconcile a changing corpus. If an edit produces fewer chunks, upserts leave obsolete points behind. A new strategy version also leaves its predecessor stored. This sample isolates those experiments through filtering. Replacing a production index needs an explicit update and deletion path.

The document update is not atomic either. If a write or embedding request fails partway through, earlier writes remain. A timeout can have the same effect. Updating an existing document under the same chunking version can leave previous and current text searchable together. A production replacement needs a strategy for partial ingestion failures as well as obsolete points.

Inspect evidence before scores

The query asks whether all queued work should drain during shutdown. Check whether its top results contain the deadline condition, not just words related to queues. Exact ranks and scores depend on the embedding provider and model version.

For this controlled example, the result loop assumes the payload shape written by the indexing loop. A shared collection with other writers needs payload validation.

The local sample retains the earlier setup without payload indexes. For larger filtered workloads, create keyword payload indexes on document_id and chunking_version. This also matters on Qdrant Cloud, where strict mode is enabled by default for new collections and filtering on unindexed payload fields can be rejected.

Try these questions and inspect the returned text:

QuestionEvidence the result should contain
Should I drain all queued work during shutdown?Draining is bounded by the shutdown deadline
What happens when the shutdown deadline expires?Cancel active I/O and leave unfinished work for the retry path
How do I keep the in-memory queue from growing?A bounded channel and producers waiting for capacity

Then compare the same source with maxWords: 40, overlap: false and a new version such as section-paragraphs-40w-overlap0-v1. Preview it first. The first two shutdown paragraphs now occupy separate chunks. Record whether the retrieved set still contains both the action and its condition.

Keep the model, query text, filters, and result limit constant between runs. Changing several of those at once makes a boundary comparison hard to interpret. Record the chunking version, returned chunk IDs, and whether the required evidence appeared. This is a small inspection exercise, not a retrieval benchmark.

Overlap may return the same deadline paragraph twice. That can help preserve local context, but it can also spend two result slots on similar text. More retrieved chunks do not necessarily mean more distinct evidence. Inspect the set and leave deduplication and the final context budget to the context-assembly stage.

When to use this chunking path

Use this approach for a small corpus whose source adapter already exposes meaningful sections and short paragraphs. It gives you explicit boundaries, deterministic point IDs, and a way to compare chunking policies through the existing .NET query path.

Do not apply it unchanged to arbitrary PDFs, long code blocks, tables, or continuously changing sources. Those need suitable parsing, token-aware input checks, and update or deletion handling. A short note that already contains one complete answer can stay as one chunk.

Sources