<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Lukas Walter </title>
    <description>The latest articles on DEV Community by Lukas Walter  (@lukaswalter).</description>
    <link>https://dev.to/lukaswalter</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3783973%2F8171c4c5-d69c-4059-b5d9-7b7af32a8962.png</url>
      <title>DEV Community: Lukas Walter </title>
      <link>https://dev.to/lukaswalter</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/lukaswalter"/>
    <language>en</language>
    <item>
      <title>Querying Qdrant from .NET</title>
      <dc:creator>Lukas Walter </dc:creator>
      <pubDate>Tue, 29 Sep 2026 15:30:00 +0000</pubDate>
      <link>https://dev.to/lukaswalter/querying-qdrant-from-net-20a4</link>
      <guid>https://dev.to/lukaswalter/querying-qdrant-from-net-20a4</guid>
      <description>&lt;p&gt;This example turns a question into an embedding, applies a payload filter, and asks Qdrant for up to three closest eligible points. It then prints the stored text and score for each result.&lt;/p&gt;

&lt;p&gt;The code continues the same &lt;code&gt;QdrantCollections&lt;/code&gt; console project and three support notes from &lt;a href="https://www.lukaswalter.dev/posts/indexing-your-first-documents/" rel="noopener noreferrer"&gt;Indexing Your First Documents&lt;/a&gt;. It reuses the collection, embedding profile, point IDs, and payload fields from that example. This time, the program reads instead of writing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Continue from the indexed collection
&lt;/h2&gt;

&lt;p&gt;Start the Qdrant container from &lt;a href="https://www.lukaswalter.dev/posts/your-first-qdrant-collection/" rel="noopener noreferrer"&gt;Your First Qdrant Collection&lt;/a&gt; if it is not already running:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker start &lt;span class="nt"&gt;-a&lt;/span&gt; qdrant-local
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep it running and use another terminal for the .NET application. The collection should contain the three points added at the end of the indexing article:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Document ID&lt;/th&gt;
&lt;th&gt;Title&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;worker-shutdown&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Stop a background service cleanly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;http-client-lifetime&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Create HTTP clients through the factory&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;bounded-work-queue&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Limit queued background work&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If the collection is empty, run the indexing program first. This query example does not create the collection or add points.&lt;/p&gt;

&lt;p&gt;Continue with the same embedding setup that produced the stored vectors:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Setting&lt;/th&gt;
&lt;th&gt;Azure OpenAI&lt;/th&gt;
&lt;th&gt;Ollama locally&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Model&lt;/td&gt;
&lt;td&gt;&lt;code&gt;text-embedding-3-small&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;embeddinggemma:300m&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dimensions&lt;/td&gt;
&lt;td&gt;1,536&lt;/td&gt;
&lt;td&gt;768&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Collection&lt;/td&gt;
&lt;td&gt;&lt;code&gt;support-documents-v1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;support-documents-embeddinggemma-v1&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Profile ID&lt;/td&gt;
&lt;td&gt;&lt;code&gt;support-text-v1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;support-embeddinggemma-v1&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Query format&lt;/td&gt;
&lt;td&gt;Plain question text&lt;/td&gt;
&lt;td&gt;`task: search result \&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The query vector and stored document vectors must belong to the same embedding space. A 768-dimensional EmbeddingGemma query cannot search the 1,536-dimensional Azure OpenAI collection. Matching dimensions would not be enough either. Two models can return the same number of values while producing incompatible vectors.&lt;/p&gt;

&lt;h2&gt;
  
  
  Configure the query embedding
&lt;/h2&gt;

&lt;p&gt;Replace {% raw %}&lt;code&gt;Program.cs&lt;/code&gt; with one of the following setup blocks. Then append the shared query code from the next section.&lt;/p&gt;

&lt;h3&gt;
  
  
  Azure OpenAI
&lt;/h3&gt;

&lt;p&gt;Use this block if you indexed the notes into &lt;code&gt;support-documents-v1&lt;/code&gt;. It reuses the packages, deployment, and environment variables from the indexing article.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Microsoft.Extensions.AI&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Azure.AI.OpenAI&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;System.ClientModel&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Qdrant.Client&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Qdrant.Client.Grpc&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;static&lt;/span&gt; &lt;span class="n"&gt;Qdrant&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Grpc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Conditions&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;collectionName&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"support-documents-v1"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;embeddingProfileId&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"support-text-v1"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;vectorSize&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="m"&gt;1536&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;endpoint&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Environment&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GetEnvironmentVariable&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"AZURE_OPENAI_ENDPOINT"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;??&lt;/span&gt; &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;InvalidOperationException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Set AZURE_OPENAI_ENDPOINT."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;deploymentName&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Environment&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GetEnvironmentVariable&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="s"&gt;"AZURE_OPENAI_EMBEDDING_DEPLOYMENT"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;??&lt;/span&gt; &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;InvalidOperationException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Set AZURE_OPENAI_EMBEDDING_DEPLOYMENT."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;apiKey&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Environment&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GetEnvironmentVariable&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"AZURE_OPENAI_API_KEY"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;??&lt;/span&gt; &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;InvalidOperationException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Set AZURE_OPENAI_API_KEY."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;IEmbeddingGenerator&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Embedding&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;float&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;generator&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt;
    &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;AzureOpenAIClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;Uri&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;endpoint&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;ApiKeyCredential&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;apiKey&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
        &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GetEmbeddingClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;deploymentName&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;AsIEmbeddingGenerator&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="n"&gt;Func&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;prepareQuery&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The query is embedded as plain text because the documents were embedded as plain text. Keep the deployment on the same &lt;code&gt;text-embedding-3-small&lt;/code&gt; model and compatible model version used during indexing.&lt;/p&gt;

&lt;h3&gt;
  
  
  Ollama with EmbeddingGemma
&lt;/h3&gt;

&lt;p&gt;Use this block if you indexed the notes into &lt;code&gt;support-documents-embeddinggemma-v1&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Microsoft.Extensions.AI&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;OllamaSharp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Qdrant.Client&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Qdrant.Client.Grpc&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;static&lt;/span&gt; &lt;span class="n"&gt;Qdrant&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Grpc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Conditions&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;collectionName&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"support-documents-embeddinggemma-v1"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;embeddingProfileId&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"support-embeddinggemma-v1"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;vectorSize&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="m"&gt;768&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;IEmbeddingGenerator&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Embedding&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;float&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;generator&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt;
    &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;OllamaApiClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;Uri&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"http://localhost:11434"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="s"&gt;"embeddinggemma:300m"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="n"&gt;Func&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;prepareQuery&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt;
    &lt;span class="s"&gt;$"task: search result | query: &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The indexing example formatted EmbeddingGemma documents as &lt;code&gt;title: ... | text: ...&lt;/code&gt;. Retrieval queries use the corresponding &lt;code&gt;task: search result | query: ...&lt;/code&gt; format from the model card. This is not decorative prompt text. It is part of the embedding profile and affects the vector the model produces.&lt;/p&gt;

&lt;h2&gt;
  
  
  Embed the question and query Qdrant
&lt;/h2&gt;

&lt;p&gt;Append this shared code after the chosen setup block:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;var&lt;/span&gt; &lt;span class="n"&gt;executionTimeout&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt;
    &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;CancellationTokenSource&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;TimeSpan&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;FromMinutes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="m"&gt;2&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;span class="n"&gt;CancellationToken&lt;/span&gt; &lt;span class="n"&gt;token&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;executionTimeout&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Token&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;var&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;QdrantClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;host&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"127.0.0.1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;port&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="m"&gt;6334&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="n"&gt;CollectionInfo&lt;/span&gt; &lt;span class="n"&gt;collection&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GetCollectionInfoAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;collectionName&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="n"&gt;VectorsConfig&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="n"&gt;config&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;collection&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Params&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;VectorsConfig&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;config&lt;/span&gt; &lt;span class="k"&gt;is&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt; &lt;span class="p"&gt;||&lt;/span&gt;
    &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ConfigCase&lt;/span&gt; &lt;span class="p"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;VectorsConfig&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ConfigOneofCase&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Params&lt;/span&gt; &lt;span class="p"&gt;||&lt;/span&gt;
    &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Params&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Size&lt;/span&gt; &lt;span class="p"&gt;!=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;ulong&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="n"&gt;vectorSize&lt;/span&gt; &lt;span class="p"&gt;||&lt;/span&gt;
    &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Params&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Distance&lt;/span&gt; &lt;span class="p"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;Distance&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Cosine&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;InvalidOperationException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="s"&gt;$"Collection '&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;collectionName&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;' does not use the expected "&lt;/span&gt; &lt;span class="p"&gt;+&lt;/span&gt;
        &lt;span class="s"&gt;$"unnamed &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;vectorSize&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;-dimension cosine vector."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(!&lt;/span&gt;&lt;span class="n"&gt;collection&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Metadata&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;TryGetValue&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="s"&gt;"embedding-profile-id"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;out&lt;/span&gt; &lt;span class="n"&gt;Value&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="n"&gt;storedProfile&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;||&lt;/span&gt;
    &lt;span class="n"&gt;storedProfile&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;StringValue&lt;/span&gt; &lt;span class="p"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;embeddingProfileId&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;InvalidOperationException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="s"&gt;$"Collection '&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;collectionName&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;' does not match profile '&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;embeddingProfileId&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;'."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;question&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt;
    &lt;span class="s"&gt;"How do I stop a background service without abandoning I/O?"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="n"&gt;ReadOnlyMemory&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;float&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;queryVector&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;generator&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GenerateVectorAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="nf"&gt;prepareQuery&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;queryVector&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Length&lt;/span&gt; &lt;span class="p"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;vectorSize&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;InvalidOperationException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="s"&gt;$"Expected &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;vectorSize&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt; values, received &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;queryVector&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Length&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt; &lt;span class="n"&gt;eligibleDocumentIds&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"worker-shutdown"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"bounded-work-queue"&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;
&lt;span class="n"&gt;Filter&lt;/span&gt; &lt;span class="n"&gt;eligibilityFilter&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;Match&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"document_id"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;eligibleDocumentIds&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="n"&gt;IReadOnlyList&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;ScoredPoint&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;results&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;QueryAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;collectionName&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;collectionName&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;queryVector&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ToArray&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="n"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;eligibilityFilter&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="m"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;payloadSelector&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="s"&gt;"document_id"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"text"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="n"&gt;vectorsSelector&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="n"&gt;Console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;WriteLine&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;$"Question: &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;foreach&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ScoredPoint&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(!&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Payload&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;TryGetValue&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="s"&gt;"document_id"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;out&lt;/span&gt; &lt;span class="n"&gt;Value&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="n"&gt;documentId&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;||&lt;/span&gt;
        &lt;span class="p"&gt;!&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Payload&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;TryGetValue&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;out&lt;/span&gt; &lt;span class="n"&gt;Value&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;||&lt;/span&gt;
        &lt;span class="p"&gt;!&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Payload&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;TryGetValue&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;out&lt;/span&gt; &lt;span class="n"&gt;Value&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;InvalidOperationException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="s"&gt;$"Point '&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Id&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;' is missing the expected payload."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="n"&gt;Console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;WriteLine&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="n"&gt;Console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;WriteLine&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;$"Score: &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Score&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;F4&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="n"&gt;Console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;WriteLine&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;$"Document: &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;documentId&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;StringValue&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="n"&gt;Console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;WriteLine&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;$"Title: &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;StringValue&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="n"&gt;Console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;WriteLine&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;$"Text: &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;StringValue&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run the program:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dotnet run
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;worker-shutdown&lt;/code&gt; is expected to rank first for this sample. I have not hardcoded the scores because they can differ between embedding providers and model versions. With the current filter, &lt;code&gt;http-client-lifetime&lt;/code&gt; cannot appear even if its vector is close to the query.&lt;/p&gt;

&lt;p&gt;The metadata check above has a deliberate limit. The profile ID says which embedding profile the collection is supposed to contain. It does not verify the model artifact that produced the stored vectors. That distinction matters for mutable model tags such as &lt;code&gt;embeddinggemma:300m&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the query does
&lt;/h2&gt;

&lt;p&gt;The read path is short:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Validate the expected unnamed vector size and distance, plus the declared embedding profile.&lt;/li&gt;
&lt;li&gt;Prepare the question according to that profile.&lt;/li&gt;
&lt;li&gt;Generate one query vector.&lt;/li&gt;
&lt;li&gt;Apply the payload eligibility constraint during vector search.&lt;/li&gt;
&lt;li&gt;Return the highest-scoring eligible points and the payload fields needed by the caller.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The call uses Qdrant's &lt;code&gt;QueryAsync&lt;/code&gt; API. Passing the float array as &lt;code&gt;query&lt;/code&gt; requests nearest-neighbor search against the collection's unnamed dense vector. The collection already defines cosine distance, so the request does not choose another metric.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;payloadSelector&lt;/code&gt; returns only &lt;code&gt;document_id&lt;/code&gt;, &lt;code&gt;title&lt;/code&gt;, and &lt;code&gt;text&lt;/code&gt;. &lt;code&gt;vectorsSelector: false&lt;/code&gt; makes it explicit that the application does not want the stored 768- or 1,536-value vectors returned. Qdrant query responses already omit vectors by default, so the argument documents the intended response shape. A retrieval result normally needs a stable source identity and displayable content, not another copy of the indexed vector.&lt;/p&gt;

&lt;p&gt;The payload guard in the loop checks that those three fields exist. It does not validate their Qdrant value types. That is enough for this controlled sample, where the indexing code owns the payload. If several writers can produce points, validate the payload schema at that boundary instead of assuming every stored value is a string.&lt;/p&gt;

&lt;p&gt;The result count is an upper bound. &lt;code&gt;limit: 3&lt;/code&gt; does not guarantee three points. Only two of the sample documents pass the allowlist, so at most two can be returned.&lt;/p&gt;

&lt;h2&gt;
  
  
  Filters decide eligibility and vectors decide order
&lt;/h2&gt;

&lt;p&gt;I keep the allowlist in the example instead of hiding it behind a helper method:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt; &lt;span class="n"&gt;eligibleDocumentIds&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"worker-shutdown"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"bounded-work-queue"&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;
&lt;span class="n"&gt;Filter&lt;/span&gt; &lt;span class="n"&gt;eligibilityFilter&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;Match&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"document_id"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;eligibleDocumentIds&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In an application, that list would come from trusted policy and application state rather than raw user input. The same boundary can represent tenant membership, permissions, language, publication state, or document type.&lt;/p&gt;

&lt;p&gt;Only points that satisfy the filter are eligible for the returned results. Qdrant integrates that constraint into its query planning and search instead of retrieving an unfiltered top-k and removing points afterward. With HNSW, Qdrant can use filter-aware graph traversal rather than building a filtered set and starting a separate vector search.&lt;/p&gt;

&lt;p&gt;Filtering after retrieval is a different design. It can leave the caller with too few allowed results because ineligible points have already occupied places in the unfiltered top-k. It also puts a mandatory policy check outside the retrieval request.&lt;/p&gt;

&lt;p&gt;This local collection has three points and no payload index, which is enough to demonstrate the request shape. For a larger collection, create payload indexes for fields used frequently in filters. A keyword index on &lt;code&gt;document_id&lt;/code&gt; supports efficient exact matching and gives Qdrant better information for query planning.&lt;/p&gt;

&lt;p&gt;Create those indexes before bulk ingestion when you already know the filter contract. Payload indexes also inform Qdrant's filter-aware vector indexing, so defining them before the HNSW index is built avoids rebuilding that structure later. Qdrant Cloud also enables strict mode by default for new collections and rejects retrieval filters on unindexed fields.&lt;/p&gt;

&lt;p&gt;The allowlist is an example, not a complete authorization system. The application still owns how it derives the eligible IDs and whether the caller may see the returned payload.&lt;/p&gt;

&lt;h2&gt;
  
  
  A score is not a relevance decision
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;ScoredPoint.Score&lt;/code&gt; lets you compare the returned points for this query. With the cosine collection used here, a higher score ranks ahead of a lower score. It does not prove that the first document answers the question.&lt;/p&gt;

&lt;p&gt;Avoid copying a threshold from another project. For this dense-vector stage, its useful range depends on the embedding model, corpus, query distribution, and distance metric. Calibrate it with labeled queries and expected results, including cases where the correct outcome is no result.&lt;/p&gt;

&lt;p&gt;Thresholds are stage-specific. A value calibrated for dense cosine search does not carry over to hybrid-fusion or reranker scores. Those stages produce different score spaces.&lt;/p&gt;

&lt;p&gt;The three-note corpus is useful for checking the mechanics, not for measuring retrieval quality. Try at least these changes before building context for a model:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Rephrase the shutdown question and confirm that &lt;code&gt;worker-shutdown&lt;/code&gt; remains near the top.&lt;/li&gt;
&lt;li&gt;Ask about HTTP connection reuse, then add &lt;code&gt;http-client-lifetime&lt;/code&gt; to the eligible IDs and inspect its rank.&lt;/li&gt;
&lt;li&gt;Remove the relevant document from the allowlist and confirm that it does not appear.&lt;/li&gt;
&lt;li&gt;Ask an unrelated question and observe that Qdrant still returns the nearest eligible points unless a calibrated acceptance rule rejects them.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last case matters. A vector search asks which eligible stored vectors are closest under the configured metric. Depending on the collection and search path, the returned points may be approximate nearest neighbors. This query has no acceptance criterion, so it cannot determine whether any result is relevant enough to use. Even with a calibrated score threshold, the cutoff remains an application decision rather than proof of relevance.&lt;/p&gt;

&lt;h2&gt;
  
  
  When to use this query path
&lt;/h2&gt;

&lt;p&gt;Use this example when the application already has a small Qdrant collection populated with one compatible embedding profile and needs a direct dense-vector query with mandatory payload filtering. It is also a good boundary to wrap behind an application retrieval interface later.&lt;/p&gt;

&lt;p&gt;Do not treat the code as a complete RAG pipeline. It does not chunk long documents, deduplicate overlapping results, group chunks by source, rerank candidates, calibrate a relevance threshold, or assemble a context budget for a model. It also does not create the payload indexes needed for a larger filtered workload.&lt;/p&gt;

&lt;p&gt;The next implementation step is to replace the whole-document points with deliberate chunks and see how those boundaries change what the query can retrieve. After that, the ranked chunks still need to be selected, cited, and fitted into model context.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/posts/indexing-your-first-documents/" rel="noopener noreferrer"&gt;Indexing Your First Documents&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/posts/your-first-qdrant-collection/" rel="noopener noreferrer"&gt;Your First Qdrant Collection&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/tips/keep-vector-search-filters-separate-from-semantic-ranking/" rel="noopener noreferrer"&gt;Keep vector search filters separate from semantic ranking&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/posts/quick-tip-1/" rel="noopener noreferrer"&gt;Stop RAG Hallucinations with the Short-Circuit Pattern&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Qdrant: &lt;a href="https://qdrant.tech/documentation/search/" rel="noopener noreferrer"&gt;Similarity search and the Query API&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Qdrant: &lt;a href="https://qdrant.tech/documentation/search/filtering/" rel="noopener noreferrer"&gt;Filtering&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Qdrant: &lt;a href="https://qdrant.tech/documentation/manage-data/indexing/" rel="noopener noreferrer"&gt;Payload indexing&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Qdrant: &lt;a href="https://qdrant.tech/course/beginners/module-2/hnsw/" rel="noopener noreferrer"&gt;Fast approximate search with HNSW&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Qdrant: &lt;a href="https://qdrant.tech/documentation/cloud/configure-cluster/" rel="noopener noreferrer"&gt;Configure a Cloud cluster&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Qdrant: &lt;a href="https://github.com/qdrant/qdrant-dotnet" rel="noopener noreferrer"&gt;.NET client&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/ai/iembeddinggenerator" rel="noopener noreferrer"&gt;Use the &lt;code&gt;IEmbeddingGenerator&lt;/code&gt; interface&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Google: &lt;a href="https://ai.google.dev/gemma/docs/embeddinggemma/model_card" rel="noopener noreferrer"&gt;EmbeddingGemma model card and input formats&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>dotnet</category>
      <category>csharp</category>
      <category>qdrant</category>
      <category>rag</category>
    </item>
    <item>
      <title>Building the .NET AI App Foundation</title>
      <dc:creator>Lukas Walter </dc:creator>
      <pubDate>Thu, 24 Sep 2026 15:30:00 +0000</pubDate>
      <link>https://dev.to/lukaswalter/building-the-net-ai-app-foundation-9m9</link>
      <guid>https://dev.to/lukaswalter/building-the-net-ai-app-foundation-9m9</guid>
      <description>&lt;p&gt;If you already keep provider SDKs at the edge and model calls behind a use-case service, the next problem is assembling those decisions into a working ASP.NET Core application without inventing an AI platform.&lt;/p&gt;

&lt;p&gt;This article builds the smallest foundation I would use for that application: one concrete operation, validated model configuration, an instrumented &lt;code&gt;IChatClient&lt;/code&gt; pipeline, and explicit buffered and streaming HTTP contracts.&lt;/p&gt;

&lt;p&gt;Retrieval, tools, evaluation, and runtime controls stay out of this first version. The code still needs a sensible place for each one when it becomes a real requirement.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep the dependency direction small
&lt;/h2&gt;

&lt;p&gt;The request path in this example is deliberately boring:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;HTTP endpoint
    -&amp;gt; SupportReplyService
        -&amp;gt; IChatClient

composition root
    -&amp;gt; configuration
    -&amp;gt; Azure OpenAI client
    -&amp;gt; IChatClient pipeline
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The articles on &lt;a href="https://www.lukaswalter.dev/posts/stop-letting-provider-sdks-define-your-dotnet-ai-architecture/" rel="noopener noreferrer"&gt;provider abstraction&lt;/a&gt; and &lt;a href="https://www.lukaswalter.dev/posts/designing-an-ai-service-layer/" rel="noopener noreferrer"&gt;AI service layers&lt;/a&gt; explain these two boundaries in detail. I will treat them as settled here and focus on the host around them.&lt;/p&gt;

&lt;p&gt;I care more about those arrows than the number of projects. Everything can live in one ASP.NET Core project at first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AiAppFoundation/
├── Program.cs
├── AI/
│   └── SupportModelOptions.cs
└── SupportReplies/
    ├── SupportReplyContracts.cs
    └── SupportReplyService.cs
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Five class libraries would add more project references and more files to navigate without improving this small dependency graph. Keep the direction clear in code first. Add an assembly when you need to enforce dependencies at compile time, reuse a component independently, or reflect a clear ownership boundary. An independently deployed application will usually have its own assembly, but a class library is not a deployment boundary by itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with one actual operation
&lt;/h2&gt;

&lt;p&gt;The sample drafts a reply to a customer message. It is intentionally unglamorous. Put the application contract in &lt;code&gt;SupportReplies/SupportReplyContracts.cs&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;namespace&lt;/span&gt; &lt;span class="nn"&gt;AiAppFoundation.SupportReplies&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;sealed&lt;/span&gt; &lt;span class="k"&gt;record&lt;/span&gt; &lt;span class="nc"&gt;DraftReplyRequest&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;CustomerMessage&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;sealed&lt;/span&gt; &lt;span class="k"&gt;record&lt;/span&gt; &lt;span class="nc"&gt;DraftReply&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;Text&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The endpoint does not accept a system prompt, arbitrary tools, a model name, or provider options. A caller asks for a support reply. It does not get to reconfigure the feature.&lt;/p&gt;

&lt;p&gt;Put the service in &lt;code&gt;SupportReplies/SupportReplyService.cs&lt;/code&gt;. It owns the instructions and builds fresh request objects for each operation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;System.Runtime.CompilerServices&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Microsoft.Extensions.AI&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;namespace&lt;/span&gt; &lt;span class="nn"&gt;AiAppFoundation.SupportReplies&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;sealed&lt;/span&gt; &lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;SupportReplyService&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;IChatClient&lt;/span&gt; &lt;span class="n"&gt;chatClient&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;private&lt;/span&gt; &lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;Instructions&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"""
&lt;/span&gt;        &lt;span class="n"&gt;Draft&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="n"&gt;concise&lt;/span&gt; &lt;span class="n"&gt;support&lt;/span&gt; &lt;span class="n"&gt;reply&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;
        &lt;span class="n"&gt;Use&lt;/span&gt; &lt;span class="n"&gt;only&lt;/span&gt; &lt;span class="n"&gt;facts&lt;/span&gt; &lt;span class="n"&gt;present&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;customer&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;
        &lt;span class="n"&gt;If&lt;/span&gt; &lt;span class="n"&gt;information&lt;/span&gt; &lt;span class="k"&gt;is&lt;/span&gt; &lt;span class="n"&gt;missing&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ask&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;it&lt;/span&gt; &lt;span class="n"&gt;instead&lt;/span&gt; &lt;span class="n"&gt;of&lt;/span&gt; &lt;span class="n"&gt;inventing&lt;/span&gt; &lt;span class="n"&gt;it&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;
        &lt;span class="s"&gt;""";
&lt;/span&gt;
    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="n"&gt;Task&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;DraftReply&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;DraftAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;DraftReplyRequest&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;CancellationToken&lt;/span&gt; &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;ChatResponse&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;chatClient&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GetResponseAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="nf"&gt;CreateMessages&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
            &lt;span class="nf"&gt;CreateOptions&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
            &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;DraftReply&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Text&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="n"&gt;IAsyncEnumerable&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;StreamAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;DraftReplyRequest&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;EnumeratorCancellation&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="n"&gt;CancellationToken&lt;/span&gt; &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="k"&gt;foreach&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ChatResponseUpdate&lt;/span&gt; &lt;span class="n"&gt;update&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt;
            &lt;span class="n"&gt;chatClient&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GetStreamingResponseAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="nf"&gt;CreateMessages&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
                &lt;span class="nf"&gt;CreateOptions&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
                &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(!&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;IsNullOrEmpty&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;update&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Text&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;update&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Text&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
            &lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;private&lt;/span&gt; &lt;span class="k"&gt;static&lt;/span&gt; &lt;span class="n"&gt;List&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;ChatMessage&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;CreateMessages&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;DraftReplyRequest&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="k"&gt;new&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ChatRole&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;System&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Instructions&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="k"&gt;new&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ChatRole&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;User&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;CustomerMessage&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;];&lt;/span&gt;

    &lt;span class="k"&gt;private&lt;/span&gt; &lt;span class="k"&gt;static&lt;/span&gt; &lt;span class="n"&gt;ChatOptions&lt;/span&gt; &lt;span class="nf"&gt;CreateOptions&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;MaxOutputTokens&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="m"&gt;500&lt;/span&gt;
    &lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;SupportReplyService&lt;/code&gt; is registered as a concrete type on purpose. The operation, request, and result already give callers a useful boundary. Add a feature-specific interface when you need to replace the service itself, not simply to hide &lt;code&gt;IChatClient&lt;/code&gt;. &lt;a href="https://www.lukaswalter.dev/posts/designing-an-ai-service-layer/" rel="noopener noreferrer"&gt;Designing an AI Service Layer&lt;/a&gt; explains why a generic &lt;code&gt;IAiService&lt;/code&gt; adds little.&lt;/p&gt;

&lt;p&gt;Create the messages and &lt;code&gt;ChatOptions&lt;/code&gt; for every call because client implementations may mutate them. &lt;a href="https://www.lukaswalter.dev/posts/dependency-injection-for-ai-components/" rel="noopener noreferrer"&gt;Dependency Injection for AI Components&lt;/a&gt; covers the broader lifetime and scoped-dependency rules.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wire configuration, model access, and tracing
&lt;/h2&gt;

&lt;p&gt;The code here targets .NET 10. I tested it with &lt;code&gt;Microsoft.Extensions.AI&lt;/code&gt; 10.10.0, &lt;code&gt;Microsoft.Extensions.AI.OpenAI&lt;/code&gt; 10.10.0, &lt;code&gt;Azure.AI.OpenAI&lt;/code&gt; 2.1.0, &lt;code&gt;Azure.Identity&lt;/code&gt; 1.21.0, and OpenTelemetry 1.18.0. These are tested pins, not a claim that they will still be current when you start a new project.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dotnet new web &lt;span class="nt"&gt;-n&lt;/span&gt; AiAppFoundation &lt;span class="nt"&gt;-f&lt;/span&gt; net10.0
&lt;span class="nb"&gt;cd &lt;/span&gt;AiAppFoundation
dotnet add package Microsoft.Extensions.AI &lt;span class="nt"&gt;--version&lt;/span&gt; 10.10.0
dotnet add package Microsoft.Extensions.AI.OpenAI &lt;span class="nt"&gt;--version&lt;/span&gt; 10.10.0
dotnet add package Azure.AI.OpenAI &lt;span class="nt"&gt;--version&lt;/span&gt; 2.1.0
dotnet add package Azure.Identity &lt;span class="nt"&gt;--version&lt;/span&gt; 1.21.0
dotnet add package OpenTelemetry.Extensions.Hosting &lt;span class="nt"&gt;--version&lt;/span&gt; 1.18.0
dotnet add package OpenTelemetry.Instrumentation.AspNetCore &lt;span class="nt"&gt;--version&lt;/span&gt; 1.18.0
dotnet add package OpenTelemetry.Exporter.OpenTelemetryProtocol &lt;span class="nt"&gt;--version&lt;/span&gt; 1.18.0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Put the provider-facing settings in &lt;code&gt;AI/SupportModelOptions.cs&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;System.ComponentModel.DataAnnotations&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;namespace&lt;/span&gt; &lt;span class="nn"&gt;AiAppFoundation.AI&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;sealed&lt;/span&gt; &lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;SupportModelOptions&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;SectionName&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"AI:SupportReplies"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;Required&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="n"&gt;Uri&lt;/span&gt; &lt;span class="n"&gt;Endpoint&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;get&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;init&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt;&lt;span class="p"&gt;!;&lt;/span&gt;

    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;Required&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;Deployment&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;get&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;init&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;""&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Add the corresponding routing information to &lt;code&gt;appsettings.json&lt;/code&gt;. It does not contain a credential:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"AI"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"SupportReplies"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"Endpoint"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://example.openai.azure.com/"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"Deployment"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"support-replies"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Register the options and client in &lt;code&gt;Program.cs&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;AiAppFoundation.AI&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;AiAppFoundation.SupportReplies&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Azure.AI.OpenAI&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Azure.Identity&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Microsoft.Extensions.AI&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Microsoft.Extensions.Options&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;OpenTelemetry.Trace&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="n"&gt;WebApplicationBuilder&lt;/span&gt; &lt;span class="n"&gt;builder&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;WebApplication&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;CreateBuilder&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;AiActivitySourceName&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"AiAppFoundation.AI"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="n"&gt;builder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Services&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;AddOptions&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;SupportModelOptions&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;()&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;BindConfiguration&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;SupportModelOptions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;SectionName&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ValidateDataAnnotations&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ValidateOnStart&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="n"&gt;builder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Services&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;AddOpenTelemetry&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;WithTracing&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tracing&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;tracing&lt;/span&gt;
        &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;AddSource&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;AiActivitySourceName&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;AddAspNetCoreInstrumentation&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;AddOtlpExporter&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt;

&lt;span class="n"&gt;builder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Services&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;AddChatClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;services&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;SupportModelOptions&lt;/span&gt; &lt;span class="n"&gt;options&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;services&lt;/span&gt;
            &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;GetRequiredService&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;IOptions&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;SupportModelOptions&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&amp;gt;()&lt;/span&gt;
            &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Value&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

        &lt;span class="kt"&gt;var&lt;/span&gt; &lt;span class="n"&gt;azureClient&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;AzureOpenAIClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Endpoint&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;DefaultAzureCredential&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt;

        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;azureClient&lt;/span&gt;
            &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GetChatClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Deployment&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;AsIChatClient&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;UseOpenTelemetry&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;sourceName&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;AiActivitySourceName&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;configure&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;telemetry&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt;
            &lt;span class="n"&gt;telemetry&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;EnableSensitiveData&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;false&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="n"&gt;builder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Services&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;AddScoped&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;SupportReplyService&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;();&lt;/span&gt;

&lt;span class="n"&gt;WebApplication&lt;/span&gt; &lt;span class="n"&gt;app&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;builder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Build&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;ValidateOnStart&lt;/code&gt; catches missing values before the first user request. It cannot prove that Azure is reachable or that the deployment supports every capability the application might request later. Cover those questions with deployment checks and integration tests.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;DefaultAzureCredential&lt;/code&gt; keeps secrets out of this configuration. For local development, authenticate with a supported developer credential before starting the application. With Azure CLI:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;az login
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If your account belongs to more than one Microsoft Entra tenant, sign in to the tenant that contains the Azure OpenAI resource:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;az login &lt;span class="nt"&gt;--tenant&lt;/span&gt; &amp;lt;tenant-id&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The signed-in identity also needs access to the resource. For inference, &lt;code&gt;Cognitive Services OpenAI User&lt;/code&gt; is one suitable role. Check the operations your application performs and grant only the permissions they require. Microsoft's &lt;a href="https://learn.microsoft.com/en-us/dotnet/ai/azure-ai-services-authentication" rel="noopener noreferrer"&gt;keyless authentication guidance&lt;/a&gt; covers other local and deployed credential options.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;AddChatClient&lt;/code&gt; registers the pipeline as a singleton by default. Keep request state, scoped database contexts, and caller-specific tools out of its factory. Bind those dependencies inside the operation that owns them.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;UseOpenTelemetry&lt;/code&gt; is easy to overread. It instruments the chat-client pipeline, but it does not collect or export those activities by itself. A host without a tracing provider subscribed to that source still receives nothing. In this sample, &lt;code&gt;AddSource&lt;/code&gt; subscribes to the MEAI source, ASP.NET Core instrumentation supplies the surrounding request span, and the OTLP exporter forwards both. Without exporter settings, it uses the local default for the selected OTLP protocol. Point it to a remote collector through the normal OpenTelemetry configuration for your environment, such as &lt;code&gt;OTEL_EXPORTER_OTLP_ENDPOINT&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The explicit &lt;code&gt;EnableSensitiveData = false&lt;/code&gt; keeps raw prompts, responses, function arguments, and function results out of MEAI telemetry even if &lt;code&gt;OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT&lt;/code&gt; is set elsewhere in the environment. Metadata such as token counts can still be recorded.&lt;/p&gt;

&lt;p&gt;The generative AI semantic conventions are still experimental, so individual telemetry attributes should not become application contracts. This sample traces the HTTP request and model call. Retrieval, tools, retries, and persistence will need their own spans when they arrive.&lt;/p&gt;

&lt;h2&gt;
  
  
  Choose the HTTP contract: buffered or streaming
&lt;/h2&gt;

&lt;p&gt;Continue in &lt;code&gt;Program.cs&lt;/code&gt; with the route mappings. My default for this operation is a buffered JSON response:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="n"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;MapPost&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"/support/replies"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;DraftReplyRequest&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;SupportReplyService&lt;/span&gt; &lt;span class="n"&gt;replies&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;CancellationToken&lt;/span&gt; &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;IsNullOrWhiteSpace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;CustomerMessage&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;Results&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;BadRequest&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;new&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="n"&gt;error&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"customerMessage is required."&lt;/span&gt;
        &lt;span class="p"&gt;});&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="n"&gt;DraftReply&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;replies&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;DraftAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;Results&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Ok&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The complete draft exists before the response starts. The application can validate it, then choose the final status code and JSON response as one unit.&lt;/p&gt;

&lt;p&gt;Streaming can make a longer generation feel faster. It also changes the API contract. I would give it a separate endpoint rather than quietly changing the buffered one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="n"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;MapPost&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"/support/replies/stream"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;DraftReplyRequest&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;SupportReplyService&lt;/span&gt; &lt;span class="n"&gt;replies&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;HttpResponse&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;CancellationToken&lt;/span&gt; &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;IsNullOrWhiteSpace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;CustomerMessage&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;StatusCode&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;StatusCodes&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Status400BadRequest&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;WriteAsJsonAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="n"&gt;error&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"customerMessage is required."&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
            &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ContentType&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"text/plain; charset=utf-8"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="k"&gt;foreach&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="n"&gt;replies&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;StreamAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;WriteAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;FlushAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="n"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Run&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Flushing after every update favors immediate delivery in this small example. Model updates can contain only a few characters, so a production endpoint may buffer tiny deltas and flush them in larger chunks to reduce write and frame overhead. Measure the effect against the latency the UI needs.&lt;/p&gt;

&lt;p&gt;Once the response starts, the endpoint cannot replace it with a neat JSON error. The client also needs framing when a stream carries citations, tool progress, usage, or a final outcome. Plain text works here only because text is the entire contract.&lt;/p&gt;

&lt;p&gt;Server-Sent Events or another framed protocol will often suit a browser application better. The important decision is earlier: stream because the product benefits from partial output, not because the model API happens to expose it.&lt;/p&gt;

&lt;p&gt;Both endpoints pass ASP.NET Core's request cancellation token through &lt;code&gt;SupportReplyService&lt;/code&gt; to the model call. Keep that chain intact. Cancellation is cooperative: it cannot guarantee that provider work or billing stops immediately, and it cannot undo a side effect that has already happened.&lt;/p&gt;

&lt;p&gt;Run the application on a fixed local URL:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dotnet run &lt;span class="nt"&gt;--urls&lt;/span&gt; http://localhost:5000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This starts the HTTP server. It does not print a model response in the terminal. Once the application is listening, send a buffered request from another terminal:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST http://localhost:5000/support/replies &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"customerMessage":"My order has not arrived yet. What should I do?"}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The response is JSON with a &lt;code&gt;text&lt;/code&gt; property. Its wording depends on the configured model. For the streaming route, tell &lt;code&gt;curl&lt;/code&gt; not to buffer the response:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;--no-buffer&lt;/span&gt; &lt;span class="nt"&gt;-X&lt;/span&gt; POST http://localhost:5000/support/replies/stream &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"customerMessage":"My order has not arrived yet."}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Leave room without building placeholders
&lt;/h2&gt;

&lt;p&gt;I would not create interfaces for retrieval, evaluation, and tools on day one. I would know where each concern belongs before adding it.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Capability&lt;/th&gt;
&lt;th&gt;Where it attaches&lt;/th&gt;
&lt;th&gt;Boundary to preserve&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Retrieval&lt;/td&gt;
&lt;td&gt;A capability injected into &lt;code&gt;SupportReplyService&lt;/code&gt; before message construction&lt;/td&gt;
&lt;td&gt;Application code chooses allowed data and filters before anything reaches the model.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tools&lt;/td&gt;
&lt;td&gt;Per-operation &lt;code&gt;ChatOptions&lt;/code&gt;, with function-invocation middleware in the client pipeline&lt;/td&gt;
&lt;td&gt;A singleton client must not capture scoped authorization or data services.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Structured output&lt;/td&gt;
&lt;td&gt;The service's response contract and model call&lt;/td&gt;
&lt;td&gt;Validate the deserialized result before returning or persisting it.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Evaluation&lt;/td&gt;
&lt;td&gt;Tests and production feedback around the feature contract&lt;/td&gt;
&lt;td&gt;Keep deterministic application tests separate from model-quality evaluation.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Observability&lt;/td&gt;
&lt;td&gt;The request trace and the &lt;code&gt;IChatClient&lt;/code&gt; pipeline&lt;/td&gt;
&lt;td&gt;Correlate the complete operation, including work outside the provider call.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Runtime controls&lt;/td&gt;
&lt;td&gt;The endpoint, service, and client middleware&lt;/td&gt;
&lt;td&gt;Carry one request budget through retries, tools, and model calls.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Treat this as an attachment map, not a backlog. Add a capability when the feature requires it. Retrieval should arrive as an authorized application dependency. Bind tools for the current operation. Start evaluation with a concrete quality question rather than an evaluator package.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this foundation deliberately does not solve
&lt;/h2&gt;

&lt;p&gt;This sample is an application shape, not a production-ready AI module. It does not yet define:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;authentication or authorization for the endpoint&lt;/li&gt;
&lt;li&gt;maximum input size and request rate&lt;/li&gt;
&lt;li&gt;structured output and response acceptance rules&lt;/li&gt;
&lt;li&gt;provider timeouts, recovery, or overload behavior&lt;/li&gt;
&lt;li&gt;content-safety policy&lt;/li&gt;
&lt;li&gt;tool authorization and side-effect handling&lt;/li&gt;
&lt;li&gt;retrieval filters and data-access rules&lt;/li&gt;
&lt;li&gt;evaluation datasets and release gates&lt;/li&gt;
&lt;li&gt;telemetry redaction and retention&lt;/li&gt;
&lt;li&gt;deployment health checks or fallback behavior&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those omissions are deliberate. Their design depends on the product, its data, and the consequences of failure. This foundation does not answer those questions. It gives their eventual answers a clear place in the application.&lt;/p&gt;

&lt;h2&gt;
  
  
  When this shape earns its keep
&lt;/h2&gt;

&lt;p&gt;Use this foundation when a model experiment is becoming an application feature with an HTTP contract, shared configuration, operational requirements, or a likely path toward data and tools. A disposable console experiment can remain smaller and call &lt;code&gt;IChatClient&lt;/code&gt; directly.&lt;/p&gt;

&lt;p&gt;Before adding retrieval or tools, I would check these points:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Provider SDK types appear only at the composition edge.&lt;/li&gt;
&lt;li&gt;Missing configuration stops the application at startup.&lt;/li&gt;
&lt;li&gt;The endpoint accepts an application request, not an arbitrary prompt or model name.&lt;/li&gt;
&lt;li&gt;One use-case service owns the instructions, messages, and result contract.&lt;/li&gt;
&lt;li&gt;Messages and call options are local to one operation.&lt;/li&gt;
&lt;li&gt;Buffered and streaming endpoints have explicit response contracts.&lt;/li&gt;
&lt;li&gt;The caller's cancellation token reaches the model call.&lt;/li&gt;
&lt;li&gt;The tracing provider subscribes to the MEAI activity source and has an exporter.&lt;/li&gt;
&lt;li&gt;MEAI sensitive-data capture is disabled explicitly.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;At that point, add the next real requirement where it belongs. The application does not need an AI platform first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.lukaswalter.dev/posts/stop-letting-provider-sdks-define-your-dotnet-ai-architecture/" rel="noopener noreferrer"&gt;Stop Letting Provider SDKs Define Your .NET AI Architecture&lt;/a&gt; explains why provider clients stay at the composition edge.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.lukaswalter.dev/posts/designing-an-ai-service-layer/" rel="noopener noreferrer"&gt;Designing an AI Service Layer&lt;/a&gt; explains why application operations should not become a generic &lt;code&gt;IAiService&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.lukaswalter.dev/posts/dependency-injection-for-ai-components/" rel="noopener noreferrer"&gt;Dependency Injection for AI Components&lt;/a&gt; covers lifetimes, captured dependencies, and per-operation state.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/ai/microsoft-extensions-ai" rel="noopener noreferrer"&gt;Microsoft.Extensions.AI libraries&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/ai/ichatclient" rel="noopener noreferrer"&gt;Use the &lt;code&gt;IChatClient&lt;/code&gt; interface&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.ai.ichatclient?view=net-10.0-pp" rel="noopener noreferrer"&gt;&lt;code&gt;IChatClient&lt;/code&gt; API and concurrency contract&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.dependencyinjection.chatclientbuilderservicecollectionextensions.addchatclient?view=net-10.0-pp" rel="noopener noreferrer"&gt;&lt;code&gt;AddChatClient&lt;/code&gt; registration and default lifetime&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.ai.chatclientextensions.getstreamingresponseasync?view=net-10.0-pp" rel="noopener noreferrer"&gt;&lt;code&gt;GetStreamingResponseAsync&lt;/code&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/aspnet/core/fundamentals/configuration/options?view=aspnetcore-10.0" rel="noopener noreferrer"&gt;Options validation&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/aspnet/core/fundamentals/use-http-context?view=aspnetcore-10.0" rel="noopener noreferrer"&gt;Use &lt;code&gt;HttpContext&lt;/code&gt; and &lt;code&gt;RequestAborted&lt;/code&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/cli/azure/authenticate-azure-cli-interactively" rel="noopener noreferrer"&gt;Sign in interactively with Azure CLI&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/ai/azure-ai-services-authentication" rel="noopener noreferrer"&gt;Azure OpenAI with keyless authentication&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.ai.opentelemetrychatclient?view=net-10.0-pp" rel="noopener noreferrer"&gt;&lt;code&gt;OpenTelemetryChatClient&lt;/code&gt; and the experimental GenAI conventions&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.ai.opentelemetrychatclient.enablesensitivedata?view=net-10.0-pp" rel="noopener noreferrer"&gt;&lt;code&gt;OpenTelemetryChatClient.EnableSensitiveData&lt;/code&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;OpenTelemetry: &lt;a href="https://opentelemetry.io/docs/languages/dotnet/traces/getting-started-aspnetcore/" rel="noopener noreferrer"&gt;Getting started with traces in ASP.NET Core&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;NuGet: &lt;a href="https://www.nuget.org/packages/Microsoft.Extensions.AI/10.10.0" rel="noopener noreferrer"&gt;&lt;code&gt;Microsoft.Extensions.AI&lt;/code&gt; 10.10.0&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;NuGet: &lt;a href="https://www.nuget.org/packages/Microsoft.Extensions.AI.OpenAI/10.10.0" rel="noopener noreferrer"&gt;&lt;code&gt;Microsoft.Extensions.AI.OpenAI&lt;/code&gt; 10.10.0&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;NuGet: &lt;a href="https://www.nuget.org/packages/OpenTelemetry.Extensions.Hosting/1.18.0" rel="noopener noreferrer"&gt;&lt;code&gt;OpenTelemetry.Extensions.Hosting&lt;/code&gt; 1.18.0&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;NuGet: &lt;a href="https://www.nuget.org/packages/OpenTelemetry.Exporter.OpenTelemetryProtocol/1.18.0" rel="noopener noreferrer"&gt;&lt;code&gt;OpenTelemetry.Exporter.OpenTelemetryProtocol&lt;/code&gt; 1.18.0&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>dotnet</category>
      <category>csharp</category>
      <category>ai</category>
    </item>
    <item>
      <title>Indexing Your First Documents</title>
      <dc:creator>Lukas Walter </dc:creator>
      <pubDate>Tue, 22 Sep 2026 15:30:00 +0000</pubDate>
      <link>https://dev.to/lukaswalter/indexing-your-first-documents-j4h</link>
      <guid>https://dev.to/lukaswalter/indexing-your-first-documents-j4h</guid>
      <description>&lt;p&gt;This example writes a short support document into Qdrant, reads it back, and lets you run the same code again without creating another point for that document.&lt;/p&gt;

&lt;p&gt;In &lt;a href="https://www.lukaswalter.dev/posts/your-first-qdrant-collection/" rel="noopener noreferrer"&gt;Your First Qdrant Collection&lt;/a&gt;, we created &lt;code&gt;support-documents-v1&lt;/code&gt; with an unnamed dense vector and checked its embedding profile. That collection is ready for the Azure OpenAI path below. If you prefer local embeddings, the optional Ollama setup creates a separate collection on the same Qdrant server. Both paths use the same indexing code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Continue from the collection example
&lt;/h2&gt;

&lt;p&gt;Use the &lt;code&gt;QdrantCollections&lt;/code&gt; console project from the collection article.&lt;/p&gt;

&lt;p&gt;If you stopped the existing container, start it again:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker start &lt;span class="nt"&gt;-a&lt;/span&gt; qdrant-local
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep it running and use another terminal for the .NET commands. If you are starting here, follow the &lt;a href="https://www.lukaswalter.dev/posts/your-first-qdrant-collection/#run-qdrant-locally" rel="noopener noreferrer"&gt;local setup and collection example&lt;/a&gt; first. It includes the Docker command, project creation, and collection validation.&lt;/p&gt;

&lt;p&gt;Choose one embedding setup. Azure OpenAI continues with the existing collection; Ollama needs the preparation in the next section.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Setting&lt;/th&gt;
&lt;th&gt;Azure OpenAI&lt;/th&gt;
&lt;th&gt;Ollama locally&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Model&lt;/td&gt;
&lt;td&gt;&lt;code&gt;text-embedding-3-small&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;embeddinggemma:300m&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dimensions&lt;/td&gt;
&lt;td&gt;1,536&lt;/td&gt;
&lt;td&gt;768&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Distance&lt;/td&gt;
&lt;td&gt;Cosine&lt;/td&gt;
&lt;td&gt;Cosine&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Collection&lt;/td&gt;
&lt;td&gt;&lt;code&gt;support-documents-v1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;support-documents-embeddinggemma-v1&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Profile ID&lt;/td&gt;
&lt;td&gt;&lt;code&gt;support-text-v1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;support-embeddinggemma-v1&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The local path uses a separate collection because EmbeddingGemma produces a different vector space. The later query must target the collection for its embedding profile.&lt;/p&gt;

&lt;h2&gt;
  
  
  Use Azure OpenAI in Microsoft Foundry
&lt;/h2&gt;

&lt;p&gt;The package commands pin exact versions so you can reproduce the setup. The Microsoft.Extensions.AI dependencies stay at 10.9.0 to match the earlier embedding example. Newer compatible versions may also work.&lt;/p&gt;

&lt;p&gt;In the project directory, add the Microsoft.Extensions.AI OpenAI integration used in &lt;a href="https://www.lukaswalter.dev/posts/embeddings-in-dotnet/" rel="noopener noreferrer"&gt;Embeddings in .NET&lt;/a&gt;, plus the Azure OpenAI client:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dotnet add package Microsoft.Extensions.AI.OpenAI &lt;span class="nt"&gt;--version&lt;/span&gt; 10.9.0
dotnet add package Azure.AI.OpenAI &lt;span class="nt"&gt;--version&lt;/span&gt; 2.1.0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Deploy &lt;code&gt;text-embedding-3-small&lt;/code&gt; in Microsoft Foundry. This example connects through the Azure OpenAI resource endpoint shown below. Keep its default 1,536-dimensional output. Set these environment variables in the terminal that runs the application:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Variable&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;AZURE_OPENAI_ENDPOINT&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Your Azure OpenAI resource endpoint, such as &lt;code&gt;https://your-resource.openai.azure.com/&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;AZURE_OPENAI_EMBEDDING_DEPLOYMENT&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The name you assigned to the &lt;code&gt;text-embedding-3-small&lt;/code&gt; deployment, such as &lt;code&gt;support-embeddings&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;AZURE_OPENAI_API_KEY&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;A key for that Azure OpenAI resource&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Use the resource endpoint for this &lt;code&gt;AzureOpenAIClient&lt;/code&gt; example, not a Foundry project URL or an endpoint with &lt;code&gt;/openai/v1&lt;/code&gt; appended. The client builds the deployment-specific request path. Keep the key outside source control. This example uses key authentication. The Azure SDK also supports &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/overview/azure/ai.openai-readme?view=azure-dotnet#authenticate-the-client" rel="noopener noreferrer"&gt;Microsoft Entra credentials&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The embedding request sends the text to your Azure OpenAI deployment and incurs Azure usage charges. Start with the sample notes below.&lt;/p&gt;

&lt;p&gt;Replace &lt;code&gt;Program.cs&lt;/code&gt; with this setup block, then append the shared code under &lt;strong&gt;Index and read back the document&lt;/strong&gt;. Together, the two blocks form the complete program.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Microsoft.Extensions.AI&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Azure.AI.OpenAI&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;System.ClientModel&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Qdrant.Client&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Qdrant.Client.Grpc&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;collectionName&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"support-documents-v1"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;embeddingProfileId&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"support-text-v1"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;vectorSize&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="m"&gt;1536&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;endpoint&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Environment&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GetEnvironmentVariable&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"AZURE_OPENAI_ENDPOINT"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;??&lt;/span&gt; &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;InvalidOperationException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Set AZURE_OPENAI_ENDPOINT."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;deploymentName&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Environment&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GetEnvironmentVariable&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="s"&gt;"AZURE_OPENAI_EMBEDDING_DEPLOYMENT"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;??&lt;/span&gt; &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;InvalidOperationException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Set AZURE_OPENAI_EMBEDDING_DEPLOYMENT."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;apiKey&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Environment&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GetEnvironmentVariable&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"AZURE_OPENAI_API_KEY"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;??&lt;/span&gt; &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;InvalidOperationException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Set AZURE_OPENAI_API_KEY."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;IEmbeddingGenerator&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Embedding&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;float&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;generator&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt;
    &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;AzureOpenAIClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;Uri&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;endpoint&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;ApiKeyCredential&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;apiKey&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
        &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GetEmbeddingClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;deploymentName&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;AsIEmbeddingGenerator&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="n"&gt;Func&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;SourceDocument&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;prepareDocument&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;document&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Text&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://learn.microsoft.com/en-us/dotnet/api/azure.ai.openai.azureopenaiclient.getembeddingclient?view=azure-dotnet" rel="noopener noreferrer"&gt;GetEmbeddingClient&lt;/a&gt; takes the Azure deployment name, which may differ from the model name. Confirm that the deployment uses &lt;code&gt;text-embedding-3-small&lt;/code&gt;. Pointing the variable at another model does not make its vectors compatible.&lt;/p&gt;

&lt;p&gt;The Azure OpenAI path embeds &lt;code&gt;document.Text&lt;/code&gt; exactly as supplied. That keeps the existing &lt;code&gt;support-text-v1&lt;/code&gt; profile: plain text with the default 1,536-dimensional output. We store the title in the payload for display. The empty collection from the first article can be reused with this configuration. If you have already populated it through another endpoint, verify model-version compatibility before mixing vectors. Matching dimensions alone is insufficient.&lt;/p&gt;

&lt;h2&gt;
  
  
  Optional: generate embeddings locally with Ollama
&lt;/h2&gt;

&lt;p&gt;Skip this section if you chose Azure OpenAI.&lt;/p&gt;

&lt;p&gt;Install &lt;a href="https://ollama.com/download" rel="noopener noreferrer"&gt;Ollama&lt;/a&gt; and start it. On a command-line installation, run &lt;code&gt;ollama serve&lt;/code&gt; in a separate terminal. If the desktop app already runs the server, leave that instance running.&lt;/p&gt;

&lt;p&gt;Download the embedding model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ollama pull embeddinggemma:300m
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://ollama.com/library/embeddinggemma" rel="noopener noreferrer"&gt;EmbeddingGemma on Ollama&lt;/a&gt; requires Ollama 0.11.10 or later. After downloading the model, this path generates embeddings on your machine without an Azure subscription or cloud API key. Use the explicit &lt;code&gt;300m&lt;/code&gt; tag for this sample. Tags can change, so record the model digest shown by &lt;code&gt;ollama list&lt;/code&gt; when reproducing results.&lt;/p&gt;

&lt;p&gt;Before replacing the collection program, change its three constants to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;collectionName&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"support-documents-embeddinggemma-v1"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;embeddingProfileId&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"support-embeddinggemma-v1"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;ulong&lt;/span&gt; &lt;span class="n"&gt;vectorSize&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="m"&gt;768&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Leave cosine distance and the rest of that program unchanged. Run &lt;code&gt;dotnet run&lt;/code&gt; to create and validate the local model's collection. The indexing program below does not create collections. It expects this setup step to have completed successfully. The different name preserves &lt;code&gt;support-documents-v1&lt;/code&gt; if it already contains Azure OpenAI vectors.&lt;/p&gt;

&lt;p&gt;Here, &lt;code&gt;support-embeddinggemma-v1&lt;/code&gt; declares &lt;code&gt;embeddinggemma:300m&lt;/code&gt;, the full 768-dimensional output, cosine distance, and the document/query formatting described below. For the local Ollama path, the sample checks the declared profile ID, but it does not enforce model-artifact identity. If the tag later points to a different digest, the collection check still passes. Before deploying this pipeline, record the expected model digest in the profile and compare it with the locally installed model before writing or querying. Stop on a mismatch and review whether the new artifact requires reindexing.&lt;/p&gt;

&lt;p&gt;Add OllamaSharp and pin the abstraction package used by this sample:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dotnet add package Microsoft.Extensions.AI &lt;span class="nt"&gt;--version&lt;/span&gt; 10.9.0
dotnet add package OllamaSharp &lt;span class="nt"&gt;--version&lt;/span&gt; 5.4.30
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;OllamaSharp already depends on &lt;code&gt;Microsoft.Extensions.AI&lt;/code&gt;. The explicit reference pins it to 10.9.0.&lt;/p&gt;

&lt;p&gt;Some .NET 10 SDK versions may report CS9057 for OllamaSharp's optional source generator. This example does not use its generated &lt;code&gt;[OllamaTool]&lt;/code&gt; support.&lt;/p&gt;

&lt;p&gt;Use this setup block instead of the Azure OpenAI block at the top of &lt;code&gt;Program.cs&lt;/code&gt;. Append the same shared indexing code that follows.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Microsoft.Extensions.AI&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;OllamaSharp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Qdrant.Client&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Qdrant.Client.Grpc&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;collectionName&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"support-documents-embeddinggemma-v1"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;embeddingProfileId&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"support-embeddinggemma-v1"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;vectorSize&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="m"&gt;768&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;IEmbeddingGenerator&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Embedding&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;float&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;generator&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt;
    &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;OllamaApiClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;Uri&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"http://localhost:11434"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="s"&gt;"embeddinggemma:300m"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="n"&gt;Func&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;SourceDocument&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;prepareDocument&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;document&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt;
    &lt;span class="s"&gt;$"title: &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Title&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt; | text: &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Text&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://github.com/awaescher/OllamaSharp#usage-with-microsoftextensionsai" rel="noopener noreferrer"&gt;OllamaApiClient&lt;/a&gt; implements &lt;code&gt;IEmbeddingGenerator&amp;lt;string, Embedding&amp;lt;float&amp;gt;&amp;gt;&lt;/code&gt; directly. The indexing code uses that interface in either setup.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;title: ... | text: ...&lt;/code&gt; prefix follows &lt;a href="https://ai.google.dev/gemma/docs/embeddinggemma/model_card" rel="noopener noreferrer"&gt;Google's document format&lt;/a&gt;. Later, a search query should use &lt;code&gt;task: search result | query: ...&lt;/code&gt; with the same model. These prefixes belong to the embedding profile. They are not part of the original document text stored in the payload.&lt;/p&gt;

&lt;p&gt;EmbeddingGemma has a 2,048-token input limit, including the title and document prefix. &lt;a href="https://docs.ollama.com/api/embed" rel="noopener noreferrer"&gt;Ollama's embedding API&lt;/a&gt; defaults to &lt;code&gt;truncate: true&lt;/code&gt;. This sample leaves that default in place, so an oversized input can lose its ending without the request failing. The short notes below fit, but before accepting larger documents, add token-aware size checks or chunking. You can also set the API's &lt;code&gt;truncate&lt;/code&gt; option to &lt;code&gt;false&lt;/code&gt; to reject oversized inputs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Give the source a stable identity
&lt;/h2&gt;

&lt;p&gt;The first document is a short note about stopping a background service. Its source ID is &lt;code&gt;worker-shutdown&lt;/code&gt;, and it has one chunk named &lt;code&gt;whole-document&lt;/code&gt;. Keeping the entire note in one point makes the first write easy to inspect.&lt;/p&gt;

&lt;p&gt;The Qdrant point ID is a fixed UUID stored with that source record. Qdrant accepts UUIDs or unsigned 64-bit integers for point IDs, so a source slug such as &lt;code&gt;worker-shutdown&lt;/code&gt; belongs in the payload rather than directly in the point ID field.&lt;/p&gt;

&lt;p&gt;For this one-chunk example, each document has one persistent point ID. If a document later produces several chunks, each chunk needs its own stable point ID. The &lt;code&gt;document_id&lt;/code&gt; and &lt;code&gt;chunk_id&lt;/code&gt; payload fields describe the relationship. They do not enforce uniqueness.&lt;/p&gt;

&lt;p&gt;I would keep the first example this small. A file parser and a chunking strategy can come after we have a document we can write and retrieve.&lt;/p&gt;

&lt;h2&gt;
  
  
  Index and read back the document
&lt;/h2&gt;

&lt;p&gt;Append this code directly after your chosen setup block. Do not keep the old collection program in the same file.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;var&lt;/span&gt; &lt;span class="n"&gt;executionTimeout&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt;
    &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;CancellationTokenSource&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;TimeSpan&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;FromMinutes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="m"&gt;2&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;span class="n"&gt;CancellationToken&lt;/span&gt; &lt;span class="n"&gt;token&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;executionTimeout&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Token&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;var&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;QdrantClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;host&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"127.0.0.1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;port&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="m"&gt;6334&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="n"&gt;CollectionInfo&lt;/span&gt; &lt;span class="n"&gt;collection&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GetCollectionInfoAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;collectionName&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="n"&gt;VectorsConfig&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="n"&gt;config&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;collection&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Params&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;VectorsConfig&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;config&lt;/span&gt; &lt;span class="k"&gt;is&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt; &lt;span class="p"&gt;||&lt;/span&gt;
    &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ConfigCase&lt;/span&gt; &lt;span class="p"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;VectorsConfig&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ConfigOneofCase&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Params&lt;/span&gt; &lt;span class="p"&gt;||&lt;/span&gt;
    &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Params&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Size&lt;/span&gt; &lt;span class="p"&gt;!=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;ulong&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="n"&gt;vectorSize&lt;/span&gt; &lt;span class="p"&gt;||&lt;/span&gt;
    &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Params&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Distance&lt;/span&gt; &lt;span class="p"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;Distance&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Cosine&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;InvalidOperationException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="s"&gt;$"Collection '&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;collectionName&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;' does not match the vector configuration."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(!&lt;/span&gt;&lt;span class="n"&gt;collection&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Metadata&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;TryGetValue&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="s"&gt;"embedding-profile-id"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;out&lt;/span&gt; &lt;span class="n"&gt;Value&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="n"&gt;storedProfile&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;||&lt;/span&gt;
    &lt;span class="n"&gt;storedProfile&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;StringValue&lt;/span&gt; &lt;span class="p"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;embeddingProfileId&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;InvalidOperationException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="s"&gt;$"Collection '&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;collectionName&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;' does not match profile '&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;embeddingProfileId&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;'."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="n"&gt;SourceDocument&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt; &lt;span class="n"&gt;documents&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt;
&lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="k"&gt;new&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;PointId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Guid&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"19e8e940-0562-4b6b-bca5-1f99ea154f2e"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;DocumentId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"worker-shutdown"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;Title&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"Stop a background service cleanly"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;Text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"Pass the stopping token to asynchronous I/O in a BackgroundService. "&lt;/span&gt; &lt;span class="p"&gt;+&lt;/span&gt;
              &lt;span class="s"&gt;"When shutdown starts, let the operation observe cancellation "&lt;/span&gt; &lt;span class="p"&gt;+&lt;/span&gt;
              &lt;span class="s"&gt;"and finish within the host shutdown timeout."&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;];&lt;/span&gt;

&lt;span class="k"&gt;foreach&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;SourceDocument&lt;/span&gt; &lt;span class="n"&gt;document&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="n"&gt;documents&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;IsNullOrWhiteSpace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Text&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;InvalidOperationException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="s"&gt;$"Document '&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;DocumentId&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;' has no text."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="n"&gt;ReadOnlyMemory&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;float&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;vector&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;generator&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GenerateVectorAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="nf"&gt;prepareDocument&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;document&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;vector&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Length&lt;/span&gt; &lt;span class="p"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;vectorSize&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;InvalidOperationException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="s"&gt;$"Expected &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;vectorSize&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt; values, received &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;vector&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Length&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="kt"&gt;var&lt;/span&gt; &lt;span class="n"&gt;point&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="n"&gt;PointStruct&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;Id&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;PointId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;Vectors&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;vector&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ToArray&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
        &lt;span class="n"&gt;Payload&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"document_id"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;DocumentId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"chunk_id"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"whole-document"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Text&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;};&lt;/span&gt;

    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;UpsertAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;collectionName&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;collectionName&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;points&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;point&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="n"&gt;wait&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="kt"&gt;var&lt;/span&gt; &lt;span class="n"&gt;stored&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;RetrieveAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;collectionName&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;collectionName&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;PointId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;withPayload&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;withVectors&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;stored&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Count&lt;/span&gt; &lt;span class="p"&gt;!=&lt;/span&gt; &lt;span class="m"&gt;1&lt;/span&gt; &lt;span class="p"&gt;||&lt;/span&gt;
        &lt;span class="p"&gt;!&lt;/span&gt;&lt;span class="n"&gt;stored&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="m"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;Payload&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;TryGetValue&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;out&lt;/span&gt; &lt;span class="n"&gt;Value&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="n"&gt;storedText&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;||&lt;/span&gt;
        &lt;span class="n"&gt;storedText&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;StringValue&lt;/span&gt; &lt;span class="p"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Text&lt;/span&gt; &lt;span class="p"&gt;||&lt;/span&gt;
        &lt;span class="p"&gt;!&lt;/span&gt;&lt;span class="n"&gt;stored&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="m"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;Payload&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;TryGetValue&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;out&lt;/span&gt; &lt;span class="n"&gt;Value&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="n"&gt;storedTitle&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;InvalidOperationException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="s"&gt;$"Read-back failed for '&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;DocumentId&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;'."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="n"&gt;Console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;WriteLine&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;$"Indexed '&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;DocumentId&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;' into '&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;collectionName&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;'."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="n"&gt;Console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;WriteLine&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;$"Point ID: &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;PointId&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="n"&gt;Console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;WriteLine&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;$"Title: &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;storedTitle&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;StringValue&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="n"&gt;Console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;WriteLine&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;$"Text: &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;storedText&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;StringValue&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;sealed&lt;/span&gt; &lt;span class="k"&gt;record&lt;/span&gt; &lt;span class="nc"&gt;SourceDocument&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;Guid&lt;/span&gt; &lt;span class="n"&gt;PointId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;DocumentId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;Title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;Text&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The program checks the existing collection before making an embedding request. A missing collection fails at &lt;code&gt;GetCollectionInfoAsync&lt;/code&gt;. Create it with the setup program first. A wrong size, distance, or declared profile ID stops the write. Matching that ID does not prove which model artifact generated a vector. As in the collection article, these checks cover the simple unnamed dense-vector setup, not every schema option Qdrant supports.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;GenerateVectorAsync&lt;/code&gt; returns &lt;code&gt;ReadOnlyMemory&amp;lt;float&amp;gt;&lt;/code&gt;. The official Qdrant client accepts the array assigned to &lt;code&gt;Vectors&lt;/code&gt; as the unnamed vector. Each point also carries the source text because an embedding cannot reconstruct the document when we need to display it later.&lt;/p&gt;

&lt;p&gt;The upsert explicitly uses &lt;code&gt;wait: true&lt;/code&gt;, so the read-back follows a completed update rather than merely an acknowledged request. This verifies that the point can be retrieved by ID and its text payload matches the source. The code also checks that the title field exists before printing it. It does not read the vector back or validate the other payload values. The check says nothing about search quality and does not require an HNSW index for this tiny collection.&lt;/p&gt;

&lt;p&gt;The two-minute timeout bounds this small console run, including model loading for the local path. It is a sample setting, not a recommended ingestion deadline. The code passes the same cancellation token through the embedding and database calls.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run it, then change the text
&lt;/h2&gt;

&lt;p&gt;Run the program:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dotnet run
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For the Azure OpenAI setup, expected output is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Indexed 'worker-shutdown' into 'support-documents-v1'.
Point ID: 19e8e940-0562-4b6b-bca5-1f99ea154f2e
Title: Stop a background service cleanly
Text: Pass the stopping token to asynchronous I/O in a BackgroundService. When shutdown starts, let the operation observe cancellation and finish within the host shutdown timeout.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The Ollama setup prints &lt;code&gt;support-documents-embeddinggemma-v1&lt;/code&gt; instead. Open the &lt;a href="http://127.0.0.1:6333/dashboard" rel="noopener noreferrer"&gt;local Qdrant dashboard&lt;/a&gt;, select that collection, and inspect the point's payload.&lt;/p&gt;

&lt;p&gt;Run the program again with the same input. Qdrant replaces the point with that ID, so the collection still has one point if it was empty before this exercise. Now change the note's text and run it once more. The program generates a new vector and replaces the point's vector and payload together.&lt;/p&gt;

&lt;p&gt;Keep the UUID unchanged during that edit. Calling &lt;code&gt;Guid.NewGuid()&lt;/code&gt; inside the indexing loop would turn each run into another point. In a real source system, persist the assigned UUID or derive point IDs reproducibly from stable source and chunk identities.&lt;/p&gt;

&lt;p&gt;An upsert replaces the existing point, including its payload. Send the complete payload you want to retain. This sample does not use a partial payload update. Also, the repeat run still calls the embedding model. Stable IDs prevent duplicate points, but they do not skip embedding costs or implement change detection.&lt;/p&gt;

&lt;h2&gt;
  
  
  Add two more documents
&lt;/h2&gt;

&lt;p&gt;Replace the &lt;code&gt;documents&lt;/code&gt; array with this version. The loop stays unchanged:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="n"&gt;SourceDocument&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt; &lt;span class="n"&gt;documents&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt;
&lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="k"&gt;new&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;Guid&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"19e8e940-0562-4b6b-bca5-1f99ea154f2e"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="s"&gt;"worker-shutdown"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="s"&gt;"Stop a background service cleanly"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="s"&gt;"Pass the stopping token to asynchronous I/O in a BackgroundService. "&lt;/span&gt; &lt;span class="p"&gt;+&lt;/span&gt;
        &lt;span class="s"&gt;"When shutdown starts, let the operation observe cancellation "&lt;/span&gt; &lt;span class="p"&gt;+&lt;/span&gt;
        &lt;span class="s"&gt;"and finish within the host shutdown timeout."&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="k"&gt;new&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;Guid&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"81852b6d-d848-4a62-a777-e45b75f56755"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="s"&gt;"http-client-lifetime"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="s"&gt;"Create HTTP clients through the factory"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="s"&gt;"Use IHttpClientFactory to create named HTTP clients. "&lt;/span&gt; &lt;span class="p"&gt;+&lt;/span&gt;
        &lt;span class="s"&gt;"The factory manages handler lifetimes so callers can dispose "&lt;/span&gt; &lt;span class="p"&gt;+&lt;/span&gt;
        &lt;span class="s"&gt;"their clients without creating a new connection pool each time."&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="k"&gt;new&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;Guid&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"fcad478c-cdae-428e-8784-c07d7ec9b8c3"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="s"&gt;"bounded-work-queue"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="s"&gt;"Limit queued background work"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="s"&gt;"Use a bounded Channel to limit queued work. With FullMode set "&lt;/span&gt; &lt;span class="p"&gt;+&lt;/span&gt;
        &lt;span class="s"&gt;"to Wait, producers await available capacity instead of letting "&lt;/span&gt; &lt;span class="p"&gt;+&lt;/span&gt;
        &lt;span class="s"&gt;"an in-memory backlog grow without a limit."&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;];&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After a successful run, the collection contains these three points, plus any unrelated points already present. Repeating the run preserves the same three IDs.&lt;/p&gt;

&lt;p&gt;Each pass through the loop embeds one note, writes it, and checks the stored text. If the second note fails, the first point remains in Qdrant. There is no transaction around the whole array. For a larger job, I would handle batching and recovery explicitly.&lt;/p&gt;

&lt;h2&gt;
  
  
  When to use this indexing path
&lt;/h2&gt;

&lt;p&gt;Use this example to populate a local collection with a few short documents and verify that the application can write and retrieve their payloads. It also leaves a small corpus for the next step: embedding a question and searching the chosen collection.&lt;/p&gt;

&lt;p&gt;Do not use the loop unchanged for long files or continuous source synchronization. Long documents need chunking and input-size checks. A changing corpus needs a policy for updates, removed documents, and interrupted runs. Removing an item from this array does not delete its existing Qdrant point.&lt;/p&gt;

&lt;p&gt;When you add search, use the same embedding profile that produced these points. That includes the model identity, dimensions, distance, and the document/query formatting defined by the profile. Azure OpenAI queries stay with plain text. EmbeddingGemma queries need its retrieval query prefix. Then try a question whose answer you know is in one of the notes. Reading a point by ID cannot tell you whether semantic search will find it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/posts/your-first-qdrant-collection/" rel="noopener noreferrer"&gt;Your First Qdrant Collection&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/posts/embeddings-in-dotnet/" rel="noopener noreferrer"&gt;Embeddings in .NET&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/tips/separate-ingestion-health-from-retrieval-readiness/" rel="noopener noreferrer"&gt;Separate ingestion health from retrieval readiness&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Qdrant: &lt;a href="https://qdrant.tech/documentation/manage-data/points/" rel="noopener noreferrer"&gt;Points, IDs, and updates&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Qdrant: &lt;a href="https://github.com/qdrant/qdrant-dotnet" rel="noopener noreferrer"&gt;.NET client&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/ai/iembeddinggenerator" rel="noopener noreferrer"&gt;Use the IEmbeddingGenerator interface&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Ollama: &lt;a href="https://docs.ollama.com/capabilities/embeddings" rel="noopener noreferrer"&gt;Embeddings&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Google: &lt;a href="https://ai.google.dev/gemma/docs/embeddinggemma/model_card" rel="noopener noreferrer"&gt;EmbeddingGemma model card and input formats&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>dotnet</category>
      <category>csharp</category>
      <category>rag</category>
      <category>ai</category>
    </item>
    <item>
      <title>What Should You Actually Evaluate?</title>
      <dc:creator>Lukas Walter </dc:creator>
      <pubDate>Thu, 17 Sep 2026 15:30:00 +0000</pubDate>
      <link>https://dev.to/lukaswalter/what-should-you-actually-evaluate-41d1</link>
      <guid>https://dev.to/lukaswalter/what-should-you-actually-evaluate-41d1</guid>
      <description>&lt;p&gt;Evaluate whether the application delivered the promised outcome, within its permissions and runtime limits. A good answer is one part of that result.&lt;/p&gt;

&lt;p&gt;A support assistant can write an accurate reply and save it to the wrong ticket. Or it can cite an obsolete policy and give advice that sounds well supported. Even a correct save can finish after the caller has timed out. The final text alone will not tell you which of these happened.&lt;/p&gt;

&lt;p&gt;I would write down what the feature promises, then decide what evidence would show that it kept that promise. Choose the metrics from there.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with one outcome you can verify
&lt;/h2&gt;

&lt;p&gt;Consider a support assistant with this contract:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Prepare a reply from the ticket and the support policy applicable to that ticket. Save it as a draft only when requested. Never send it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The requester must have access to the ticket. If the available policy does not support an answer, the assistant should explain the gap and request review. If saving fails, it must not claim that a draft exists.&lt;/p&gt;

&lt;p&gt;For this worked example, each evaluation case identifies the applicable policy version and the effective timestamp used to select it. "Current" means current for that case, not whatever document happens to be newest when the suite runs. Keep those reference facts fixed when comparing candidates. Historical cases keep their original policy version. Review any new or changed cases intended to represent a newer policy. The contract gives us these checks:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Boundary&lt;/th&gt;
&lt;th&gt;What to verify&lt;/th&gt;
&lt;th&gt;Evidence to inspect&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Reply&lt;/td&gt;
&lt;td&gt;Addresses the request, includes required conditions, makes no unsupported commitments&lt;/td&gt;
&lt;td&gt;Ticket, applicable policy, generated reply&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Retrieval quality&lt;/td&gt;
&lt;td&gt;Surfaces the information needed to answer&lt;/td&gt;
&lt;td&gt;Labeled relevant passages, retrieved IDs, final model context&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Policy applicability&lt;/td&gt;
&lt;td&gt;Uses the policy that applies to this ticket at the case's effective time&lt;/td&gt;
&lt;td&gt;Policy version, effective timestamp, ticket facts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Authorization&lt;/td&gt;
&lt;td&gt;Permits disclosure to the caller or model and any protected state change&lt;/td&gt;
&lt;td&gt;Requester identity, resource metadata, authorization decisions, disclosure and write records&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tools&lt;/td&gt;
&lt;td&gt;Saves only when requested, against the right ticket, without sending&lt;/td&gt;
&lt;td&gt;Tool requests, authorization decisions, persisted state&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Workflow&lt;/td&gt;
&lt;td&gt;Returns an outcome consistent with what actually happened&lt;/td&gt;
&lt;td&gt;Application result, dependency outcomes, state after execution&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Safety&lt;/td&gt;
&lt;td&gt;Keeps protected data and disallowed actions outside the permitted path&lt;/td&gt;
&lt;td&gt;Adversarial cases, boundary decisions, exposed context and outputs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Runtime&lt;/td&gt;
&lt;td&gt;Meets the caller's deadline and the backend execution budget&lt;/td&gt;
&lt;td&gt;Caller and backend outcomes and timings, usage, retries, and tool calls&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These checks need different methods. Some require judgment about language. Others are ordinary assertions over IDs, state, and side effects.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://www.lukaswalter.dev/posts/eval-first-in-.net/" rel="noopener noreferrer"&gt;eval-first article&lt;/a&gt; covers how to bring these checks into the development workflow. First, decide which behaviors need them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Separate answer correctness from groundedness
&lt;/h2&gt;

&lt;p&gt;A grounded answer follows the supplied context. That does not establish that the context is current, complete, or applicable to this customer.&lt;/p&gt;

&lt;p&gt;Suppose an old policy says customers have 30 days to request a replacement. The current policy says 14 days, with an exception for damaged deliveries. An answer based on the old document might be perfectly grounded and still wrong for the task.&lt;/p&gt;

&lt;p&gt;I would score these questions separately:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Criterion&lt;/th&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Correctness&lt;/td&gt;
&lt;td&gt;Do the material claims and conclusion match the applicable policy and ticket facts?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Groundedness&lt;/td&gt;
&lt;td&gt;Does the evidence supplied to the model support its factual claims?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Completeness&lt;/td&gt;
&lt;td&gt;Does the reply include the conditions or exceptions the customer needs to act?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Relevance&lt;/td&gt;
&lt;td&gt;Does it answer this customer's request?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A reference answer can help, but acceptable wording should not be limited to one sentence. Write the required facts and forbidden claims explicitly.&lt;/p&gt;

&lt;p&gt;For a damaged-delivery case, a useful rubric might require the reply to identify the exception, ask for missing evidence required by the policy, and avoid promising approval before that evidence is reviewed. A fluent reply that promises an immediate replacement fails the task.&lt;/p&gt;

&lt;p&gt;Evaluate abstention separately. On cases where the policy cannot answer the question, does the assistant acknowledge the gap? On answerable cases, does it unnecessarily refuse? Combining those groups can reward an assistant that avoids mistakes by declining everything.&lt;/p&gt;

&lt;p&gt;Tone matters when the product requires it. For this assistant, I would prioritize unsupported commitments and omitted conditions before spending time distinguishing polished prose from slightly awkward prose.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check retrieval before blaming the prompt
&lt;/h2&gt;

&lt;p&gt;An incorrect answer leaves several possible causes. The relevant policy may be absent from the corpus, missed by search, removed during context assembly, or ignored by the model.&lt;/p&gt;

&lt;p&gt;Inspect both the search results and the context actually passed to generation. Finding the right document is insufficient if the passage containing the exception never reaches the model.&lt;/p&gt;

&lt;p&gt;With labeled relevant passages, recall at a chosen result count measures the fraction of relevant passages retrieval found. Precision measures the fraction of returned passages that are relevant. Define the relevance unit first: document, chunk, or required fact. Ten overlapping chunks from one paragraph should not look like ten independent pieces of useful evidence.&lt;/p&gt;

&lt;p&gt;For the support example, I would also check whether the final context contains every policy fact needed to answer the ticket. A missing exception then counts as a failure even if the rest of the result list looks reasonable.&lt;/p&gt;

&lt;p&gt;Check authorization and policy applicability separately from retrieval quality. A relevant passage from another tenant must not reach the caller or model without permission. A policy that was obsolete at the case's effective time is unsuitable evidence even if it ranks first.&lt;/p&gt;

&lt;p&gt;To investigate a failure, compare the original context with a manually verified context in repeated, paired runs. Keep the model, instructions, generation settings, and scoring rubric fixed. A consistent improvement with verified context is stronger evidence of a retrieval or context-assembly problem than one successful rerun. If both conditions keep failing, inspect the instructions and generation behavior too. This experiment helps narrow the search. It does not establish a single cause or replace evaluation of the complete application.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check what the tools changed
&lt;/h2&gt;

&lt;p&gt;Calling &lt;code&gt;SaveDraft&lt;/code&gt; looks correct when the user asks to save a reply. The tool name alone tells us little.&lt;/p&gt;

&lt;p&gt;Check which ticket ID and body reached the tool, whether access was enforced, and what was stored. A valid JSON payload can contain the wrong ticket ID. A successful tool response can also be followed by an inaccurate message to the user.&lt;/p&gt;

&lt;p&gt;For this assistant, the important assertions include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A request to prepare a reply without saving produces no write.&lt;/li&gt;
&lt;li&gt;A save request stores the intended body against the authorized ticket.&lt;/li&gt;
&lt;li&gt;Retrying the same logical save operation with the same operation or idempotency identity does not create another draft. Independent save requests follow the product's rules for new drafts or revisions.&lt;/li&gt;
&lt;li&gt;No execution sends a reply.&lt;/li&gt;
&lt;li&gt;A failed or uncertain save does not produce a confirmed-success result.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Run side-effecting scenarios against isolated test resources or controlled tool implementations. A recording fake can show what the application attempted. An integration test is needed to verify the real adapter and storage behavior. These are ordinary software tests, but they still belong in the evaluation map when their results determine whether the AI feature fulfilled its contract.&lt;/p&gt;

&lt;p&gt;Inspect the sequence where order matters. Authorization must succeed before protected contents reach the caller or model, and before a protected state change. The application may need to load a resource or its authorization metadata to make that decision. ASP.NET Core supports this through &lt;a href="https://learn.microsoft.com/en-us/aspnet/core/security/authorization/resource-based?view=aspnetcore-10.0" rel="noopener noreferrer"&gt;resource-based authorization&lt;/a&gt;. When trusted tenant identity is already available, constrain retrieval queries to that tenant before fetching content. Independent lookups can still run in either order. One successful trace should not become the only permitted sequence.&lt;/p&gt;

&lt;p&gt;Microsoft's &lt;a href="https://learn.microsoft.com/en-us/azure/foundry/concepts/evaluation-evaluators/agent-evaluators?preserve-view=true&amp;amp;view=foundry" rel="noopener noreferrer"&gt;agent evaluation documentation&lt;/a&gt; distinguishes task outcomes from the process used to reach them, including tool usage. That is a useful distinction, but scoring tool messages does not replace checking the application's actual side effects.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make failure outcomes part of task success
&lt;/h2&gt;

&lt;p&gt;Successful task completion depends on the scenario. When policy evidence is missing, requesting review can be the correct outcome. When the user lacks access, refusing the operation is correct.&lt;/p&gt;

&lt;p&gt;Build these cases from the feature's &lt;a href="https://www.lukaswalter.dev/posts/failure-modes-before-happy-paths/" rel="noopener noreferrer"&gt;failure modes&lt;/a&gt;: empty retrieval, provider timeout, malformed output, unavailable storage, cancellation, and a write whose outcome is uncertain.&lt;/p&gt;

&lt;p&gt;For each case, assert the expected application result and what may already have happened. If storage commits but the acknowledgment is lost, a generic retry test that only counts exceptions misses the important question: did recovery create a duplicate?&lt;/p&gt;

&lt;p&gt;Keep two views in the report. One measures whether the system handled each scenario according to its contract. The other measures how often users received the useful outcome they wanted. A service that handles every outage gracefully can still be unavailable too often.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep safety failures visible
&lt;/h2&gt;

&lt;p&gt;For this support assistant, safety cases should include a ticket containing instructions to send data elsewhere, a request for another tenant's ticket, and a request to send the reply despite the feature's draft-only contract.&lt;/p&gt;

&lt;p&gt;Inspect the evidence at each boundary. Did the application apply the known tenant restriction to the query and authorize disclosure before ticket content reached the caller or model? Did instructions inside the ticket trigger an unauthorized tool attempt, and did the executor block it?&lt;/p&gt;

&lt;p&gt;Record an unsafe proposal separately from an executed action. A blocked attempt shows that a control worked, while still revealing behavior worth investigating. A reassuring final answer does not erase an unauthorized disclosure or prohibited operation earlier in the workflow.&lt;/p&gt;

&lt;p&gt;Content safety checks may also matter, depending on the feature. A classifier for harmful language cannot verify tenant isolation or write authorization. Those require application-level evidence.&lt;/p&gt;

&lt;p&gt;I would make any observed unauthorized disclosure or prohibited side effect block this feature's release. Keep those failures visible rather than averaging them with relevance or tone scores. Zero observed failures in a test set is useful evidence, not proof that no other input can cross the boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measure latency and cost across the complete task
&lt;/h2&gt;

&lt;p&gt;Measure the caller's wait and outcome separately from backend completion time and outcome. A caller might time out after 12 seconds while the backend saves the draft at 19 seconds. The timeout is a user-visible failure. The late save is a side effect that recovery must account for. Correlate both records with the same logical operation. Include retrieval, retries, tool execution, and persistence in the backend measurement.&lt;/p&gt;

&lt;p&gt;For an interactive feature, inspect median and tail latency, such as the 95th percentile, under stated load. For streaming, distinguish the first visible output from completion of the usable result. A quick opening sentence does not mean the draft is ready.&lt;/p&gt;

&lt;p&gt;Report failures and timeouts beside latency. Separate user-initiated cancellation, deadline or budget cancellation, and system or dependency cancellation. Record an unknown cause when you cannot establish it. A user choosing to stop is different from an execution exhausting its budget. Show which outcomes the latency sample includes so a candidate cannot look faster merely by dropping slow, failed requests from the report.&lt;/p&gt;

&lt;p&gt;Cost should cover the boundaries the feature pays for. Name model-only estimates accordingly, and keep missing usage visible. Include failed attempts and retries rather than counting only the final successful call.&lt;/p&gt;

&lt;p&gt;For cost per successful task, first decide which outcomes count as success. If the denominator counts cases handled according to contract, a correct access denial counts. If it counts useful user outcomes, that denial does not. Report the two ratios separately when you need both, alongside their success rates and per-execution cost. A ratio alone can hide expensive outliers or changes in the case mix.&lt;/p&gt;

&lt;p&gt;Choose limits from the product promise. The &lt;a href="https://www.lukaswalter.dev/posts/ai-systems-need-runtime-budgets/" rel="noopener noreferrer"&gt;runtime budgets article&lt;/a&gt; covers enforcement. Evaluation checks whether realistic workloads and failures stay within those limits while still producing useful outcomes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Match the evaluator to the evidence
&lt;/h2&gt;

&lt;p&gt;Use deterministic checks for schema validity, identifiers, permissions, persisted state, and counts. Use an explicit rubric with human review or a model judge for semantic questions such as whether the reply omitted a policy condition.&lt;/p&gt;

&lt;p&gt;In .NET, &lt;a href="https://learn.microsoft.com/en-us/dotnet/ai/evaluation/libraries" rel="noopener noreferrer"&gt;Microsoft.Extensions.AI.Evaluation&lt;/a&gt; supplies relevance, completeness, and groundedness evaluators that work with existing test infrastructure. Its &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.ai.evaluation.quality.completenessevaluator?view=net-10.0-pp" rel="noopener noreferrer"&gt;CompletenessEvaluator&lt;/a&gt; compares the response against supplied ground truth. For groundedness, explicitly pass the grounding evidence from the final context the generation model actually received through &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.ai.evaluation.quality.groundednessevaluatorcontext.-ctor?view=net-10.0-pp" rel="noopener noreferrer"&gt;GroundednessEvaluatorContext&lt;/a&gt;. The evaluator does not recover that context automatically. Decide what evidence and pass criteria you need before choosing either.&lt;/p&gt;

&lt;p&gt;Give the judge the evidence its rubric asks it to assess: the supplied context for groundedness, or the applicable reference facts for correctness. You still need to inspect storage to verify that a draft exists. The phrase "Draft saved" proves nothing about database state.&lt;/p&gt;

&lt;p&gt;Review a sample of judge decisions against human judgments, including failures and borderline cases. Record disagreements and tighten ambiguous rubrics. If the evaluator errors or lacks required evidence, report the result as unscored. Require a valid result that meets the pass criteria at the release gate. Missing or inconclusive results need an explicit decision.&lt;/p&gt;

&lt;p&gt;Keep the component results beside the result for the complete workflow. They should let you trace a failed task back to a missing passage, an unsupported claim, or an incorrect write without losing sight of what the user received.&lt;/p&gt;

&lt;h2&gt;
  
  
  Decide what would block the release
&lt;/h2&gt;

&lt;p&gt;A single overall score hides the reason a candidate improved or regressed. Report results by scenario: routine answers, policy exceptions, missing evidence, denied access, and dependency failures. Include the number of cases in each group.&lt;/p&gt;

&lt;p&gt;For example, a higher average answer score is not enough to accept a candidate that now invents commitments on missing-policy cases. A cheaper candidate is not an improvement if it saves the wrong draft. Define these acceptance rules before reviewing the candidate's results.&lt;/p&gt;

&lt;p&gt;Compare candidates on the same evaluation protocol and inspect per-case changes. Repeat cases where model variation could change the decision. Keep results tied to their versions using the &lt;a href="https://www.lukaswalter.dev/tips/version-prompts-models-datasets-and-evaluation-results-together/" rel="noopener noreferrer"&gt;evaluation provenance guidance&lt;/a&gt; rather than relying on a score copied into a pull request.&lt;/p&gt;

&lt;p&gt;Production evaluation can reveal cases the offline suite missed. Track corrections, abandoned tasks, and sampled quality judgments with their context. User acceptance alone does not establish correctness. The &lt;a href="https://www.lukaswalter.dev/tips/run-production-evaluations-outside-the-user-response-path/" rel="noopener noreferrer"&gt;production evaluation tip&lt;/a&gt; covers how to collect those judgments outside the user response path.&lt;/p&gt;

&lt;h2&gt;
  
  
  When to use this approach
&lt;/h2&gt;

&lt;p&gt;Use this system-level evaluation map when an AI feature retrieves information, invokes tools, changes state, or has meaningful failure and runtime constraints. Start with one workflow and the failures that would change a release decision. Add dimensions as the feature gains responsibilities.&lt;/p&gt;

&lt;h2&gt;
  
  
  When a smaller evaluation is enough
&lt;/h2&gt;

&lt;p&gt;A low-risk text transformation may need only output validity, preservation of required meaning, and latency. It has no retrieval or tool path to evaluate. A local experiment may begin with a few manually reviewed examples, provided that limited evidence is not treated as production readiness.&lt;/p&gt;

&lt;p&gt;For your next change, write one sentence describing the promised outcome. List the unacceptable results, identify the evidence that would expose each one, and choose a check for each. Then add the quality and runtime measures needed to compare acceptable candidates.&lt;/p&gt;

&lt;p&gt;For the support assistant, that means checking the reply against the applicable policy, inspecting the saved draft, and confirming what the caller saw. Those results give you something concrete to review before shipping.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/posts/eval-first-in-.net/" rel="noopener noreferrer"&gt;Eval-first: Why "It Worked Once" Is Not a Sign of Quality&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/posts/quick-tip-2/" rel="noopener noreferrer"&gt;Stop Guessing: Use Golden Datasets for Prompt Evals&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/tips/version-prompts-models-datasets-and-evaluation-results-together/" rel="noopener noreferrer"&gt;Version prompts, models, datasets, and evaluation results together&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/aspnet/core/security/authorization/resource-based?view=aspnetcore-10.0" rel="noopener noreferrer"&gt;Resource-based authorization in ASP.NET Core&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/azure/foundry/concepts/evaluation-evaluators/agent-evaluators?preserve-view=true&amp;amp;view=foundry" rel="noopener noreferrer"&gt;Agent evaluators in Microsoft Foundry&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/ai/evaluation/libraries" rel="noopener noreferrer"&gt;The Microsoft.Extensions.AI.Evaluation libraries&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.ai.evaluation.quality.completenessevaluator?view=net-10.0-pp" rel="noopener noreferrer"&gt;CompletenessEvaluator&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.ai.evaluation.quality.groundednessevaluatorcontext.-ctor?view=net-10.0-pp" rel="noopener noreferrer"&gt;GroundednessEvaluatorContext constructor&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>evals</category>
      <category>testing</category>
    </item>
    <item>
      <title>Dependency Injection for AI Components</title>
      <dc:creator>Lukas Walter </dc:creator>
      <pubDate>Tue, 15 Sep 2026 15:30:00 +0000</pubDate>
      <link>https://dev.to/lukaswalter/dependency-injection-for-ai-components-38o3</link>
      <guid>https://dev.to/lukaswalter/dependency-injection-for-ai-components-38o3</guid>
      <description>&lt;p&gt;An &lt;code&gt;IChatClient&lt;/code&gt; is expected to support concurrent requests unless its implementation documents otherwise. A tool bound to one user's database context cannot safely follow it into a singleton registration.&lt;/p&gt;

&lt;p&gt;I would check that boundary before adding more registrations. A tool delegate can quietly retain a caller's state for as long as the object that holds it.&lt;/p&gt;

&lt;p&gt;This article assumes you already know the .NET DI container. It uses &lt;code&gt;Microsoft.Extensions.AI&lt;/code&gt; to show where chat clients, prompts, retrieval, and tools belong in the dependency graph. The lifetime decisions also apply when an agent framework sits above the client.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the objects that hold state
&lt;/h2&gt;

&lt;p&gt;Consider a support assistant with a shared model connection and an order lookup tool. The lookup reads through a scoped service that knows which orders the current caller may access.&lt;/p&gt;

&lt;p&gt;If startup code captures that service in a tool delegate and saves the tool on a singleton, the delegate retains a reference to the scoped service. If the service came from a temporary scope, later calls can reach it after that scope has disposed its disposable services. Holding a reference does not prevent that disposal.&lt;/p&gt;

&lt;p&gt;The intended graph is small:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Application lifetime
    IChatClient pipeline
    immutable prompt templates

Request or job scope
    support service
        -&amp;gt; shared IChatClient
        -&amp;gt; authorized retrieval / order lookup

One model operation
    fresh messages
    fresh ChatOptions
    tools bound to this scope
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;a href="https://www.lukaswalter.dev/posts/designing-an-ai-service-layer/" rel="noopener noreferrer"&gt;service-layer article&lt;/a&gt; explains which behavior belongs behind the support feature. Once those dependencies exist, each needs an owner and a lifetime.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Starting lifetime&lt;/th&gt;
&lt;th&gt;What changes the decision&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Chat client and shared middleware&lt;/td&gt;
&lt;td&gt;Singleton&lt;/td&gt;
&lt;td&gt;Every component must support concurrent use and avoid retaining caller state.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fixed prompt template&lt;/td&gt;
&lt;td&gt;Constant or immutable singleton&lt;/td&gt;
&lt;td&gt;A renderer that depends on scoped services or retains scoped state must follow that scope.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Application service&lt;/td&gt;
&lt;td&gt;Scoped when it uses scoped capabilities&lt;/td&gt;
&lt;td&gt;A stateless service with suitable dependencies can be transient.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Retrieval or tool that depends on scoped caller or database state&lt;/td&gt;
&lt;td&gt;Scoped&lt;/td&gt;
&lt;td&gt;Follow the lifetime and concurrency limits of those dependencies.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Messages, call options, bound tool list&lt;/td&gt;
&lt;td&gt;Local to one operation&lt;/td&gt;
&lt;td&gt;Conversation history needs explicit storage and ownership across operations.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;code&gt;IChatClient&lt;/code&gt; expects concurrent use unless an implementation says otherwise. Its arguments have a different contract: implementations may mutate messages or options. Allocate those per operation. The &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.ai.ichatclient?view=net-10.0-pp" rel="noopener noreferrer"&gt;interface remarks&lt;/a&gt; describe both points.&lt;/p&gt;

&lt;h2&gt;
  
  
  Register one client at the composition root
&lt;/h2&gt;

&lt;p&gt;The following sample uses .NET 10, &lt;a href="https://www.nuget.org/packages/Microsoft.Extensions.Hosting/10.0.11" rel="noopener noreferrer"&gt;&lt;code&gt;Microsoft.Extensions.Hosting&lt;/code&gt; 10.0.11&lt;/a&gt;, &lt;code&gt;Microsoft.Extensions.AI&lt;/code&gt; 10.9.0, and OllamaSharp 5.4.30. Ollama keeps credentials out of this local example. Wire a hosted provider's adapter and authentication at the composition root. The client factory can resolve a separately registered application-wide credential.&lt;/p&gt;

&lt;p&gt;Create a console project:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dotnet new console &lt;span class="nt"&gt;-n&lt;/span&gt; AiComposition &lt;span class="nt"&gt;-f&lt;/span&gt; net10.0
&lt;span class="nb"&gt;cd &lt;/span&gt;AiComposition
dotnet add package Microsoft.Extensions.Hosting &lt;span class="nt"&gt;--version&lt;/span&gt; 10.0.11
dotnet add package Microsoft.Extensions.AI &lt;span class="nt"&gt;--version&lt;/span&gt; 10.9.0
dotnet add package OllamaSharp &lt;span class="nt"&gt;--version&lt;/span&gt; 5.4.30
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Replace &lt;code&gt;Program.cs&lt;/code&gt; with this startup code, then append the types from the next two sections:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Microsoft.Extensions.AI&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Microsoft.Extensions.DependencyInjection&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Microsoft.Extensions.Hosting&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Microsoft.Extensions.Options&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;OllamaSharp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;IHost&lt;/span&gt; &lt;span class="n"&gt;host&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Host&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;CreateDefaultBuilder&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;UseDefaultServiceProvider&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;options&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ValidateScopes&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ValidateOnBuild&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ConfigureServices&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;services&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;services&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;AddOptions&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;SupportModelOptions&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;()&lt;/span&gt;
            &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Bind&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Configuration&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GetSection&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"AI:Support"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
            &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Validate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;o&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;Uri&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;TryCreate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                    &lt;span class="n"&gt;o&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Endpoint&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;UriKind&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Absolute&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;out&lt;/span&gt; &lt;span class="kt"&gt;var&lt;/span&gt; &lt;span class="n"&gt;uri&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="p"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;uri&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Scheme&lt;/span&gt; &lt;span class="p"&gt;==&lt;/span&gt; &lt;span class="s"&gt;"http"&lt;/span&gt; &lt;span class="p"&gt;||&lt;/span&gt; &lt;span class="n"&gt;uri&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Scheme&lt;/span&gt; &lt;span class="p"&gt;==&lt;/span&gt; &lt;span class="s"&gt;"https"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
                &lt;span class="s"&gt;"AI:Support:Endpoint must be an HTTP(S) URI."&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Validate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;o&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;!&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;IsNullOrWhiteSpace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;o&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Model&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
                &lt;span class="s"&gt;"AI:Support:Model is required."&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ValidateOnStart&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

        &lt;span class="n"&gt;services&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;AddSingleton&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;IChatClient&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;(&lt;/span&gt;&lt;span class="n"&gt;sp&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="kt"&gt;var&lt;/span&gt; &lt;span class="n"&gt;settings&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;sp&lt;/span&gt;
                &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;GetRequiredService&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;IOptions&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;SupportModelOptions&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&amp;gt;()&lt;/span&gt;
                &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Value&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

            &lt;span class="n"&gt;IChatClient&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;OllamaApiClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;Uri&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;settings&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Endpoint&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;settings&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Model&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;AsBuilder&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
                &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;UseFunctionInvocation&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
                &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Build&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
        &lt;span class="p"&gt;});&lt;/span&gt;

        &lt;span class="n"&gt;services&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;AddScoped&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;IOrderStatusReader&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;DemoOrderStatusReader&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;();&lt;/span&gt;
        &lt;span class="n"&gt;services&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;AddScoped&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;SupportAssistant&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;();&lt;/span&gt;
    &lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Build&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;host&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;StartAsync&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;var&lt;/span&gt; &lt;span class="n"&gt;scope&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;host&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Services&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;CreateAsyncScope&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kt"&gt;var&lt;/span&gt; &lt;span class="n"&gt;assistant&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;scope&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ServiceProvider&lt;/span&gt;
        &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;GetRequiredService&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;SupportAssistant&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;();&lt;/span&gt;

    &lt;span class="n"&gt;Console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;WriteLine&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"AI service graph resolved."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Contains&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"--ask"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;var&lt;/span&gt; &lt;span class="n"&gt;timeout&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;CancellationTokenSource&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;TimeSpan&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;FromSeconds&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="m"&gt;30&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;

        &lt;span class="n"&gt;Console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;WriteLine&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;assistant&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;AnswerAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="s"&gt;"What is the status of order DEMO-42?"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Token&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;host&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;StopAsync&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;sealed&lt;/span&gt; &lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;SupportModelOptions&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;Endpoint&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;get&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;set&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;""&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;Model&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;get&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;set&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;""&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The factory constructs the complete client pipeline. It resolves application-wide settings, but never an order reader or current user. &lt;code&gt;UseFunctionInvocation&lt;/code&gt; supplies the execution mechanism. The feature supplies tools for each call. Microsoft's &lt;a href="https://learn.microsoft.com/en-us/dotnet/ai/ichatclient#tool-calling" rel="noopener noreferrer"&gt;tool-calling example&lt;/a&gt; shows this same separation between the pipeline and call options.&lt;/p&gt;

&lt;p&gt;I use an explicit singleton factory to make ownership visible. &lt;code&gt;AddChatClient&lt;/code&gt; is also available when you want the library's DI builder API.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;ValidateOnStart&lt;/code&gt; checks the configured values when the host starts. It does not prove that the endpoint is reachable or that the model supports tools. The &lt;a href="https://learn.microsoft.com/en-us/dotnet/core/extensions/options#options-validation" rel="noopener noreferrer"&gt;options documentation&lt;/a&gt; explains that validation lifecycle. This client reads its settings once. Changing configuration later will not rebuild it. Use a deliberate replacement strategy if you need live model changes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bind tools inside the operation
&lt;/h2&gt;

&lt;p&gt;Append this service to &lt;code&gt;Program.cs&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;sealed&lt;/span&gt; &lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;SupportAssistant&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;IChatClient&lt;/span&gt; &lt;span class="n"&gt;chatClient&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;IOrderStatusReader&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;private&lt;/span&gt; &lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;Instructions&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"""
&lt;/span&gt;        &lt;span class="n"&gt;Help&lt;/span&gt; &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;order&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt; &lt;span class="n"&gt;Use&lt;/span&gt; &lt;span class="n"&gt;lookup_order_status&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;order&lt;/span&gt; &lt;span class="n"&gt;facts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;
        &lt;span class="n"&gt;If&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;lookup&lt;/span&gt; &lt;span class="n"&gt;returns&lt;/span&gt; &lt;span class="n"&gt;no&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;say&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="k"&gt;is&lt;/span&gt; &lt;span class="n"&gt;unavailable&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;
        &lt;span class="n"&gt;Do&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="n"&gt;invent&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;
        &lt;span class="s"&gt;""";
&lt;/span&gt;
    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="n"&gt;Task&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;AnswerAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;CancellationToken&lt;/span&gt; &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;ArgumentException&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ThrowIfNullOrWhiteSpace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

        &lt;span class="n"&gt;List&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;ChatMessage&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;messages&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt;
        &lt;span class="p"&gt;[&lt;/span&gt;
            &lt;span class="k"&gt;new&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ChatRole&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;System&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Instructions&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
            &lt;span class="k"&gt;new&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ChatRole&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;User&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;];&lt;/span&gt;

        &lt;span class="n"&gt;ChatOptions&lt;/span&gt; &lt;span class="n"&gt;options&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="n"&gt;Tools&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt;
            &lt;span class="p"&gt;[&lt;/span&gt;
                &lt;span class="n"&gt;AIFunctionFactory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                    &lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;FindVisibleAsync&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                    &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"lookup_order_status"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                    &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"Look up an order status by order number."&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="p"&gt;]&lt;/span&gt;
        &lt;span class="p"&gt;};&lt;/span&gt;

        &lt;span class="n"&gt;ChatResponse&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;chatClient&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GetResponseAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Text&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The method creates a tool bound to this scope's &lt;code&gt;orders&lt;/code&gt; instance. The caller awaits the entire operation before disposing its scope, so the tool can use its scoped dependencies until the call completes. A streaming implementation must defer scope disposal until enumeration finishes too.&lt;/p&gt;

&lt;p&gt;The fixed instructions need no DI registration. If several features share an immutable template catalog, registering it as a singleton can be reasonable. Rendered messages still belong to the operation. Putting them on the catalog would mix one caller's content with another's.&lt;/p&gt;

&lt;p&gt;A scoped tool is not automatically safe for parallel invocation. For example, two tools in one scope could share the same EF Core &lt;code&gt;DbContext&lt;/code&gt;. Keep their execution sequential or give each parallel operation its own correctly initialized scope and context. &lt;code&gt;FunctionInvokingChatClient.AllowConcurrentInvocation&lt;/code&gt; defaults to &lt;code&gt;false&lt;/code&gt;, so the middleware executes function calls within a request sequentially. Separate requests can still invoke their tools concurrently, as the &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.ai.functioninvokingchatclient" rel="noopener noreferrer"&gt;concurrency remarks&lt;/a&gt; explain.&lt;/p&gt;

&lt;p&gt;This sample returns plain text to keep the lifetime example readable. A real support feature still needs response acceptance rules and a tool-loop budget. The prompt cannot enforce either requirement.&lt;/p&gt;

&lt;h2&gt;
  
  
  Give retrieval and tools their application dependencies
&lt;/h2&gt;

&lt;p&gt;Append a small lookup contract and a demo implementation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;interface&lt;/span&gt; &lt;span class="nc"&gt;IOrderStatusReader&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;Task&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;?&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;FindVisibleAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;orderNumber&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;CancellationToken&lt;/span&gt; &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;sealed&lt;/span&gt; &lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;DemoOrderStatusReader&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;IOrderStatusReader&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="n"&gt;Task&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;?&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;FindVisibleAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;orderNumber&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;CancellationToken&lt;/span&gt; &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ThrowIfCancellationRequested&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;Task&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;FromResult&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;?&amp;gt;(&lt;/span&gt;
            &lt;span class="n"&gt;orderNumber&lt;/span&gt; &lt;span class="p"&gt;==&lt;/span&gt; &lt;span class="s"&gt;"DEMO-42"&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="s"&gt;"Dispatched"&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The demo contains one public fixture and has no authentication or database. In an application, replace it with a service that checks the authenticated caller's access before loading an order. Keep that service scoped when it depends on scoped caller context or a scoped database context. The tool accepts an order number. Trusted application context supplies identity and tenant scope.&lt;/p&gt;

&lt;p&gt;Register an authorized retrieval service the same way when it depends on a scoped database context or caller context. Its lower-level vector client or embedding generator may be reusable across scopes if the implementation supports that. A shared connection does not require a shared authorization context. Database access alone does not dictate a scoped lifetime: a service can use &lt;a href="https://learn.microsoft.com/en-us/ef/core/dbcontext-configuration/#use-a-dbcontext-factory" rel="noopener noreferrer"&gt;&lt;code&gt;IDbContextFactory&amp;lt;TContext&amp;gt;&lt;/code&gt;&lt;/a&gt; to create and dispose a context for each operation. Its own lifetime still depends on the other state and dependencies it holds.&lt;/p&gt;

&lt;p&gt;Scope also has a precise limit: in ASP.NET Core, it normally means one HTTP request. It does not mean one conversation. Store conversation history under an authorized conversation ID and load the allowed messages for each operation. A scoped &lt;code&gt;List&amp;lt;ChatMessage&amp;gt;&lt;/code&gt; will disappear with the request. A singleton list will be shared across callers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Let the owner dispose the client
&lt;/h2&gt;

&lt;p&gt;The DI container owns the client returned by the singleton factory. At shutdown, disposing the outer pipeline also disposes its inner client. Consumers should await their work and leave disposal to that owner. See the &lt;a href="https://learn.microsoft.com/en-us/dotnet/core/extensions/dependency-injection/guidelines#disposal-of-services" rel="noopener noreferrer"&gt;.NET disposal guidance&lt;/a&gt; and &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.ai.delegatingchatclient.dispose" rel="noopener noreferrer"&gt;&lt;code&gt;DelegatingChatClient.Dispose&lt;/code&gt;&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Do not put an injected &lt;code&gt;IChatClient&lt;/code&gt; in a &lt;code&gt;using&lt;/code&gt; block inside &lt;code&gt;AnswerAsync&lt;/code&gt;. That would dispose the shared instance while other requests may still need it. For the same reason, avoid separately registering and owning the inner client when a disposing wrapper already owns it.&lt;/p&gt;

&lt;p&gt;Creating a scope around a tool delegate and then returning the delegate does not transfer ownership. The delegate still refers to the scoped instance after disposal. Keep the operation inside the scope instead.&lt;/p&gt;

&lt;p&gt;A background worker needs to create that scope explicitly. Hosted services have no automatic per-job scope. Resolve the feature inside a fresh async scope for each job and await completion before disposing it, as in Microsoft's &lt;a href="https://learn.microsoft.com/en-us/dotnet/core/extensions/scoped-service" rel="noopener noreferrer"&gt;scoped worker example&lt;/a&gt;. A new scope does not supply an HTTP user. Rebuild execution context from validated job identity data and re-evaluate authorization where permissions may have changed since enqueueing. Stored identity alone does not establish current permission.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check composition without asking a model
&lt;/h2&gt;

&lt;p&gt;Set the endpoint and model through configuration:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;AI__Support__Endpoint&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;http://localhost:11434
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;AI__Support__Model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;llama3.1
dotnet run
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Without &lt;code&gt;--ask&lt;/code&gt;, the sample starts the host and resolves the scoped feature. It constructs the client without sending a model request. You should see &lt;code&gt;AI service graph resolved.&lt;/code&gt; among the host logs.&lt;/p&gt;

&lt;p&gt;Remove the model value and startup should fail options validation. Temporarily register &lt;code&gt;SupportAssistant&lt;/code&gt; as a singleton and scope validation should reject its scoped order reader. These checks answer useful composition questions before a provider enters the picture.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;ValidateOnBuild&lt;/code&gt; cannot inspect arbitrary code inside registration factories. Resolve the feature graph as well, as the sample does, and test any factory branches your application uses. Avoid calling &lt;code&gt;BuildServiceProvider&lt;/code&gt; inside registration code to get dependencies early. Use the provider passed to the registration factory.&lt;/p&gt;

&lt;p&gt;For an optional live check, run Ollama locally, pull &lt;code&gt;llama3.1&lt;/code&gt; (or whatever model you like), then run &lt;code&gt;dotnet run -- --ask&lt;/code&gt;. That exercises the model path. If the model is missing, Ollama returns a &lt;a href="https://docs.ollama.com/api/errors" rel="noopener noreferrer"&gt;404 error&lt;/a&gt;. This sample does not pull it automatically. The live check is separate from the composition check and does not establish answer quality.&lt;/p&gt;

&lt;p&gt;When adding a second client, choose it explicitly at the feature boundary with a keyed registration or a typed adapter. Keep keys fixed in composition code. A user-supplied string should not choose a privileged model connection.&lt;/p&gt;

&lt;h2&gt;
  
  
  When to use this structure
&lt;/h2&gt;

&lt;p&gt;Use a shared client with scoped application capabilities when an AI feature reads caller-specific data or exposes tools backed by scoped services. Build fresh messages and tool bindings for every operation, and dispose the scope only after all work finishes.&lt;/p&gt;

&lt;p&gt;Before shipping, trace one tool delegate back to its target object. Check who owns that object, when it is disposed, and whether two invocations can reach it concurrently. That small review catches problems a successful chat response will never reveal.&lt;/p&gt;

&lt;h2&gt;
  
  
  When not to add more DI machinery
&lt;/h2&gt;

&lt;p&gt;A short console experiment can construct and dispose its client directly. A fixed prompt can remain a constant. Add a factory, interface, or registration extension when it owns a real variation or lifecycle decision.&lt;/p&gt;

&lt;p&gt;If the lifetime is hard to choose, look at the fields and captured delegates. A stored message list or a reference to the current caller often explains why an otherwise reusable component cannot be shared.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/posts/designing-an-ai-service-layer/" rel="noopener noreferrer"&gt;Designing an AI Service Layer&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/posts/agentframework_1_7/" rel="noopener noreferrer"&gt;Tools and Dependency Injection in Microsoft Agent Framework&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/posts/provider-independence-from-day-one/" rel="noopener noreferrer"&gt;Provider Independence from Day One&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.ai.ichatclient?view=net-10.0-pp" rel="noopener noreferrer"&gt;&lt;code&gt;IChatClient&lt;/code&gt; concurrency and argument-mutation contract&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/ai/ichatclient#tool-calling" rel="noopener noreferrer"&gt;Tool calling with &lt;code&gt;IChatClient&lt;/code&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.ai.functioninvokingchatclient" rel="noopener noreferrer"&gt;&lt;code&gt;FunctionInvokingChatClient&lt;/code&gt; and tool concurrency&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/core/extensions/dependency-injection/guidelines#disposal-of-services" rel="noopener noreferrer"&gt;Dependency injection and disposal ownership&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.ai.delegatingchatclient.dispose" rel="noopener noreferrer"&gt;&lt;code&gt;DelegatingChatClient.Dispose&lt;/code&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/core/extensions/options#options-validation" rel="noopener noreferrer"&gt;Options validation and &lt;code&gt;ValidateOnStart&lt;/code&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/core/extensions/scoped-service" rel="noopener noreferrer"&gt;Using scoped services in a background worker&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/ef/core/dbcontext-configuration/#use-a-dbcontext-factory" rel="noopener noreferrer"&gt;Creating EF Core contexts with &lt;code&gt;IDbContextFactory&lt;/code&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Ollama: &lt;a href="https://docs.ollama.com/api/errors" rel="noopener noreferrer"&gt;API errors, including missing models&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;NuGet: &lt;a href="https://www.nuget.org/packages/Microsoft.Extensions.Hosting/10.0.11" rel="noopener noreferrer"&gt;&lt;code&gt;Microsoft.Extensions.Hosting&lt;/code&gt; 10.0.11&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>dotnet</category>
      <category>csharp</category>
      <category>ai</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Human Approval as a System Boundary</title>
      <dc:creator>Lukas Walter </dc:creator>
      <pubDate>Thu, 10 Sep 2026 12:30:00 +0000</pubDate>
      <link>https://dev.to/lukaswalter/human-approval-as-a-system-boundary-31o6</link>
      <guid>https://dev.to/lukaswalter/human-approval-as-a-system-boundary-31o6</guid>
      <description>&lt;p&gt;A user approves a support email. Before the worker sends it, the agent rewrites the body and adds another recipient. The application still has &lt;code&gt;Approved = true&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Human approval is an authorization control at the execution boundary. It must bind the decision to an exact operation: its destination, payload, and execution conditions. The flag says nothing about which version the user accepted.&lt;/p&gt;

&lt;p&gt;I would design that agreement before adding the approval button. Once the reviewer clicks it, the worker needs enough stored information to finish without asking the model what the user meant.&lt;/p&gt;

&lt;h2&gt;
  
  
  Decide where human judgment helps
&lt;/h2&gt;

&lt;p&gt;Suppose a support assistant can read a ticket, draft a reply, and send it. Company policy requires a support lead to review outbound replies because they can make commitments on the company's behalf. The application enforces that requirement.&lt;/p&gt;

&lt;p&gt;Authorization first establishes which operations the requester may propose and for which resources. Application policy then chooses &lt;code&gt;Allow&lt;/code&gt;, &lt;code&gt;Deny&lt;/code&gt;, or &lt;code&gt;RequireApproval&lt;/code&gt;. Approval is part of authorization: a support lead can authorize a reply that the requester cannot send alone. That decision must stay within the operation's permitted scope and cannot override a hard policy denial. This example uses one reviewer. Workflows with several reviewers need separate decisions and an aggregate policy result.&lt;/p&gt;

&lt;p&gt;The broader enforcement design is in &lt;a href="https://www.lukaswalter.dev/posts/trust-boundaries-around-ai-features/" rel="noopener noreferrer"&gt;Trust Boundaries Around AI Features&lt;/a&gt;. This article follows one operation through review and dispatch.&lt;/p&gt;

&lt;p&gt;The useful human decision is whether this reply should go to this customer. Routine authorized lookups need no interruption. The &lt;a href="https://www.lukaswalter.dev/tips/use-approval-for-side-effects-not-for-every-tool-call/" rel="noopener noreferrer"&gt;approval tip&lt;/a&gt; covers that distinction, though a read that exports protected data can still require stricter controls.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prepare something the reviewer can approve
&lt;/h2&gt;

&lt;p&gt;The assistant should finish the permitted preparation and validation before asking. That means a reply with resolved recipients and the intended attachments. Asking "May I send a response?" before one exists leaves the consequential details undecided.&lt;/p&gt;

&lt;p&gt;Store the proposal as an immutable operation revision and use it for both the review screen and executor. That lets the application check the content and destinations it submits to the provider against what the person approved. Downstream processing and the recipient's mail client can still change the delivered message or its appearance.&lt;/p&gt;

&lt;p&gt;For a hypothetical ticket, the reviewer might see:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Field&lt;/th&gt;
&lt;th&gt;Value shown for review&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Operation&lt;/td&gt;
&lt;td&gt;Send a customer reply&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Account and environment&lt;/td&gt;
&lt;td&gt;Example Support, production&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ticket&lt;/td&gt;
&lt;td&gt;T-1842, version 17&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sender&lt;/td&gt;
&lt;td&gt;&lt;a href="mailto:support@example.com"&gt;support@example.com&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Recipients&lt;/td&gt;
&lt;td&gt;To: &lt;a href="mailto:customer@example.net"&gt;customer@example.net&lt;/a&gt;; no CC or BCC; provider envelope has the same single destination&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Subject&lt;/td&gt;
&lt;td&gt;Replacement for order O-731&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Body&lt;/td&gt;
&lt;td&gt;Full text, including the promised replacement date&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Attachments&lt;/td&gt;
&lt;td&gt;replacement-details.pdf, with access to the exact stored file&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Approval scope&lt;/td&gt;
&lt;td&gt;Authorize one logical dispatch of this revision&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Decide by&lt;/td&gt;
&lt;td&gt;2026-09-06 10:15 Europe/Berlin (UTC+02:00)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Commit for dispatch by&lt;/td&gt;
&lt;td&gt;2026-09-06 10:20 Europe/Berlin (UTC+02:00)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These fields come from the stored operation. A generated explanation can help the reviewer understand them, but must not replace or hide them. The review screen needs safe text rendering and access to the attachments.&lt;/p&gt;

&lt;p&gt;Store attachments as immutable versions or protected snapshots. Finish transformations that can change approved details before review. Later transformations, such as provider footers or tracking-link rewriting, need fixed behavior within the approval scope or documented treatment outside it. A filename alone cannot bind approval to a file's contents.&lt;/p&gt;

&lt;p&gt;Show actual addresses as well as display names. Header recipients and the SMTP or API envelope can differ, so the review must cover both. If approval depends on individual recipients, resolve mutable groups or aliases into fixed destinations before review. Approving a group address alone cannot establish who its members will be later.&lt;/p&gt;

&lt;p&gt;This follows the transaction-review principle in the &lt;a href="https://cheatsheetseries.owasp.org/cheatsheets/Transaction_Authorization_Cheat_Sheet.html" rel="noopener noreferrer"&gt;OWASP Transaction Authorization Cheat Sheet&lt;/a&gt;: show the significant operation data, protect it from modification, and enforce the decision on the server.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bind the decision to a stored revision
&lt;/h2&gt;

&lt;p&gt;An approval record needs a stable ID and a reference to the immutable operation revision. Record the requester and tenant, policy version, reviewer requirement, and decision. Store &lt;code&gt;DecideBy&lt;/code&gt; and &lt;code&gt;ExecuteBy&lt;/code&gt; as absolute UTC instants. Display the reviewer's time zone and offset. Keep the logical operation ID stable across attempts to dispatch that revision.&lt;/p&gt;

&lt;p&gt;An executor deployment must preserve the meaning of stored operations. Record the operation type and a schema or handler-contract version when backward-compatible interpretation is not otherwise guaranteed. The review client submits identifiers, and the server loads the stored revision rather than accepting a replacement payload.&lt;/p&gt;

&lt;p&gt;The model supplies candidate content. Authenticated context supplies identity and tenant scope. Application policy sets the approval requirement. The review request only identifies the stored operation and decision:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;enum&lt;/span&gt; &lt;span class="n"&gt;ReviewDecision&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;Approve&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;Reject&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;sealed&lt;/span&gt; &lt;span class="k"&gt;record&lt;/span&gt; &lt;span class="nc"&gt;SubmitReview&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;Guid&lt;/span&gt; &lt;span class="n"&gt;ApprovalId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="kt"&gt;long&lt;/span&gt; &lt;span class="n"&gt;ExpectedOperationRevision&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;ReviewDecision&lt;/span&gt; &lt;span class="n"&gt;Decision&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This DTO sketches the request. The endpoint must validate the decision, load the record inside the authenticated tenant, and authorize the reviewer, including any rule against self-approval. The client supplies neither reviewer identity nor replacement tool arguments.&lt;/p&gt;

&lt;p&gt;Accept the decision only while the record is pending, the operation revision matches, and server time is strictly before &lt;code&gt;DecideBy&lt;/code&gt;. Enforce this at the write, with separate approval-record concurrency control. &lt;code&gt;ExpectedOperationRevision&lt;/code&gt; identifies the payload, not the ticket version, policy version, or concurrency token. Concurrent approve and reject submissions must commit at most one decision. Knowing an approval ID grants no authority to use it.&lt;/p&gt;

&lt;p&gt;If the reviewer edits the body, save a new operation revision and supersede the old approval request. Present the final edited payload for confirmation. An explicit "Save and approve" flow can combine those steps only if it binds the decision to the exact edited revision the reviewer sees.&lt;/p&gt;

&lt;p&gt;After approval, the worker loads that revision directly. Asking the model to reconstruct the email from conversation history would create another proposal.&lt;/p&gt;

&lt;h2&gt;
  
  
  Give waiting its own lifecycle
&lt;/h2&gt;

&lt;p&gt;A reviewer may answer after a restart. Persist the pending request, return control instead of keeping a model request open, and provide a page or endpoint for checking its status. Give resumed execution its own &lt;a href="https://www.lukaswalter.dev/posts/ai-systems-need-runtime-budgets/" rel="noopener noreferrer"&gt;runtime budget&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;CommitForDispatch&lt;/code&gt; is the durable business transition that admits one logical operation to the dispatch process. A worker later claims an individual attempt. That gives the worker ownership of work already authorized by this gate:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Decision:
Pending -&amp;gt; Approved | Rejected | DecisionExpired | Superseded | Withdrawn

CommitForDispatch requires:
decision == Approved
and operation revision is still current and intact
and server time &amp;lt; ExecuteBy
and current policy and resource conditions permit execution
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;An approved decision stays in the history even when &lt;code&gt;ExecuteBy&lt;/code&gt; passes or a new proposal supersedes its operation revision. If the operation has not yet been admitted, it can no longer pass this gate. The record still shows who approved it. &lt;code&gt;DecisionExpired&lt;/code&gt; means nobody decided in time. &lt;code&gt;Withdrawn&lt;/code&gt; means an authorized requester withdrew the pending review without submitting a replacement.&lt;/p&gt;

&lt;p&gt;Withdrawing the pending review prevents later admission. If the withdrawal ends the proposed operation, its logical state becomes &lt;code&gt;Cancelled&lt;/code&gt;. The review remains &lt;code&gt;Withdrawn&lt;/code&gt; to show that no reviewer decided.&lt;/p&gt;

&lt;p&gt;The deadlines can differ: approval at 10:14 may authorize admission at 10:19, but not at 10:20. Check each deadline synchronously at its transition. A delayed cleanup job must not extend either one.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;ExecuteBy&lt;/code&gt; bounds admission, not when the provider receives the request. An operation admitted at 10:19 could remain queued past 10:20. If freshness must hold until provider handoff, enforce a separate deadline and the relevant freshness checks before starting the first provider call. If nobody approves, the reply stays unsent. Reminders cannot authorize it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recheck before committing to dispatch
&lt;/h2&gt;

&lt;p&gt;At 10:00, a support lead approves the reply for ticket version 17. At 10:02, another employee corrects the customer's email address. At 10:03, the worker picks up the approved send.&lt;/p&gt;

&lt;p&gt;The stored decision is still useful history. The worker must now establish whether it can execute under the approved conditions.&lt;/p&gt;

&lt;p&gt;Before admission, the application should:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Load the approved operation and check its integrity, revision, decision, and &lt;code&gt;ExecuteBy&lt;/code&gt; deadline.&lt;/li&gt;
&lt;li&gt;Reload the affected resources through the trusted tenant path and apply the configured requester and reviewer authority rules.&lt;/li&gt;
&lt;li&gt;Compare relevant resource versions and evaluate current policy, including whether it still accepts the recorded approval.&lt;/li&gt;
&lt;li&gt;Persist &lt;code&gt;CommitForDispatch&lt;/code&gt; atomically with the approval and relevant resource checks that share its transactional store.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;In this example, a relevant ticket change blocks the old send and requires a refreshed proposal. Do not silently send to the new address under an approval for the old one. If you choose a narrower version check than the entire ticket row, document which fields invalidate approval and make sure changes to them advance that version.&lt;/p&gt;

&lt;p&gt;For this support workflow, I would require both requester and reviewer to retain their relevant authority until admission commits. Another workflow may accept the request and approval as historical acts, then execute under organizational or service authority. Each identity needs an explicit policy rule, and the executor needs authority to act. A current hard denial still stops execution.&lt;/p&gt;

&lt;p&gt;When approval and relevant resource state share a transactional store, the admission transaction can conditionally check both. When the business effect and result also live in that transaction, they can commit together. &lt;a href="https://learn.microsoft.com/en-us/ef/core/saving/concurrency" rel="noopener noreferrer"&gt;EF Core concurrency tokens&lt;/a&gt; make updates conditional on the original version and report conflicts. They do not pull another service's state into the transaction. Blindly retrying a conflict with new values would discard the approved precondition.&lt;/p&gt;

&lt;p&gt;If the ticket belongs to another service, use its conditional-operation or reservation support where available. A local admission transaction cannot atomically validate that remote ticket or an independent identity service. Without coordination, a time-of-check-to-time-of-use window remains. Document the risk and restrict or block the operation if policy cannot tolerate it.&lt;/p&gt;

&lt;p&gt;Sending the email is external even when ticket and approval share a database. In this example, &lt;code&gt;CommitForDispatch&lt;/code&gt; commits the business decision to dispatch the fixed revision. Coordinate conflicting local changes through that point. If policy must allow cancellation until provider handoff, the dispatch path needs an additional coordinated cancellation check. A queued record alone cannot provide that guarantee.&lt;/p&gt;

&lt;p&gt;Once dispatch has begun, revoking approval cannot reliably recall the email. The UI should report the actual execution state and offer only cancellation or recovery actions the system can still honor.&lt;/p&gt;

&lt;h2&gt;
  
  
  Approval does not make retries safe
&lt;/h2&gt;

&lt;p&gt;Two clicks on "Approve" must not create two logical dispatches. Record one decision, admit the operation once, and return its existing status when a request repeats. Each worker attempt gets a separate ID under that same logical operation.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Logical operation:
NotStarted --CommitForDispatch--&amp;gt; InProgress
NotStarted -&amp;gt; Blocked
NotStarted -&amp;gt; Cancelled
InProgress -&amp;gt; Succeeded | Failed | OutcomeUnknown
InProgress -&amp;gt; Blocked | Cancelled (before first provider handoff)
OutcomeUnknown -&amp;gt; Succeeded | Failed (new evidence)
OutcomeUnknown -&amp;gt; InProgress (safe retry permitted)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;Succeeded&lt;/code&gt; means confirmed provider acceptance of the approved request. Mailbox delivery and bounce tracking are separate. SMTP also distinguishes acceptance from later delivery or failure notification in &lt;a href="https://www.rfc-editor.org/rfc/rfc5321.html#section-6.1" rel="noopener noreferrer"&gt;RFC 5321, section 6.1&lt;/a&gt;. &lt;code&gt;Failed&lt;/code&gt; means failure is established and the workflow will make no more attempts. A definitely failed network attempt does not by itself fail the logical operation. A new attempt may run while the logical operation remains &lt;code&gt;InProgress&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;Blocked&lt;/code&gt; means policy, freshness, or preconditions prevent the first provider call. &lt;code&gt;Cancelled&lt;/code&gt; records a supported cancellation that succeeds before handoff. Both transitions must coordinate with dispatch so no worker can still start that call. Once a call may have been accepted, a policy change or cancellation request cannot establish either outcome. Keep an unresolved call in &lt;code&gt;OutcomeUnknown&lt;/code&gt; until evidence establishes what happened.&lt;/p&gt;

&lt;p&gt;Suppose the provider accepts the request and the worker crashes before saving the response. After restart, local state cannot establish acceptance. Blindly scheduling another attempt could submit the email again. Worker ownership does not resolve that uncertainty: an expired lease is no proof that a provider call stopped.&lt;/p&gt;

&lt;p&gt;Recovery needs durable provider evidence or an idempotency contract that makes another attempt safe. Reuse the logical operation's idempotency key with the same payload, within the provider's scope and retention window. A new attempt ID must not become a new key. An outbox cannot supply deduplication that the provider lacks.&lt;/p&gt;

&lt;p&gt;A saved provider reference or queryable client operation ID may establish what happened. The lost response may also have contained the only usable reference. Without evidence or safe idempotency, keep &lt;code&gt;OutcomeUnknown&lt;/code&gt; and stop automatic retries. Investigation may never resolve it.&lt;/p&gt;

&lt;p&gt;Admission before &lt;code&gt;ExecuteBy&lt;/code&gt; can permit safe recovery of the same logical operation afterward if policy allows it. Current hard denials may stop future attempts, but they cannot undo a call the provider may already have accepted or change &lt;code&gt;OutcomeUnknown&lt;/code&gt; to &lt;code&gt;Blocked&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.lukaswalter.dev/posts/retries-are-not-a-recovery-strategy/" rel="noopener noreferrer"&gt;Retries Are Not a Recovery Strategy&lt;/a&gt; covers attempt recovery in more detail. In this example, the approval record establishes permission to dispatch the reply. It cannot tell the restarted worker whether the provider accepted the request whose response was lost.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test what happens after the click
&lt;/h2&gt;

&lt;p&gt;I would test the approval service and executor directly, without a model in the test. A prompt change should not affect any of these results:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Test&lt;/th&gt;
&lt;th&gt;Required result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Bypass approval or submit as a reviewer from another tenant&lt;/td&gt;
&lt;td&gt;Reject without sending or exposing the operation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Edit or tamper with the reviewed payload&lt;/td&gt;
&lt;td&gt;Require approval of a valid new revision, or reject an integrity failure&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Duplicate approval, or concurrent approve and reject&lt;/td&gt;
&lt;td&gt;Commit one decision and admit at most one logical dispatch&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Approve at or after &lt;code&gt;DecideBy&lt;/code&gt;, or after supersession&lt;/td&gt;
&lt;td&gt;Reject even if cleanup has not run&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Withdraw a pending review, then approve from a stale review page&lt;/td&gt;
&lt;td&gt;Reject the decision and never admit the operation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Approve in time; attempt admission at or after &lt;code&gt;ExecuteBy&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Preserve &lt;code&gt;Approved&lt;/code&gt; history; reject admission&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Requester or reviewer loses required authority before admission&lt;/td&gt;
&lt;td&gt;Block under this example's policy; test historical-authority policies separately&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Change local ticket state between validation and admission&lt;/td&gt;
&lt;td&gt;Reject the stale version in the shared transaction&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Change remote state where conditional execution is supported&lt;/td&gt;
&lt;td&gt;The remote operation rejects the stale version&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Change remote state after final validation but before local admission, with no coordination primitive&lt;/td&gt;
&lt;td&gt;Block if policy requires atomic freshness; otherwise assert admission succeeds against the validated snapshot despite the change, as policy explicitly permits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Restart after a provider call whose result was not saved&lt;/td&gt;
&lt;td&gt;Recover using durable evidence or valid idempotency; otherwise preserve &lt;code&gt;OutcomeUnknown&lt;/code&gt; and do not resend&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deny the operation or miss its configured pre-handoff deadline after admission, but before the first provider call&lt;/td&gt;
&lt;td&gt;Record &lt;code&gt;Blocked&lt;/code&gt; and prevent the first provider call&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cancel an admitted operation through a supported path before first handoff&lt;/td&gt;
&lt;td&gt;Record &lt;code&gt;Cancelled&lt;/code&gt; only after preventing dispatch&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deny or request cancellation while a provider call has an unknown outcome&lt;/td&gt;
&lt;td&gt;Stop future attempts; preserve &lt;code&gt;OutcomeUnknown&lt;/code&gt; until evidence resolves the call&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For diagnosis, record approval, logical operation, and attempt IDs alongside revision, policy version, identities, deadlines, and state transitions. Keep sensitive payloads in protected storage with appropriate retention rather than copying email bodies into general application logs.&lt;/p&gt;

&lt;p&gt;Also inspect the review experience. Track waiting time, edits and rejections, and approvals that miss &lt;code&gt;ExecuteBy&lt;/code&gt;. Those observations help identify a queue that nobody can keep up with. An approval rate alone says little about whether anyone had enough context to review the action.&lt;/p&gt;

&lt;h2&gt;
  
  
  When human approval helps
&lt;/h2&gt;

&lt;p&gt;Use it when an authorized person can assess a concrete operation before its consequences occur. A support lead can check whether the promised replacement is justified and whether the reply says what the company intends to promise. Give that reviewer enough context and time to decide, and define what the application does when nobody responds.&lt;/p&gt;

&lt;p&gt;Do not add it to routine work already covered by an explicit automatic-execution policy. Do not offer an approval button for an action policy forbids. If the reviewer cannot inspect the payload or understand its effect, reduce the operation's scope or keep execution in an established manual process.&lt;/p&gt;

&lt;p&gt;For an existing agent feature, start with one consequential tool. Write down the exact operation a person will review, what changes invalidate that decision, and how the executor proves it is carrying out the approved revision. Then test a changed payload and a worker crash: the first must block the old approval, and the second must preserve enough state to investigate or safely resume the original operation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/posts/agentframework_1_15/" rel="noopener noreferrer"&gt;Human-in-the-Loop Agents: When the AI Must Ask Before Acting&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/tips/use-approval-for-side-effects-not-for-every-tool-call/" rel="noopener noreferrer"&gt;Use approval for side effects, not for every tool call&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/tips/separate-prompts-from-authorization/" rel="noopener noreferrer"&gt;Separate prompts from authorization&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;OWASP: &lt;a href="https://cheatsheetseries.owasp.org/cheatsheets/Transaction_Authorization_Cheat_Sheet.html" rel="noopener noreferrer"&gt;Transaction Authorization Cheat Sheet&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/ef/core/saving/concurrency" rel="noopener noreferrer"&gt;Handling concurrency conflicts in EF Core&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;IETF: &lt;a href="https://www.rfc-editor.org/rfc/rfc5321.html#section-6.1" rel="noopener noreferrer"&gt;RFC 5321, section 6.1: Reliable delivery and replies by email&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Your First Qdrant Collection</title>
      <dc:creator>Lukas Walter </dc:creator>
      <pubDate>Tue, 08 Sep 2026 12:30:00 +0000</pubDate>
      <link>https://dev.to/lukaswalter/your-first-qdrant-collection-1f25</link>
      <guid>https://dev.to/lukaswalter/your-first-qdrant-collection-1f25</guid>
      <description>&lt;p&gt;In &lt;em&gt;&lt;a href="https://www.lukaswalter.dev/posts/embeddings-in-dotnet/" rel="noopener noreferrer"&gt;Embeddings in .NET&lt;/a&gt;&lt;/em&gt;, I stopped at the vector-store boundary. The article generated embeddings and treated the model, dimensions, distance, and input preparation as one contract. It did not create an index for those vectors.&lt;/p&gt;

&lt;p&gt;This article picks up there with a smaller goal: run Qdrant locally, connect from .NET, create one collection, and verify the parts of its configuration that the application depends on. Nothing gets indexed yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a collection owns
&lt;/h2&gt;

&lt;p&gt;A Qdrant collection is the container for points that you want to search together. Each point can hold one or more vector representations and an optional payload. This article uses one unnamed dense vector. Later, the payload can carry application data such as the source ID, text, language, or tenant ID.&lt;/p&gt;

&lt;p&gt;This first example will not write any points. It only creates the collection that will receive them later.&lt;/p&gt;

&lt;p&gt;For one unnamed dense vector, the collection schema needs two decisions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;vector size&lt;/li&gt;
&lt;li&gt;distance function&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Both come from the embedding contract. The example uses 1,536 dimensions and cosine distance because that matches the default &lt;code&gt;text-embedding-3-small&lt;/code&gt; profile from the previous article. If your model returns 768 or 3,072 values, use that number instead. If the model documentation recommends a different distance function, follow it.&lt;/p&gt;

&lt;p&gt;Matching the vector length is necessary, but it does not prove that two vectors are compatible. Qdrant can reject a vector with the wrong length. It cannot tell whether a correctly sized vector came from the wrong embedding model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run Qdrant locally
&lt;/h2&gt;

&lt;p&gt;The example uses Qdrant 1.19.0 in Docker. Create a named volume so the collection remains available after the container stops:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker volume create qdrant_storage

docker run &lt;span class="nt"&gt;--name&lt;/span&gt; qdrant-local &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-p&lt;/span&gt; 127.0.0.1:6333:6333 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-p&lt;/span&gt; 127.0.0.1:6334:6334 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-v&lt;/span&gt; qdrant_storage:/qdrant/storage &lt;span class="se"&gt;\&lt;/span&gt;
  qdrant/qdrant:v1.19.0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Qdrant exposes its HTTP API and dashboard on port &lt;code&gt;6333&lt;/code&gt;. The gRPC API listens on &lt;code&gt;6334&lt;/code&gt;, which is the port used by the .NET client below. Both ports are bound to &lt;code&gt;127.0.0.1&lt;/code&gt;, so they are available only from the Docker host. Without that address, Docker publishes ports on every host interface by default.&lt;/p&gt;

&lt;p&gt;Leave that terminal open and use a second one for the .NET commands. Open &lt;a href="http://127.0.0.1:6333/dashboard" rel="noopener noreferrer"&gt;http://127.0.0.1:6333/dashboard&lt;/a&gt; to confirm that the server is available. With a newly created, empty volume, the Collections page should initially be empty. &lt;code&gt;docker volume create&lt;/code&gt; reuses an existing volume with the same name, so old collections can reappear.&lt;/p&gt;

&lt;p&gt;When you stop the container, start it again with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker start &lt;span class="nt"&gt;-a&lt;/span&gt; qdrant-local
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This setup is for local development. It has no API key, TLS configuration, backups, or production-grade network controls. Do not expose it to an untrusted network.&lt;/p&gt;

&lt;h2&gt;
  
  
  Create the .NET project
&lt;/h2&gt;

&lt;p&gt;Create a console application and add the official Qdrant client:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dotnet new console &lt;span class="nt"&gt;--framework&lt;/span&gt; net10.0 &lt;span class="nt"&gt;--name&lt;/span&gt; QdrantCollections
&lt;span class="nb"&gt;cd &lt;/span&gt;QdrantCollections
dotnet add package Qdrant.Client &lt;span class="nt"&gt;--version&lt;/span&gt; 1.19.0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I compiled the following sample with .NET 10 and &lt;code&gt;Qdrant.Client&lt;/code&gt; 1.19.0. The client uses gRPC for its high-level API, so it connects to port &lt;code&gt;6334&lt;/code&gt; rather than the dashboard port.&lt;/p&gt;

&lt;p&gt;Replace &lt;code&gt;Program.cs&lt;/code&gt; with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Qdrant.Client&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Qdrant.Client.Grpc&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;collectionName&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"support-documents-v1"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;embeddingProfileId&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"support-text-v1"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;ulong&lt;/span&gt; &lt;span class="n"&gt;vectorSize&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="m"&gt;1536&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;var&lt;/span&gt; &lt;span class="n"&gt;executionTimeout&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt;
    &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;CancellationTokenSource&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;TimeSpan&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;FromSeconds&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="m"&gt;10&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;var&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;QdrantClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;host&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"127.0.0.1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;port&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="m"&gt;6334&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(!&lt;/span&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;CollectionExistsAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;collectionName&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;executionTimeout&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Token&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;CreateCollectionAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;collectionName&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;collectionName&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;vectorsConfig&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="n"&gt;VectorParams&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="n"&gt;Size&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;vectorSize&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;Distance&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Distance&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Cosine&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="n"&gt;metadata&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"embedding-profile-id"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;embeddingProfileId&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;executionTimeout&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Token&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="n"&gt;CollectionInfo&lt;/span&gt; &lt;span class="n"&gt;collection&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GetCollectionInfoAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;collectionName&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;executionTimeout&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Token&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="n"&gt;VectorsConfig&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="n"&gt;vectorsConfig&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;collection&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Params&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;VectorsConfig&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;vectorsConfig&lt;/span&gt; &lt;span class="k"&gt;is&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt; &lt;span class="p"&gt;||&lt;/span&gt;
    &lt;span class="n"&gt;vectorsConfig&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ConfigCase&lt;/span&gt; &lt;span class="p"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;VectorsConfig&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ConfigOneofCase&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Params&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;InvalidOperationException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="s"&gt;$"Collection '&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;collectionName&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;' does not use an unnamed "&lt;/span&gt; &lt;span class="p"&gt;+&lt;/span&gt;
        &lt;span class="s"&gt;"dense-vector configuration."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="n"&gt;VectorParams&lt;/span&gt; &lt;span class="n"&gt;vectors&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;vectorsConfig&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Params&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;vectors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Size&lt;/span&gt; &lt;span class="p"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;vectorSize&lt;/span&gt; &lt;span class="p"&gt;||&lt;/span&gt; &lt;span class="n"&gt;vectors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Distance&lt;/span&gt; &lt;span class="p"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;Distance&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Cosine&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;InvalidOperationException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="s"&gt;$"Collection '&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;collectionName&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;' does not match the expected "&lt;/span&gt; &lt;span class="p"&gt;+&lt;/span&gt;
        &lt;span class="s"&gt;$"&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;vectorSize&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;-dimension cosine configuration."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(!&lt;/span&gt;&lt;span class="n"&gt;collection&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Metadata&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;TryGetValue&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="s"&gt;"embedding-profile-id"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="k"&gt;out&lt;/span&gt; &lt;span class="n"&gt;Value&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="n"&gt;storedProfile&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;||&lt;/span&gt;
    &lt;span class="n"&gt;storedProfile&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;StringValue&lt;/span&gt; &lt;span class="p"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;embeddingProfileId&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;InvalidOperationException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="s"&gt;$"Collection '&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;collectionName&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;' does not declare the expected "&lt;/span&gt; &lt;span class="p"&gt;+&lt;/span&gt;
        &lt;span class="s"&gt;$"embedding profile '&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;embeddingProfileId&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;'."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="n"&gt;Console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;WriteLine&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="s"&gt;$"Collection '&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;collectionName&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;' is &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;collection&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Status&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt; with "&lt;/span&gt; &lt;span class="p"&gt;+&lt;/span&gt;
    &lt;span class="s"&gt;$"&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;vectors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Size&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt; dimensions, &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;vectors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Distance&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt; distance, and "&lt;/span&gt; &lt;span class="p"&gt;+&lt;/span&gt;
    &lt;span class="s"&gt;$"embedding profile '&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;embeddingProfileId&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;'."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run it while the Qdrant container is available:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dotnet run
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The output should look similar to this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Collection 'support-documents-v1' is Green with 1536 dimensions, Cosine distance, and embedding profile 'support-text-v1'.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Refresh the Qdrant dashboard and the collection should appear there too.&lt;/p&gt;

&lt;h2&gt;
  
  
  Creating once is not enough
&lt;/h2&gt;

&lt;p&gt;The existence check makes the sample convenient to run more than once. It does not assume that an existing collection is correct. This shorter version would be unsafe:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;CollectionExistsAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;collectionName&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;An older collection might have been created with 768 dimensions, dot-product distance, or a different embedding-profile ID. Returning early would allow the application to start against the wrong collection contract.&lt;/p&gt;

&lt;p&gt;The sample reads the stored configuration back. It checks that the dense-vector configuration is unnamed and uses the expected size and distance. It also checks the application-defined embedding-profile ID stored in the collection metadata.&lt;/p&gt;

&lt;p&gt;Qdrant does not infer which embedding model or preprocessing rules produced a vector. Collection metadata can hold an application-defined profile ID, but the application still owns what that ID means. Here, &lt;code&gt;support-text-v1&lt;/code&gt; refers to OpenAI &lt;code&gt;text-embedding-3-small&lt;/code&gt;, 1,536 dimensions, cosine distance, and &lt;code&gt;plain-text-v1&lt;/code&gt; preprocessing.&lt;/p&gt;

&lt;p&gt;The validation remains deliberately narrow. It does not reject additional sparse vectors, a multivector configuration, a non-default vector datatype, or every storage setting Qdrant supports. Add checks for those properties when application behavior depends on them. The current checks cover the unnamed dense-vector variant, size, distance, and profile ID.&lt;/p&gt;

&lt;p&gt;Two application instances can also observe a missing collection and race to create it. That is acceptable for this local console app. For a deployed system, I would provision collections through a controlled deployment or migration step instead of letting every application instance own schema creation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep the first collection boring
&lt;/h2&gt;

&lt;p&gt;Qdrant exposes settings for shards, replication, HNSW, quantization, on-disk storage, named vectors, sparse vectors, and more. None of them is needed to learn the collection boundary.&lt;/p&gt;

&lt;p&gt;The sample deliberately keeps Qdrant's defaults and creates:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;one collection&lt;/li&gt;
&lt;li&gt;one unnamed dense vector&lt;/li&gt;
&lt;li&gt;one vector size&lt;/li&gt;
&lt;li&gt;one distance function&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For this local exercise, the defaults are sufficient. In production, replication and sharding follow availability, scale, and isolation requirements. HNSW and quantization choices should come from measurements against a representative workload.&lt;/p&gt;

&lt;p&gt;The collection name includes &lt;code&gt;v1&lt;/code&gt; because its vectors belong to a particular embedding space. If a later profile uses an incompatible vector space, a new collection such as &lt;code&gt;support-documents-v2&lt;/code&gt; gives the migration an explicit target. It also avoids the temptation to delete and recreate a populated collection during application startup.&lt;/p&gt;

&lt;p&gt;Do not replace this with &lt;code&gt;RecreateCollectionAsync&lt;/code&gt; as a general startup shortcut. Recreating a collection deletes the existing collection and its points. That can be handy in a disposable test, but it is the wrong default once the data matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before indexing the first point
&lt;/h2&gt;

&lt;p&gt;The application has verified the dense vector's size and distance and the collection's embedding-profile ID. It has not decided how documents become stable point IDs, which payload fields to store, or where the original text lives. Those choices belong to the indexing path.&lt;/p&gt;

&lt;p&gt;Before adding that path, write down four answers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;support-text-v1&lt;/code&gt; owns this collection and identifies the embedding contract defined above.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;support-documents-v1&lt;/code&gt; is the collection name and migration target, not the embedding profile.&lt;/li&gt;
&lt;li&gt;In a deployed system, collection creation belongs to a controlled provisioning or migration step.&lt;/li&gt;
&lt;li&gt;Original text and payload design are intentionally undecided in this article.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Qdrant is a reasonable fit when a dedicated vector database suits the workload and the application can own another data boundary. If the source data already lives in PostgreSQL and the retrieval workload is modest, keeping vectors beside that data with pgvector may be simpler. &lt;em&gt;RAG with EF Core and pgvector&lt;/em&gt; covers that route.&lt;/p&gt;

&lt;p&gt;The next implementation step is to define stable point IDs and payloads, then write the first documents. HNSW and quantization can stay unchanged until retrieval measurements give a reason to revisit them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/posts/embeddings-in-dotnet/" rel="noopener noreferrer"&gt;Embeddings in .NET&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/posts/rag-efcore-pgvector/" rel="noopener noreferrer"&gt;RAG with EF Core and pgvector&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Qdrant: &lt;a href="https://qdrant.tech/documentation/quickstart/" rel="noopener noreferrer"&gt;Local quickstart&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Qdrant: &lt;a href="https://qdrant.tech/documentation/manage-data/collections/" rel="noopener noreferrer"&gt;Collections&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Qdrant: &lt;a href="https://qdrant.tech/documentation/manage-data/collections/#collection-metadata" rel="noopener noreferrer"&gt;Collection metadata&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Qdrant: &lt;a href="https://qdrant.tech/documentation/manage-data/vectors/" rel="noopener noreferrer"&gt;Vectors&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Docker: &lt;a href="https://docs.docker.com/engine/network/port-publishing/" rel="noopener noreferrer"&gt;Port publishing and mapping&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;NuGet: &lt;a href="https://www.nuget.org/packages/Qdrant.Client/1.19.0" rel="noopener noreferrer"&gt;&lt;code&gt;Qdrant.Client&lt;/code&gt; 1.19.0&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>dotnet</category>
      <category>vectordatabase</category>
      <category>rag</category>
      <category>csharp</category>
    </item>
    <item>
      <title>AI Systems Need Runtime Budgets</title>
      <dc:creator>Lukas Walter </dc:creator>
      <pubDate>Thu, 03 Sep 2026 15:30:00 +0000</pubDate>
      <link>https://dev.to/lukaswalter/ai-systems-need-runtime-budgets-43fc</link>
      <guid>https://dev.to/lukaswalter/ai-systems-need-runtime-budgets-43fc</guid>
      <description>&lt;p&gt;A five-second model timeout does not give you a five-second AI feature.&lt;/p&gt;

&lt;p&gt;Retrieval may take one second. The first model call takes four. A tool takes another two, and the next model call gets its own five-second timeout. Every dependency stayed within its local limit, yet the request took 12 seconds and paid for two model calls.&lt;/p&gt;

&lt;p&gt;The complete use case needs one runtime budget. Its remaining allowance follows every model call, tool invocation, and retry.&lt;/p&gt;

&lt;p&gt;Timeouts still matter. They are just one line in the budget.&lt;/p&gt;

&lt;h2&gt;
  
  
  Budget the execution, not each dependency
&lt;/h2&gt;

&lt;p&gt;Consider a support-answer workflow:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;load the ticket
    -&amp;gt; retrieve policy documents
    -&amp;gt; ask the model
    -&amp;gt; call a customer-history tool
    -&amp;gt; ask the model again
    -&amp;gt; validate the answer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each dependency can have a sensible timeout while the workflow still runs too long. The same applies to tokens and cost. Three calls with a 1,000-token output cap can generate 3,000 tokens. A retry repeats input, and tool results enlarge the next prompt.&lt;/p&gt;

&lt;p&gt;A runtime budget gives the workflow one shared envelope:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;What it bounds&lt;/th&gt;
&lt;th&gt;What happens when it is exhausted&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Wall-clock time&lt;/td&gt;
&lt;td&gt;Total useful lifetime of the execution&lt;/td&gt;
&lt;td&gt;Request cancellation and return a deadline result&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Input tokens&lt;/td&gt;
&lt;td&gt;Context sent across all model calls&lt;/td&gt;
&lt;td&gt;Reduce context before the call or stop&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output tokens&lt;/td&gt;
&lt;td&gt;Generated tokens across all model calls&lt;/td&gt;
&lt;td&gt;Lower the next call's cap or stop&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model calls&lt;/td&gt;
&lt;td&gt;Initial calls, follow-ups, repairs, and retries&lt;/td&gt;
&lt;td&gt;Do not start another call&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tool calls&lt;/td&gt;
&lt;td&gt;Automatic and application-directed invocations&lt;/td&gt;
&lt;td&gt;Stop the loop before another tool runs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Retries&lt;/td&gt;
&lt;td&gt;Repeated attempts after transient failures&lt;/td&gt;
&lt;td&gt;Return the chosen failure or recovery outcome&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Estimated cost&lt;/td&gt;
&lt;td&gt;Accumulated cost of metered work&lt;/td&gt;
&lt;td&gt;Stop or use a fallback whose cost already fits&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These limits belong together because the dimensions interact. A retry can spend more time, tokens, and money. A tool call spends time and may make the next prompt larger. A longer output consumes more of the deadline.&lt;/p&gt;

&lt;p&gt;A budget will stop some work that might have succeeded. Good. A late or overpriced result is not a successful execution for that product.&lt;/p&gt;

&lt;h2&gt;
  
  
  Define one execution policy
&lt;/h2&gt;

&lt;p&gt;Keep the configured limits immutable. The live ledger belongs to one execution, which may or may not be an HTTP request.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;sealed&lt;/span&gt; &lt;span class="k"&gt;record&lt;/span&gt; &lt;span class="nc"&gt;RuntimeBudget&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;TimeSpan&lt;/span&gt; &lt;span class="n"&gt;MaxDuration&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;TimeSpan&lt;/span&gt; &lt;span class="n"&gt;MaxModelAttemptDuration&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;TimeSpan&lt;/span&gt; &lt;span class="n"&gt;MinUsefulModelAttemptDuration&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;TimeSpan&lt;/span&gt; &lt;span class="n"&gt;CompletionHeadroom&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;MaxModelCalls&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;MaxToolCalls&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;MaxRetries&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="kt"&gt;long&lt;/span&gt; &lt;span class="n"&gt;MaxInputTokens&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="kt"&gt;long&lt;/span&gt; &lt;span class="n"&gt;MaxOutputTokens&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="kt"&gt;decimal&lt;/span&gt; &lt;span class="n"&gt;MaxEstimatedCostUsd&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Validate the policy when configuration is loaded. The three attempt and execution durations must be positive. Headroom, counters, token limits, and cost limits cannot be negative. Require a positive output cap. The time limits must satisfy both invariants:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;MinUsefulModelAttemptDuration &amp;lt;= MaxModelAttemptDuration
MinUsefulModelAttemptDuration + CompletionHeadroom &amp;lt;= MaxDuration
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Apply the same checks to reservations. Negative estimates must never create allowance. Use checked arithmetic or subtraction helpers to avoid overflow.&lt;/p&gt;

&lt;p&gt;Keep counters private and protect every compound read or mutation with the same lock. Return immutable snapshots created under that lock. Making one reservation method atomic does not make unrelated public properties thread-safe.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start the deadline and ledger together
&lt;/h2&gt;

&lt;p&gt;I use four separate cancellation sources:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;caller cancellation&lt;/li&gt;
&lt;li&gt;host shutdown&lt;/li&gt;
&lt;li&gt;the execution deadline&lt;/li&gt;
&lt;li&gt;a per-attempt timeout&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The execution scope owns its deadline. Both elapsed-time accounting and cancellation use the same &lt;code&gt;TimeProvider&lt;/code&gt; and start when the scope is created:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="kt"&gt;long&lt;/span&gt; &lt;span class="n"&gt;startedAt&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;timeProvider&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GetTimestamp&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;var&lt;/span&gt; &lt;span class="n"&gt;deadlineCts&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;CancellationTokenSource&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;limits&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;MaxDuration&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;timeProvider&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="n"&gt;TimeSpan&lt;/span&gt; &lt;span class="nf"&gt;RemainingTime&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;TimeSpan&lt;/span&gt; &lt;span class="n"&gt;remaining&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;limits&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;MaxDuration&lt;/span&gt;
        &lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="n"&gt;timeProvider&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GetElapsedTime&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;startedAt&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;remaining&lt;/span&gt; &lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;TimeSpan&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Zero&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="n"&gt;remaining&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;TimeSpan&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Zero&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Using one clock keeps tests honest when they use &lt;code&gt;FakeTimeProvider&lt;/code&gt;. A system timer should not expire independently from the fake clock used by the ledger.&lt;/p&gt;

&lt;p&gt;The per-attempt timeout must fit inside the time still available. If an execution has 1.4 seconds left, starting a model call with its normal five-second timeout is dishonest. Reserve completion headroom first, then calculate:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;attempt timeout = min(
    configured attempt timeout,
    remaining execution time - completion headroom)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Reject the call when the result is below &lt;code&gt;MinUsefulModelAttemptDuration&lt;/code&gt;. A one-millisecond attempt is traffic, not a serious chance of success.&lt;/p&gt;

&lt;p&gt;For each dependency call, create a time-provider-aware attempt source and link the sources only for delivery to that dependency:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;var&lt;/span&gt; &lt;span class="n"&gt;attemptCts&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;CancellationTokenSource&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;reservation&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;AttemptTimeout&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;timeProvider&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;var&lt;/span&gt; &lt;span class="n"&gt;callCts&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;CancellationTokenSource&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;CreateLinkedTokenSource&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;callerToken&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;hostStoppingToken&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;deadlineCts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Token&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;attemptCts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Token&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep the original tokens. The linked token asks observing work to stop, but it does not record which source won. In this example, classification precedence is host shutdown, caller cancellation, execution deadline, then attempt timeout. Pick an order deliberately and use it everywhere.&lt;/p&gt;

&lt;p&gt;Cancellation is cooperative. It does not prove that local work stopped immediately, or that a remote provider stopped processing and billing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reserve before handoff and reconcile afterward
&lt;/h2&gt;

&lt;p&gt;Post-call accounting is too late. The workflow has already spent the money and time.&lt;/p&gt;

&lt;p&gt;A model reservation needs a stable ID, estimated input tokens, an output cap, estimated cost, and an attempt timeout. Create it under the ledger lock only when every required dimension has enough allowance.&lt;/p&gt;

&lt;p&gt;The reservation lifecycle matters:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Reserved -&amp;gt; Started -&amp;gt; Reconciled
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;Release &lt;code&gt;Reserved&lt;/code&gt; work only when the application never handed it to the inner client or tool.&lt;/li&gt;
&lt;li&gt;Reconcile &lt;code&gt;Started&lt;/code&gt; work when required usage fields are available.&lt;/li&gt;
&lt;li&gt;When a started call returns uncertain or partial usage, charge the unresolved fields against the reservation or stop with an unknown-usage outcome.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;At this abstraction level, &lt;code&gt;Started&lt;/code&gt; means control was handed to the next layer. Local validation, serialization, or option conversion may still fail before a network request. Precise dispatch tracking requires instrumentation inside the provider or tool adapter. Without it, charge pessimistically after handoff. A caller timeout does not prove that the provider consumed nothing.&lt;/p&gt;

&lt;p&gt;The reservation operation should validate its arguments and calculate the attempt timeout inside the same lock that reserves calls, tokens, and cost. The core calculation is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;available attempt time = remaining execution time - completion headroom
attempt timeout = min(configured attempt timeout, available attempt time)
output allowance = limit - committed output - reserved output
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Reject the reservation when the attempt timeout is below the minimum or the output allowance is zero. Subtract validated non-negative values and return zero once committed or reserved usage reaches the limit. Use the same approach for input and cost.&lt;/p&gt;

&lt;p&gt;The reservation supplies the per-call &lt;code&gt;ChatOptions.MaxOutputTokens&lt;/code&gt;. That option limits one model request when the underlying client honors it. It is not a cumulative workflow limit.&lt;/p&gt;

&lt;h2&gt;
  
  
  Put enforcement inside automatic tool loops
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;FunctionInvokingChatClient.GetResponseAsync(...)&lt;/code&gt; can make several requests to its inner client. It sends tool results back to that inner client and continues until the tool loop ends. The final response aggregates available usage from those turns.&lt;/p&gt;

&lt;p&gt;One reservation around the outer call is wrong. It counts one model call even when the wrapper makes three, and its &lt;code&gt;MaxOutputTokens&lt;/code&gt; value can be applied to each inner request. It also sees none of the locally invoked tools.&lt;/p&gt;

&lt;p&gt;Place model-call enforcement inside the function-invocation wrapper. Put any retry middleware you control outside that budgeting boundary so each attempt must reserve again:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;application
    -&amp;gt; FunctionInvokingChatClient
        -&amp;gt; application retry layer
            -&amp;gt; BudgetingChatClient
                -&amp;gt; provider SDK
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is a lifetime trap here. &lt;code&gt;AddChatClient&lt;/code&gt; registers its pipeline as a singleton by default. A budgeting adapter that captures one execution's ledger or cancellation sources must not live in that singleton pipeline. Create the budgeting layer and its &lt;code&gt;FunctionInvokingChatClient&lt;/code&gt; for each execution, or keep the adapter stateless and pass the current ledger through explicit per-call context.&lt;/p&gt;

&lt;p&gt;The following fragment is illustrative pseudocode. Types such as &lt;code&gt;PromptEstimate&lt;/code&gt;, &lt;code&gt;ModelReservation&lt;/code&gt;, and the ledger are application contracts, not a companion library. &lt;code&gt;callCts.Token&lt;/code&gt; is the linked per-attempt token from the previous section.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="n"&gt;List&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;ChatMessage&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;messageList&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[..&lt;/span&gt; &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;
&lt;span class="n"&gt;ChatOptions&lt;/span&gt; &lt;span class="n"&gt;callOptions&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nf"&gt;Clone&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;??&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;ChatOptions&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="n"&gt;PromptEstimate&lt;/span&gt; &lt;span class="n"&gt;estimate&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;promptEstimator&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Estimate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;messageList&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;callOptions&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;modelProfile&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="n"&gt;ModelReservation&lt;/span&gt; &lt;span class="n"&gt;reservation&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;budget&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ReserveModelCall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;estimate&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;callOptions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;MaxOutputTokens&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="n"&gt;callOptions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;MaxOutputTokens&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;reservation&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;OutputTokenCap&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="n"&gt;budget&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;MarkStarted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;reservation&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// Controlled handoff, not confirmed dispatch.&lt;/span&gt;

&lt;span class="n"&gt;ChatResponse&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;innerClient&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GetResponseAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;messageList&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;callOptions&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;callCts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Token&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="n"&gt;budget&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ReconcileModel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;reservation&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Usage&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Materializing &lt;code&gt;messages&lt;/code&gt; once avoids enumerating an arbitrary &lt;code&gt;IEnumerable&amp;lt;ChatMessage&amp;gt;&lt;/code&gt; twice. The estimator needs more than message text. Depending on the provider, the effective prompt can include &lt;code&gt;ChatOptions.Instructions&lt;/code&gt;, tool declarations and JSON schemas, structured-output schemas, multimodal content, and provider-specific framing. &lt;code&gt;Estimate&lt;/code&gt; is the honest name because local preflight cannot always produce an exact count.&lt;/p&gt;

&lt;p&gt;Stateful clients make that limit clearer. If &lt;code&gt;ConversationId&lt;/code&gt; refers to history stored by the provider, the application may not possess the complete context. Reserve conservatively or disable local claims of exact input accounting for that path.&lt;/p&gt;

&lt;p&gt;After the controlled handoff, charge the reservation on an exception unless reliable usage data allows reconciliation. The fragment omits that branch and the streaming override. A production adapter needs both.&lt;/p&gt;

&lt;p&gt;Guard locally executed tools through &lt;code&gt;FunctionInvokingChatClient.FunctionInvoker&lt;/code&gt;. Validate known preconditions before reserving, then mark the reservation &lt;code&gt;Started&lt;/code&gt; immediately before &lt;code&gt;context.Function.InvokeAsync(context.Arguments, cancellationToken)&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;On success or a known failure, reconcile from the request, response metadata, usage record, and failure details available to that tool adapter. If those sources cannot resolve the cost after handoff, charge the reservation. This is deliberately pessimistic. Parallel tool calls need atomic reservations before any task starts.&lt;/p&gt;

&lt;p&gt;Hard tool-call limits apply only to execution you control. Provider-managed or server-side tools can be opaque. In that case, enforce the controls the provider exposes, account for reported usage, and avoid claiming that the application counted every tool call.&lt;/p&gt;

&lt;p&gt;An application-owned tool loop is the other valid design. It is more code, but it makes every model and tool boundary visible without relying on wrapper placement.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reconcile usage field by field
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;ChatResponse.Usage&lt;/code&gt; being non-null does not mean every count is present. &lt;code&gt;InputTokenCount&lt;/code&gt;, &lt;code&gt;OutputTokenCount&lt;/code&gt;, cached input, and reasoning counts are nullable. Reconcile each field required by your policy. For a missing required field, charge its reservation or stop with unknown usage.&lt;/p&gt;

&lt;p&gt;Per-call interception avoids another ambiguity. &lt;code&gt;FunctionInvokingChatClient&lt;/code&gt; adds the non-null usage objects returned by its inner calls. If one turn omits usage and another reports it, the outer response can contain a plausible but incomplete total.&lt;/p&gt;

&lt;p&gt;Do not add cached input tokens to &lt;code&gt;InputTokenCount&lt;/code&gt; again. Do not add reasoning tokens to &lt;code&gt;OutputTokenCount&lt;/code&gt; again. Those detail counts classify pricing; the documented totals already include them.&lt;/p&gt;

&lt;p&gt;Cost estimation needs a reviewed pricing key, not only a model ID. Depending on the provider, it may include the provider, deployment or model snapshot, region, service tier, modality, batch mode, and cached-token rules.&lt;/p&gt;

&lt;p&gt;An application-estimated cost budget is not a provider billing ceiling. Rounding, missing usage, and price changes can make the invoice differ. Use provider-side quotas and billing controls separately, and know whether they alert or block spend.&lt;/p&gt;

&lt;p&gt;The total cost dimension must include every paid boundary the article claims to budget. Retrieval services, rerankers, paid tools, and external APIs need their own reservations. If the ledger only accounts for model calls, name the limit &lt;code&gt;MaxEstimatedModelCostUsd&lt;/code&gt; instead.&lt;/p&gt;

&lt;h2&gt;
  
  
  Retries spend the original execution budget
&lt;/h2&gt;

&lt;p&gt;A retry does not receive a fresh deadline or token allowance. Do not start one that cannot finish inside the remaining time or would exceed another dimension.&lt;/p&gt;

&lt;p&gt;Put application retry middleware outside &lt;code&gt;BudgetingChatClient&lt;/code&gt;, as shown above. Each retry attempt then crosses the budgeting boundary and needs a new reservation.&lt;/p&gt;

&lt;p&gt;Before repeating an attempt, the retry layer must atomically reserve one retry slot. The inner budgeting client separately reserves the repeated model or dependency call.&lt;/p&gt;

&lt;p&gt;Retries inside the provider SDK are different. Configure or disable them when possible. If the SDK retries internally without an attempt callback, the application cannot claim an exact retry or model-call count. Treat that layer as opaque and reserve conservatively.&lt;/p&gt;

&lt;p&gt;A cost-exhausted execution may use a fallback only when that fallback is free, has separately reserved capacity, or still fits inside the remaining allowance.&lt;/p&gt;

&lt;h2&gt;
  
  
  Budget exhaustion is an application result
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;OperationCanceledException&lt;/code&gt; alone does not explain what happened. It might mean the user disconnected, the host stopped, a dependency attempt timed out, or the complete runtime deadline expired.&lt;/p&gt;

&lt;p&gt;Select a stop reason from the original source tokens, not from the linked token. Apply the same precedence at every boundary. This is a policy decision based on the tokens observed at inspection time; it does not prove which token caused the cancellation first.&lt;/p&gt;

&lt;p&gt;For example, a &lt;code&gt;SelectStopReason&lt;/code&gt; helper can check host shutdown first, then caller cancellation, the execution deadline, and finally the attempt timer. Map those to unavailable, cancelled, time-budget exhaustion, and dependency timeout results. Ledger failures should carry the exhausted dimension directly rather than masquerading as cancellation.&lt;/p&gt;

&lt;p&gt;The user-facing message need not mention tokens or internal prices. Keep the reason in the application result for telemetry, support, and fallback decisions.&lt;/p&gt;

&lt;p&gt;A partial answer needs its own policy. Some features can label and return one. Others should return nothing because an incomplete answer would be misleading.&lt;/p&gt;

&lt;h2&gt;
  
  
  Observe the allowance and the spend
&lt;/h2&gt;

&lt;p&gt;Record both sides of the ledger. A trace that says &lt;code&gt;model.duration = 3.2s&lt;/code&gt; does not say whether 3.2 seconds was healthy for this feature.&lt;/p&gt;

&lt;p&gt;For one execution, I would record:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the budget policy or a stable policy version&lt;/li&gt;
&lt;li&gt;elapsed and remaining time at each boundary&lt;/li&gt;
&lt;li&gt;model calls, tool calls, and retries used&lt;/li&gt;
&lt;li&gt;reported input and output tokens&lt;/li&gt;
&lt;li&gt;estimated cost and the price configuration version&lt;/li&gt;
&lt;li&gt;the exhausted dimension, when execution stops&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Do not copy prompts, retrieved documents, or tool results into telemetry just to explain the numbers. Operational metadata usually answers the budget question without duplicating sensitive content.&lt;/p&gt;

&lt;p&gt;Per-execution budgets stop one workflow from running away. They do nothing about ten thousand executions that each stay inside budget. Rate limits, concurrency limits, tenant quotas, and dependency-level retry budgets protect that shared capacity.&lt;/p&gt;

&lt;h2&gt;
  
  
  Choose numbers from the product promise
&lt;/h2&gt;

&lt;p&gt;There is no universal runtime budget. A chat interaction and an overnight evaluation run have different definitions of "too late." Start with how long the caller will wait, what one result may cost, which tools it needs, and whether a fallback still has value.&lt;/p&gt;

&lt;p&gt;Test the complete path under normal load and injected failure. Measure tail latency as well as averages. Include prompt growth, tool output, throttling, and retry delays. Test the effects of opaque SDK retries even when the SDK does not expose an exact attempt count.&lt;/p&gt;

&lt;p&gt;Leave headroom between the internal deadline and the external timeout. The application still needs to map the result, write telemetry, and respond. If both expire together, the caller may receive a connection failure instead of the intended outcome.&lt;/p&gt;

&lt;p&gt;Treat the first values as a hypothesis. Production traces and cost data will show where the budget is too loose, too strict, or spent on the wrong step.&lt;/p&gt;

&lt;h2&gt;
  
  
  When runtime budgets are useful
&lt;/h2&gt;

&lt;p&gt;Use an execution-scoped budget when a workflow can make more than one metered or remote call, invoke tools, retry, or fan out. It gives the orchestration layer one answer to the question: may this execution start more work?&lt;/p&gt;

&lt;p&gt;For interactive work, that execution normally observes caller cancellation. Deliberate background work needs an independent budget instead of silently outliving the request. Durable workflows also need persisted reservations and idempotent reconciliation so a restart cannot reset or double-charge the ledger.&lt;/p&gt;

&lt;h2&gt;
  
  
  When a full budget object is unnecessary
&lt;/h2&gt;

&lt;p&gt;A local spike or one low-risk model call may only need cancellation, a timeout, and an output cap. Add a shared ledger when the workflow can multiply time or spend, not when the object would add ceremony without changing a decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practical takeaway
&lt;/h2&gt;

&lt;p&gt;Pick one AI use case and write down the maximum wall-clock time, cumulative tokens, model calls, tool calls, retries, and estimated cost for one execution.&lt;/p&gt;

&lt;p&gt;Put those limits in one policy. Create one ledger when the execution begins. Before every new attempt, reserve what it can spend. After each metered operation, reconcile its reservation with the available usage or cost data. When a required dimension lacks enough allowance, return a named application outcome instead of starting one more call.&lt;/p&gt;

&lt;p&gt;The next model call now starts only when the complete execution can still afford it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/posts/retries-are-not-a-recovery-strategy/" rel="noopener noreferrer"&gt;Retries Are Not a Recovery Strategy&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/posts/cancellation-token/" rel="noopener noreferrer"&gt;Why CancellationToken Matters More in .NET AI Systems&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/tips/cap-tool-call-loops-explicitly/" rel="noopener noreferrer"&gt;Cap tool-call loops explicitly&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.ai.chatoptions.maxoutputtokens?view=net-10.0-pp" rel="noopener noreferrer"&gt;&lt;code&gt;ChatOptions.MaxOutputTokens&lt;/code&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;.NET source: &lt;a href="https://github.com/dotnet/extensions/blob/main/src/Libraries/Microsoft.Extensions.AI.Abstractions/ChatCompletion/ChatOptions.cs" rel="noopener noreferrer"&gt;&lt;code&gt;ChatOptions.Clone()&lt;/code&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.ai.chatresponse.usage?view=net-10.0-pp" rel="noopener noreferrer"&gt;&lt;code&gt;ChatResponse.Usage&lt;/code&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.dependencyinjection.chatclientbuilderservicecollectionextensions.addchatclient?view=net-11.0-pp" rel="noopener noreferrer"&gt;&lt;code&gt;AddChatClient&lt;/code&gt; service lifetime&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.ai.functioninvokingchatclient?view=net-11.0-pp" rel="noopener noreferrer"&gt;&lt;code&gt;FunctionInvokingChatClient&lt;/code&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;.NET source: &lt;a href="https://github.com/dotnet/extensions/blob/main/src/Libraries/Microsoft.Extensions.AI/ChatCompletion/FunctionInvokingChatClient.cs" rel="noopener noreferrer"&gt;&lt;code&gt;FunctionInvokingChatClient.cs&lt;/code&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.ai.functioninvokingchatclient.functioninvoker?view=net-10.0-pp" rel="noopener noreferrer"&gt;&lt;code&gt;FunctionInvokingChatClient.FunctionInvoker&lt;/code&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.ai.usagedetails?view=net-11.0-pp" rel="noopener noreferrer"&gt;&lt;code&gt;UsageDetails&lt;/code&gt;&lt;/a&gt;, including &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.ai.usagedetails.cachedinputtokencount?view=net-11.0-pp" rel="noopener noreferrer"&gt;&lt;code&gt;CachedInputTokenCount&lt;/code&gt;&lt;/a&gt; and &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.ai.usagedetails.reasoningtokencount?view=net-11.0-pp" rel="noopener noreferrer"&gt;&lt;code&gt;ReasoningTokenCount&lt;/code&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/standard/datetime/timeprovider-overview" rel="noopener noreferrer"&gt;TimeProvider overview&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/standard/threading/cancellation-in-managed-threads" rel="noopener noreferrer"&gt;Cancellation in managed threads&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/azure/architecture/best-practices/transient-faults" rel="noopener noreferrer"&gt;Best practices for transient fault handling&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>dotnet</category>
      <category>architecture</category>
      <category>ai</category>
    </item>
    <item>
      <title>Provider Independence from Day One</title>
      <dc:creator>Lukas Walter </dc:creator>
      <pubDate>Mon, 31 Aug 2026 15:30:00 +0000</pubDate>
      <link>https://dev.to/lukaswalter/provider-independence-from-day-one-3fo3</link>
      <guid>https://dev.to/lukaswalter/provider-independence-from-day-one-3fo3</guid>
      <description>&lt;p&gt;Provider independence starts before the second provider arrives. Draw the provider boundary when the first integration is built. Alternate implementations can wait until a real need appears.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;IChatClient&lt;/code&gt; gives .NET applications a common API for messages, responses, streaming, tools, and structured output. That removes a lot of SDK coupling. The remaining differences still need a place in the design. JSON Schema support, tool behavior, context limits, model names, deployment names, and provider-only options do not disappear because two clients implement the same interface.&lt;/p&gt;

&lt;p&gt;Provider wiring stays at the composition boundary. A dedicated factory selects and constructs the Azure OpenAI or Ollama client. The ticket-routing feature uses one C# service for either backend. A small capability declaration selects native JSON Schema or prompt-guided schema instructions. A qualification check confirms that the configured structured-output path works before the backend is promoted.&lt;/p&gt;

&lt;p&gt;That is the level of independence I trust: one application path with the differences exposed and backed by tests against each configured backend.&lt;/p&gt;

&lt;p&gt;There is a complete example available as a &lt;a href="https://gist.github.com/ovnecron/3499a3baac48ba12923d58095155d78e" rel="noopener noreferrer"&gt;.NET 10 file-based app&lt;/a&gt;. It runs offline by default and includes configuration examples for Ollama and Azure OpenAI.&lt;/p&gt;

&lt;h2&gt;
  
  
  Define the portable behavior first
&lt;/h2&gt;

&lt;p&gt;Start with the feature, not a list of providers.&lt;/p&gt;

&lt;p&gt;This example routes a support ticket to one of three queues. The model call needs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;system and user messages&lt;/li&gt;
&lt;li&gt;a typed JSON result&lt;/li&gt;
&lt;li&gt;cancellation&lt;/li&gt;
&lt;li&gt;no tools, images, or provider-managed conversation state&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those requirements fit &lt;code&gt;IChatClient&lt;/code&gt;. The application result does not need to expose chat messages or SDK response types:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;enum&lt;/span&gt; &lt;span class="n"&gt;TicketQueue&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;Billing&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;TechnicalSupport&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;AccountAccess&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;sealed&lt;/span&gt; &lt;span class="k"&gt;record&lt;/span&gt; &lt;span class="nc"&gt;TicketRoutingDecision&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;TicketQueue&lt;/span&gt; &lt;span class="n"&gt;Queue&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="kt"&gt;bool&lt;/span&gt; &lt;span class="n"&gt;RequiresHumanReview&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;Reason&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the portability contract for the first version. If the feature later requires hosted file search or another provider-only API, the contract has changed. Pretending otherwise would only move the coupling into a loosely typed options dictionary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Install the provider integrations
&lt;/h2&gt;

&lt;p&gt;The sample uses Azure OpenAI in a hosted environment and Ollama for a local backend:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dotnet add package Microsoft.Extensions.AI
dotnet add package Microsoft.Extensions.AI.OpenAI
dotnet add package Azure.AI.OpenAI
dotnet add package Azure.Identity
dotnet add package OllamaSharp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I compiled the code in this article with .NET 10, &lt;code&gt;Microsoft.Extensions.AI&lt;/code&gt; 10.9.0, &lt;code&gt;Microsoft.Extensions.AI.OpenAI&lt;/code&gt; 10.9.0, &lt;code&gt;Azure.AI.OpenAI&lt;/code&gt; 2.1.0, &lt;code&gt;Azure.Identity&lt;/code&gt; 1.21.0, and &lt;code&gt;OllamaSharp&lt;/code&gt; 5.4.30. Check the current package documentation when using newer versions.&lt;/p&gt;

&lt;p&gt;The dependency direction stays small:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TicketRouter
    -&amp;gt; SupportModelClient
        -&amp;gt; IChatClient

composition root
    -&amp;gt; SupportModelClientFactory
        -&amp;gt; AzureOpenAIClient or OllamaApiClient
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;TicketRouter&lt;/code&gt; has no reference to either provider SDK.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep model and deployment configuration separate
&lt;/h2&gt;

&lt;p&gt;Azure OpenAI addresses a deployment. Ollama addresses a model available at its endpoint. The application should not use either value as its stable service name.&lt;/p&gt;

&lt;p&gt;Use one options type for the selected support-routing backend:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;System.ComponentModel.DataAnnotations&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;enum&lt;/span&gt; &lt;span class="n"&gt;StructuredOutputMode&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;Unspecified&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;NativeJsonSchema&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;PromptedJsonSchema&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;sealed&lt;/span&gt; &lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;SupportModelOptions&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;Required&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;Provider&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;get&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;init&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;""&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;Required&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;Model&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;get&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;init&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;""&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="n"&gt;Deployment&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;get&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;init&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;Required&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="n"&gt;Uri&lt;/span&gt; &lt;span class="n"&gt;Endpoint&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;get&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;init&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt;&lt;span class="p"&gt;!;&lt;/span&gt;

    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="n"&gt;StructuredOutputMode&lt;/span&gt; &lt;span class="n"&gt;StructuredOutput&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;get&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;init&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;An Azure OpenAI configuration can look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"AI"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"SupportRouting"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"Provider"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"AzureOpenAI"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"Endpoint"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://example.openai.azure.com/"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"Model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"your-model-id"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"Deployment"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"support-routing-prod"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"StructuredOutput"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"NativeJsonSchema"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The deployment is the Azure OpenAI deployment name used by the SDK. It identifies a model deployment inside the Azure OpenAI resource. &lt;code&gt;Model&lt;/code&gt; is the expected model identity the team records with traces, evaluations, and release decisions. They may have the same text, but they own different concerns.&lt;/p&gt;

&lt;p&gt;For local Ollama development, only the provider-specific settings change:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"AI"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"SupportRouting"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"Provider"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Ollama"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"Endpoint"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"http://localhost:11434/"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"Model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"your-qualified-model"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"StructuredOutput"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"PromptedJsonSchema"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;StructuredOutput&lt;/code&gt; describes a tested capability of the model, backend, and integration. It is not a permanent fact about the provider. Native JSON Schema may work well enough for one combination of model, backend, and integration but not another. Set &lt;code&gt;NativeJsonSchema&lt;/code&gt; only after checking the exact combination you will deploy. The probe later in this article checks that the path works, but it cannot prove schema enforcement. &lt;code&gt;Unspecified&lt;/code&gt; exists so missing configuration fails instead of silently selecting the stronger mode.&lt;/p&gt;

&lt;p&gt;Do not copy the placeholder model names and capability values into production. Qualify the model you intend to run.&lt;/p&gt;

&lt;h2&gt;
  
  
  Build one provider-neutral client holder
&lt;/h2&gt;

&lt;p&gt;The application needs the chat client and the capability declaration together. A small holder keeps them consistent. The DI container owns that holder and disposes the selected &lt;code&gt;IChatClient&lt;/code&gt; when the container shuts down:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Microsoft.Extensions.AI&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;sealed&lt;/span&gt; &lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;SupportModelClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;IChatClient&lt;/span&gt; &lt;span class="n"&gt;chatClient&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;declaredModelIdentity&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;StructuredOutputMode&lt;/span&gt; &lt;span class="n"&gt;structuredOutput&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;IDisposable&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="n"&gt;IChatClient&lt;/span&gt; &lt;span class="n"&gt;ChatClient&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;get&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;chatClient&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;DeclaredModelIdentity&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;get&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;declaredModelIdentity&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="n"&gt;StructuredOutputMode&lt;/span&gt; &lt;span class="n"&gt;StructuredOutput&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;get&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;structuredOutput&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;Dispose&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;ChatClient&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Dispose&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I would not call this class &lt;code&gt;AiClient&lt;/code&gt;. It is the selected client for one application purpose. A document summarizer or an agent may have a different model, capability set, and operating policy.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;DeclaredModelIdentity&lt;/code&gt; is configuration, not proof of what an Azure deployment currently serves. A deployment can move to another model or version while the application setting becomes stale. Record model metadata returned by the provider as an observed identity when it is available, and compare it with the declared value in telemetry or deployment checks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Create the provider clients at the edge
&lt;/h2&gt;

&lt;p&gt;The factory is allowed to know the provider SDKs. That is its job.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Azure.AI.OpenAI&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Azure.Identity&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Microsoft.Extensions.AI&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;OllamaSharp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;static&lt;/span&gt; &lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;SupportModelClientFactory&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;static&lt;/span&gt; &lt;span class="n"&gt;SupportModelClient&lt;/span&gt; &lt;span class="nf"&gt;Create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;SupportModelOptions&lt;/span&gt; &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;StructuredOutput&lt;/span&gt; &lt;span class="k"&gt;is&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;StructuredOutputMode&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;NativeJsonSchema&lt;/span&gt; &lt;span class="k"&gt;or&lt;/span&gt;
            &lt;span class="n"&gt;StructuredOutputMode&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;PromptedJsonSchema&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;InvalidOperationException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="s"&gt;"A supported structured-output mode is required."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;

        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Provider&lt;/span&gt; &lt;span class="k"&gt;switch&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="s"&gt;"AzureOpenAI"&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;CreateAzureOpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
            &lt;span class="s"&gt;"Ollama"&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;CreateOllama&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
            &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;InvalidOperationException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="s"&gt;$"Unsupported AI provider '&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Provider&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;'."&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;};&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;private&lt;/span&gt; &lt;span class="k"&gt;static&lt;/span&gt; &lt;span class="n"&gt;SupportModelClient&lt;/span&gt; &lt;span class="nf"&gt;CreateAzureOpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;SupportModelOptions&lt;/span&gt; &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;IsNullOrWhiteSpace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Deployment&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;InvalidOperationException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="s"&gt;"AI:SupportRouting:Deployment is required for Azure OpenAI."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;

        &lt;span class="kt"&gt;var&lt;/span&gt; &lt;span class="n"&gt;azureClient&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;AzureOpenAIClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Endpoint&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;DefaultAzureCredential&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt;

        &lt;span class="n"&gt;IChatClient&lt;/span&gt; &lt;span class="n"&gt;chatClient&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;azureClient&lt;/span&gt;
            &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GetChatClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Deployment&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;AsIChatClient&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;SupportModelClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;chatClient&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;StructuredOutput&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;private&lt;/span&gt; &lt;span class="k"&gt;static&lt;/span&gt; &lt;span class="n"&gt;SupportModelClient&lt;/span&gt; &lt;span class="nf"&gt;CreateOllama&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;SupportModelOptions&lt;/span&gt; &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;IChatClient&lt;/span&gt; &lt;span class="n"&gt;chatClient&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;OllamaApiClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Endpoint&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Model&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;SupportModelClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;chatClient&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;StructuredOutput&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Azure authentication belongs in &lt;code&gt;CreateAzureOpenAI&lt;/code&gt;. Ollama endpoint and model setup belong in &lt;code&gt;CreateOllama&lt;/code&gt;. &lt;code&gt;TicketRouter&lt;/code&gt; sees neither SDK.&lt;/p&gt;

&lt;p&gt;The switch runs once while the container creates its singleton. It does not appear in every request path, controller, or application service.&lt;/p&gt;

&lt;p&gt;Register the options, selected client, and feature service in &lt;code&gt;Program.cs&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Microsoft.Extensions.Options&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="n"&gt;builder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Services&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;AddOptions&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;SupportModelOptions&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;()&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;BindConfiguration&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"AI:SupportRouting"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ValidateDataAnnotations&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Validate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;options&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Provider&lt;/span&gt; &lt;span class="k"&gt;is&lt;/span&gt; &lt;span class="s"&gt;"AzureOpenAI"&lt;/span&gt; &lt;span class="k"&gt;or&lt;/span&gt; &lt;span class="s"&gt;"Ollama"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="s"&gt;"Provider must be AzureOpenAI or Ollama."&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Validate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;options&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;StructuredOutput&lt;/span&gt; &lt;span class="k"&gt;is&lt;/span&gt;
            &lt;span class="n"&gt;StructuredOutputMode&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;NativeJsonSchema&lt;/span&gt; &lt;span class="k"&gt;or&lt;/span&gt;
            &lt;span class="n"&gt;StructuredOutputMode&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;PromptedJsonSchema&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="s"&gt;"StructuredOutput must be declared explicitly."&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Validate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;options&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt;
            &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Provider&lt;/span&gt; &lt;span class="p"&gt;!=&lt;/span&gt; &lt;span class="s"&gt;"AzureOpenAI"&lt;/span&gt; &lt;span class="p"&gt;||&lt;/span&gt;
            &lt;span class="p"&gt;!&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;IsNullOrWhiteSpace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Deployment&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="s"&gt;"Deployment is required for Azure OpenAI."&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ValidateOnStart&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="n"&gt;builder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Services&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;AddSingleton&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;SupportModelClient&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;(&lt;/span&gt;&lt;span class="n"&gt;services&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;SupportModelOptions&lt;/span&gt; &lt;span class="n"&gt;options&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;services&lt;/span&gt;
        &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;GetRequiredService&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;IOptions&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;SupportModelOptions&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&amp;gt;()&lt;/span&gt;
        &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Value&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;SupportModelClientFactory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="n"&gt;builder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Services&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;AddTransient&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;TicketRouter&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The custom validators reject an unsupported provider, a missing capability declaration, and a missing Azure OpenAI deployment during startup. The factory keeps defensive checks for callers that bypass this registration path. A production application should also validate allowed endpoint schemes, the selected authentication mode, and unsupported option combinations.&lt;/p&gt;

&lt;h2&gt;
  
  
  Implement the feature once
&lt;/h2&gt;

&lt;p&gt;The model-facing DTO uses strings so the application can reject unknown queue values deliberately. Deserializing directly into an enum can make an unsupported value look like a serialization detail instead of an invalid model decision.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;internal&lt;/span&gt; &lt;span class="k"&gt;sealed&lt;/span&gt; &lt;span class="k"&gt;record&lt;/span&gt; &lt;span class="nc"&gt;TicketRoutingOutput&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="n"&gt;Queue&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="kt"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="n"&gt;RequiresHumanReview&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="n"&gt;Reason&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;RequiresHumanReview&lt;/code&gt; is nullable only at the model boundary. If the model omits the property, deserialization leaves it as &lt;code&gt;null&lt;/code&gt; and the application can reject the incomplete decision instead of treating the missing value as &lt;code&gt;false&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The router asks for a typed response, then validates every field that crosses the model boundary:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Microsoft.Extensions.AI&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;sealed&lt;/span&gt; &lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;TicketRouter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;SupportModelClient&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;private&lt;/span&gt; &lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;Instructions&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"""
&lt;/span&gt;        &lt;span class="n"&gt;Route&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;support&lt;/span&gt; &lt;span class="n"&gt;ticket&lt;/span&gt; &lt;span class="n"&gt;to&lt;/span&gt; &lt;span class="n"&gt;exactly&lt;/span&gt; &lt;span class="n"&gt;one&lt;/span&gt; &lt;span class="n"&gt;queue&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;Billing&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;TechnicalSupport&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;or&lt;/span&gt; &lt;span class="n"&gt;AccountAccess&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;
        &lt;span class="n"&gt;Set&lt;/span&gt; &lt;span class="n"&gt;requiresHumanReview&lt;/span&gt; &lt;span class="n"&gt;to&lt;/span&gt; &lt;span class="k"&gt;true&lt;/span&gt; &lt;span class="k"&gt;when&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;ticket&lt;/span&gt; &lt;span class="k"&gt;is&lt;/span&gt; &lt;span class="n"&gt;ambiguous&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;
        &lt;span class="n"&gt;Keep&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;reason&lt;/span&gt; &lt;span class="n"&gt;to&lt;/span&gt; &lt;span class="m"&gt;120&lt;/span&gt; &lt;span class="n"&gt;characters&lt;/span&gt; &lt;span class="k"&gt;or&lt;/span&gt; &lt;span class="n"&gt;fewer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;
        &lt;span class="n"&gt;Return&lt;/span&gt; &lt;span class="n"&gt;JSON&lt;/span&gt; &lt;span class="n"&gt;only&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;
        &lt;span class="s"&gt;""";
&lt;/span&gt;
    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="n"&gt;Task&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;TicketRoutingDecision&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;RouteAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;ticket&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;CancellationToken&lt;/span&gt; &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;IsNullOrWhiteSpace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ticket&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;ArgumentException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="s"&gt;"Ticket text is required."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="k"&gt;nameof&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ticket&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;

        &lt;span class="n"&gt;ChatResponse&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;TicketRoutingOutput&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt;
            &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ChatClient&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;GetResponseAsync&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;TicketRoutingOutput&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;(&lt;/span&gt;
                &lt;span class="p"&gt;[&lt;/span&gt;
                    &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;ChatMessage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ChatRole&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;System&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Instructions&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
                    &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;ChatMessage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ChatRole&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;User&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ticket&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="p"&gt;],&lt;/span&gt;
                &lt;span class="n"&gt;useJsonSchemaResponseFormat&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;StructuredOutput&lt;/span&gt; &lt;span class="k"&gt;is&lt;/span&gt;
                        &lt;span class="n"&gt;StructuredOutputMode&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;NativeJsonSchema&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(!&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;TryGetResult&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;out&lt;/span&gt; &lt;span class="n"&gt;TicketRoutingOutput&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;||&lt;/span&gt;
            &lt;span class="n"&gt;output&lt;/span&gt; &lt;span class="k"&gt;is&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt; &lt;span class="p"&gt;||&lt;/span&gt;
            &lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;RequiresHumanReview&lt;/span&gt; &lt;span class="k"&gt;is&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt; &lt;span class="p"&gt;||&lt;/span&gt;
            &lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;IsNullOrWhiteSpace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Reason&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;||&lt;/span&gt;
            &lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Reason&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Length&lt;/span&gt; &lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="m"&gt;120&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;InvalidOperationException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="s"&gt;"The model returned an invalid routing decision."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;

        &lt;span class="n"&gt;TicketQueue&lt;/span&gt; &lt;span class="n"&gt;queue&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Queue&lt;/span&gt; &lt;span class="k"&gt;switch&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;nameof&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;TicketQueue&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Billing&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;TicketQueue&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Billing&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="k"&gt;nameof&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;TicketQueue&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;TechnicalSupport&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt;
                &lt;span class="n"&gt;TicketQueue&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;TechnicalSupport&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="k"&gt;nameof&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;TicketQueue&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;AccountAccess&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;TicketQueue&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;AccountAccess&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;InvalidOperationException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="s"&gt;"The model returned an invalid queue."&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;};&lt;/span&gt;

        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;TicketRoutingDecision&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;queue&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;RequiresHumanReview&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Value&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Reason&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both modes use the same DTO and application validation. The generated schema describes the serialized DTO shape. The explicit checks still own domain rules such as required fields, the allowed queue names, and the maximum reason length.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;GetResponseAsync&amp;lt;T&amp;gt;&lt;/code&gt; derives a JSON Schema from &lt;code&gt;T&lt;/code&gt; in both modes. With &lt;code&gt;NativeJsonSchema&lt;/code&gt;, the helper sends that schema through the backend's structured-response mechanism. With &lt;code&gt;PromptedJsonSchema&lt;/code&gt;, it requests JSON and appends another user message containing the generated schema. The second mode guides the model through the prompt instead of relying on native schema enforcement. The application applies the same post-deserialization validation either way.&lt;/p&gt;

&lt;p&gt;An invalid result should become an explicit application outcome in a complete service layer. I use an exception here to keep the sample focused on provider composition. The production decision may be to return &lt;code&gt;InvalidGeneration&lt;/code&gt;, make one bounded repair attempt, or require review. It should not return the unvalidated text as a routing decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  Treat capabilities as claims that need evidence
&lt;/h2&gt;

&lt;p&gt;A capability flag can become another lie in configuration. Verify it against the exact model, deployment, integration package, and endpoint used by the environment.&lt;/p&gt;

&lt;p&gt;A small qualification probe can check the structured-output path:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Microsoft.Extensions.AI&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;internal&lt;/span&gt; &lt;span class="k"&gt;sealed&lt;/span&gt; &lt;span class="k"&gt;record&lt;/span&gt; &lt;span class="nc"&gt;CapabilityProbe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="n"&gt;Value&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;static&lt;/span&gt; &lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;SupportModelQualification&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;static&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="n"&gt;Task&lt;/span&gt; &lt;span class="nf"&gt;VerifyAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;SupportModelClient&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;CancellationToken&lt;/span&gt; &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;ChatResponse&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;CapabilityProbe&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt;
            &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ChatClient&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;GetResponseAsync&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;CapabilityProbe&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;(&lt;/span&gt;
                &lt;span class="s"&gt;"Return JSON with value set to ready."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="n"&gt;useJsonSchemaResponseFormat&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;StructuredOutput&lt;/span&gt; &lt;span class="k"&gt;is&lt;/span&gt;
                        &lt;span class="n"&gt;StructuredOutputMode&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;NativeJsonSchema&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(!&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;TryGetResult&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;out&lt;/span&gt; &lt;span class="n"&gt;CapabilityProbe&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;||&lt;/span&gt;
            &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="n"&gt;Value&lt;/span&gt; &lt;span class="p"&gt;!=&lt;/span&gt; &lt;span class="s"&gt;"ready"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;InvalidOperationException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="s"&gt;$"Model '&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;DeclaredModelIdentity&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;' failed the "&lt;/span&gt; &lt;span class="p"&gt;+&lt;/span&gt;
                &lt;span class="s"&gt;"support-routing structured-output check."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run this as a deployment smoke test or structured-output contract test. I would not make every application startup depend on a paid model call. The test should run whenever the model, deployment, provider integration, or relevant configuration changes.&lt;/p&gt;

&lt;p&gt;This probe proves that the request succeeds and the helper can deserialize a compatible result for this example. A model could return &lt;code&gt;{"value":"ready"}&lt;/code&gt; without obeying the supplied schema, so the probe does not prove native schema enforcement. It also says nothing about routing quality. Keep a small golden dataset for that question and evaluate every candidate backend against the same cases. Provider independence without comparable output quality is not useful independence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep provider-only behavior in a named adapter
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;ChatOptions&lt;/code&gt; covers common settings. It also exposes &lt;code&gt;AdditionalProperties&lt;/code&gt; and &lt;code&gt;RawRepresentationFactory&lt;/code&gt; for options understood by a specific provider. Those escape hatches are useful, but the type system cannot turn them into portable behavior.&lt;/p&gt;

&lt;p&gt;If &lt;code&gt;TicketRouter&lt;/code&gt; creates an OpenAI &lt;code&gt;ChatCompletionOptions&lt;/code&gt; through &lt;code&gt;RawRepresentationFactory&lt;/code&gt;, that call is OpenAI-specific even though the outer dependency is &lt;code&gt;IChatClient&lt;/code&gt;. Put it in an infrastructure adapter with a name that says what it requires, or accept that the feature is tied to that provider.&lt;/p&gt;

&lt;p&gt;The same rule applies to hosted file search, provider-managed conversation state, batch APIs, safety configuration, and model-specific reasoning controls. Do not flatten those features into an &lt;code&gt;AiOptions&lt;/code&gt; bag shared by the whole application.&lt;/p&gt;

&lt;p&gt;Sometimes the honest design has two paths:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;portable ticket routing
    -&amp;gt; IChatClient

provider-hosted file search
    -&amp;gt; purpose-specific application interface
        -&amp;gt; provider SDK adapter
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is still a well-contained architecture. Independence is about controlling the dependency, not eliminating every provider reference from the repository.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changes during a provider switch
&lt;/h2&gt;

&lt;p&gt;With this structure, moving the ticket router to another backend still requires engineering work. The work is contained:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Add the provider package and one factory branch.&lt;/li&gt;
&lt;li&gt;Configure endpoint, model identity, deployment identity when applicable, and the qualified capability mode.&lt;/li&gt;
&lt;li&gt;Run the structured-output contract check.&lt;/li&gt;
&lt;li&gt;Run the routing evaluation dataset and compare quality, latency, and cost.&lt;/li&gt;
&lt;li&gt;Review operational differences such as authentication, throttling, telemetry metadata, and content retention.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;code&gt;TicketRouter&lt;/code&gt; changes only if the feature contract changes. A new provider that cannot satisfy the existing contract is not a drop-in replacement. It may still be usable behind a deliberately reduced mode, but the application must name and test that behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  When this approach is enough
&lt;/h2&gt;

&lt;p&gt;Use this pattern when several backends can satisfy the same bounded operation through &lt;code&gt;IChatClient&lt;/code&gt;. Classification, extraction, summarization, and constrained generation are good candidates when their required inputs and outputs fit the shared abstraction.&lt;/p&gt;

&lt;p&gt;It also works well when production uses one provider and local development uses another. That setup catches accidental SDK leakage early, even if the production provider never changes.&lt;/p&gt;

&lt;h2&gt;
  
  
  When provider independence is the wrong goal
&lt;/h2&gt;

&lt;p&gt;Keep the provider SDK visible when the feature exists because of a provider-only capability, or when the abstraction would discard information the application needs. A thin, purpose-specific adapter is better than a generic interface full of string properties and raw objects.&lt;/p&gt;

&lt;p&gt;Do not pay a permanent complexity tax for a switch the application will never make. One provider behind a clean composition boundary is already a sensible design. Add alternate implementations and capability negotiation when an environment-specific requirement, resilience plan, procurement constraint, or local-development need justifies them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practical takeaway
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;IChatClient&lt;/code&gt; is useful only for behavior it can represent honestly. Keep provider construction at the edge and attach the evidence-based capability mode to the selected client.&lt;/p&gt;

&lt;p&gt;Provider-specific code will remain. Keep it easy to find and change, with its purpose stated in the adapter. The application feature should have no reason to care which SDK sits underneath.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/posts/stop-letting-provider-sdks-define-your-dotnet-ai-architecture/" rel="noopener noreferrer"&gt;Stop Letting Provider SDKs Define Your .NET AI Architecture&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/posts/designing-an-ai-service-layer/" rel="noopener noreferrer"&gt;Designing an AI Service Layer&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/tips/keep-model-name-and-deployment-name-in-config/" rel="noopener noreferrer"&gt;Keep model name and deployment name in config&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/posts/local-llms-in-.net/" rel="noopener noreferrer"&gt;Local LLMs in .NET with Ollama and Microsoft.Extensions.AI&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/ai/ichatclient" rel="noopener noreferrer"&gt;Use the &lt;code&gt;IChatClient&lt;/code&gt; interface&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.ai.ichatclient.getresponseasync" rel="noopener noreferrer"&gt;&lt;code&gt;IChatClient.GetResponseAsync&lt;/code&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.ai.chatclientstructuredoutputextensions" rel="noopener noreferrer"&gt;Structured-output extensions for &lt;code&gt;IChatClient&lt;/code&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;dotnet/extensions&lt;/code&gt;: &lt;a href="https://github.com/dotnet/extensions/blob/main/src/Libraries/Microsoft.Extensions.AI/ChatCompletion/ChatClientStructuredOutputExtensions.cs" rel="noopener noreferrer"&gt;&lt;code&gt;GetResponseAsync&amp;lt;T&amp;gt;&lt;/code&gt; structured-output implementation&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.ai.chatoptions.rawrepresentationfactory" rel="noopener noreferrer"&gt;&lt;code&gt;ChatOptions.RawRepresentationFactory&lt;/code&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.dependencyinjection.optionsservicecollectionextensions.addoptions" rel="noopener noreferrer"&gt;&lt;code&gt;AddOptions&lt;/code&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.dependencyinjection.optionsbuilderextensions.validateonstart" rel="noopener noreferrer"&gt;&lt;code&gt;ValidateOnStart&lt;/code&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/overview/azure/ai.openai-readme" rel="noopener noreferrer"&gt;Azure OpenAI client library for .NET&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/working-with-models" rel="noopener noreferrer"&gt;Azure OpenAI model deployments&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;dotnet/extensions&lt;/code&gt;: &lt;a href="https://github.com/dotnet/extensions/tree/main/src/Libraries/Microsoft.Extensions.AI.OpenAI" rel="noopener noreferrer"&gt;&lt;code&gt;Microsoft.Extensions.AI.OpenAI&lt;/code&gt; examples&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;OllamaSharp: &lt;a href="https://github.com/awaescher/OllamaSharp" rel="noopener noreferrer"&gt;&lt;code&gt;IChatClient&lt;/code&gt; integration&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;NuGet: &lt;a href="https://www.nuget.org/packages/Microsoft.Extensions.AI/10.9.0" rel="noopener noreferrer"&gt;&lt;code&gt;Microsoft.Extensions.AI&lt;/code&gt; 10.9.0&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;NuGet: &lt;a href="https://www.nuget.org/packages/Microsoft.Extensions.AI.OpenAI/10.9.0" rel="noopener noreferrer"&gt;&lt;code&gt;Microsoft.Extensions.AI.OpenAI&lt;/code&gt; 10.9.0&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;NuGet: &lt;a href="https://www.nuget.org/packages/OllamaSharp/5.4.30" rel="noopener noreferrer"&gt;&lt;code&gt;OllamaSharp&lt;/code&gt; 5.4.30&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>dotnet</category>
      <category>csharp</category>
      <category>ai</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Retries Are Not a Recovery Strategy</title>
      <dc:creator>Lukas Walter </dc:creator>
      <pubDate>Thu, 27 Aug 2026 15:30:00 +0000</pubDate>
      <link>https://dev.to/lukaswalter/retries-are-not-a-recovery-strategy-mb3</link>
      <guid>https://dev.to/lukaswalter/retries-are-not-a-recovery-strategy-mb3</guid>
      <description>&lt;p&gt;A retry answers a narrow question: might the same operation succeed if I attempt it again?&lt;/p&gt;

&lt;p&gt;Recovery has a harder job. It must bring the original business operation to a known, valid outcome after something went wrong. Getting there may require another attempt, a status lookup, resuming from persisted state, or compensation. If the system cannot resolve the operation safely, it must hand it to a person.&lt;/p&gt;

&lt;p&gt;This difference matters as soon as an AI workflow does more than return text. If it retrieves data, calls tools, writes state, or continues after the HTTP request ends, adding three retries around the workflow is not a recovery design. It is three more chances to spend money, repeat a side effect, or lose track of what already happened.&lt;/p&gt;

&lt;h2&gt;
  
  
  A retry repeats an attempt
&lt;/h2&gt;

&lt;p&gt;Suppose a support feature performs this workflow:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;load the ticket and approved policy
    -&amp;gt; generate a reply
    -&amp;gt; validate the reply
    -&amp;gt; save it as a draft
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The policy read returns &lt;code&gt;503 Service Unavailable&lt;/code&gt; with an applicable &lt;code&gt;Retry-After&lt;/code&gt; response, and the dependency contract classifies it as transient. No application business state changed, and the request still has time left. A delayed retry may be reasonable.&lt;/p&gt;

&lt;p&gt;Now suppose the draft save times out after the request reached the database. The caller cannot tell whether the write committed. Repeating the complete workflow creates a new model response and may save a second draft. Retrying only the write is safe when the write is naturally idempotent, or when the boundary can recognize the retry as the same logical operation. Otherwise, the second attempt may create another draft.&lt;/p&gt;

&lt;p&gt;Both failures may appear as a timeout or dependency exception in application code. They do not have the same effect.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What happened&lt;/th&gt;
&lt;th&gt;What is known&lt;/th&gt;
&lt;th&gt;Suitable response&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A transient policy read failed before returning data&lt;/td&gt;
&lt;td&gt;No application business state changed&lt;/td&gt;
&lt;td&gt;Retry the read within its budget&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The model endpoint rejected an invalid request&lt;/td&gt;
&lt;td&gt;The same request will fail again&lt;/td&gt;
&lt;td&gt;Stop and fix the request or contract&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The model call timed out&lt;/td&gt;
&lt;td&gt;No application state changed, but provider work and cost may already have occurred&lt;/td&gt;
&lt;td&gt;Retry only if the result is still useful and budget remains&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A draft save was submitted and its response was lost&lt;/td&gt;
&lt;td&gt;The draft may already exist&lt;/td&gt;
&lt;td&gt;Look up the original operation or retry it with the same idempotency identity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A multi-step workflow stopped after some steps completed&lt;/td&gt;
&lt;td&gt;The operation is partially complete&lt;/td&gt;
&lt;td&gt;Resume, reconcile, compensate, or escalate according to persisted workflow state&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Retry belongs to the first few rows. Recovery begins when the system must find out what happened or restore a valid business state.&lt;/p&gt;

&lt;h2&gt;
  
  
  A timeout does not mean the operation failed
&lt;/h2&gt;

&lt;p&gt;When the caller's timeout expires, it tells the caller that it stopped waiting. It does not prove that the callee stopped working, rolled back, or never received the request. The remote operation may never have started, may have failed, may still be running, or may have completed while its response was lost.&lt;/p&gt;

&lt;p&gt;Those states should not collapse into one &lt;code&gt;Failed&lt;/code&gt; result. I find it useful to separate the operation status from its business outcome:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;enum&lt;/span&gt; &lt;span class="n"&gt;OperationStatus&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;Pending&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;Running&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;PartiallyCompleted&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;OutcomeUnknown&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;Terminal&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;NeedsReview&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;enum&lt;/span&gt; &lt;span class="n"&gt;OperationOutcome&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;Succeeded&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;Rejected&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;Failed&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;Compensated&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;sealed&lt;/span&gt; &lt;span class="k"&gt;record&lt;/span&gt; &lt;span class="nc"&gt;OperationState&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;OperationStatus&lt;/span&gt; &lt;span class="n"&gt;Status&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;OperationOutcome&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="n"&gt;Outcome&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This type is deliberately simplified. As written, it permits invalid combinations such as &lt;code&gt;Running&lt;/code&gt; with &lt;code&gt;Succeeded&lt;/code&gt;, or &lt;code&gt;Terminal&lt;/code&gt; with no outcome. Production code should enforce those invariants through its type design or validated construction. &lt;code&gt;OperationStatus&lt;/code&gt; is a top-level recovery status, not a complete per-step workflow model. A workflow can be partially complete while the outcome of its latest step is still unknown.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;Failed&lt;/code&gt; is a terminal unsuccessful outcome supported by authoritative evidence. A payment decline, for example, is a completed request with a &lt;code&gt;Rejected&lt;/code&gt; business outcome. &lt;code&gt;OutcomeUnknown&lt;/code&gt; means the application cannot yet determine the terminal outcome. &lt;code&gt;NeedsReview&lt;/code&gt; is not a business outcome at all. It transfers ownership from automatic recovery to a person.&lt;/p&gt;

&lt;p&gt;That extra state is mildly inconvenient. Good. The uncertainty already exists in the system, whether the type admits it or not. Encoding it gives the API, UI, recovery worker, and support tooling something honest to work with.&lt;/p&gt;

&lt;p&gt;For the draft save, the application can allocate an operation ID before it submits the write:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="n"&gt;Guid&lt;/span&gt; &lt;span class="n"&gt;operationId&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Guid&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;NewGuid&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="kt"&gt;var&lt;/span&gt; &lt;span class="n"&gt;command&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;SaveDraftCommand&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;OperationId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;operationId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;TicketId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;ticketId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;Reply&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;reply&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;draftOperations&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;CreateAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;operationId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;command&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;try&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;draftStore&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;SaveAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;command&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;draftOperations&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;MarkTerminalAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;operationId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;OperationOutcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Succeeded&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="k"&gt;catch&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;DraftSaveOutcomeUnknownException&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;var&lt;/span&gt; &lt;span class="n"&gt;recoveryWriteCts&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;CancellationTokenSource&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;TimeSpan&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;FromSeconds&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="m"&gt;2&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;

    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;draftOperations&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;MarkOutcomeUnknownAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;operationId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;recoveryWriteCts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Token&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="k"&gt;throw&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The example uses &lt;code&gt;DraftSaveOutcomeUnknownException&lt;/code&gt; for a specific conclusion from the storage adapter: the attempt crossed the submission boundary, but the adapter cannot determine whether the write committed. A plain &lt;code&gt;OperationCanceledException&lt;/code&gt; is not enough. Cancellation may happen before anything was submitted, in which case the write outcome is known.&lt;/p&gt;

&lt;p&gt;Production code must still distinguish caller cancellation from an attempt timeout. It also has to handle failures while recording the state transition and prevent an older attempt from overwriting a newer result. Create the durable identity before the ambiguous boundary, then keep it for lookup and any safe retry.&lt;/p&gt;

&lt;p&gt;The persistence contract must enforce uniqueness for the operation ID and reject reuse with different command data. Merely logging a GUID does not make the write idempotent.&lt;/p&gt;

&lt;h2&gt;
  
  
  Retrying the workflow is usually the wrong scope
&lt;/h2&gt;

&lt;p&gt;Broad retry policies are attractive because they are easy to add around an endpoint or orchestration method. They also repeat every successful step before the failure.&lt;/p&gt;

&lt;p&gt;If the support workflow fails while saving the draft, a workflow-level retry may retrieve the same policy again, pay for another model call, produce a different reply, and run validation again before it reaches the uncertain write. The retry has changed the thing being recovered.&lt;/p&gt;

&lt;p&gt;AI workflows make this especially awkward:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Model calls are metered and can be slow.&lt;/li&gt;
&lt;li&gt;A repeated generation is not guaranteed to return the same output.&lt;/li&gt;
&lt;li&gt;Tool calls can hide writes behind an interface that looks like an ordinary function call.&lt;/li&gt;
&lt;li&gt;SDKs, HTTP clients, queues, and workflow code may each have their own retry behavior.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Agent loops add another failure mode. Consider an approved workflow that is allowed to charge a customer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;model selects ChargeCustomer(...)
    -&amp;gt; application validates and authorizes the tool call
    -&amp;gt; payment succeeds
    -&amp;gt; tool response is lost
    -&amp;gt; agent loop starts again
    -&amp;gt; model produces another ChargeCustomer(...) call
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second model run may use a new tool-call ID or produce slightly different arguments. Approval and argument validation do not tell the payment boundary that this is the same charge. The application needs a stable business operation ID that exists independently of the model output and survives another model turn. Otherwise, the retry has changed both the reasoning and the command being recovered.&lt;/p&gt;

&lt;p&gt;That last point can turn a small setting into a large amount of traffic. Three attempts in one layer and three in the next can produce up to nine calls to the struggling dependency. Under load, the retries may prolong the incident they were supposed to smooth over.&lt;/p&gt;

&lt;p&gt;Per-request limits may still allow too much retry traffic when many requests fail together. A dependency-level retry budget caps the aggregate retries generated by concurrent requests. Once the budget is spent, requests get no additional attempts. Their initial attempts are a separate decision. A circuit breaker or admission-control policy can make new requests fail fast.&lt;/p&gt;

&lt;p&gt;Retry the smallest operation whose failure is known to be transient and whose repetition is safe. Do not restart a workflow simply because one of its steps threw an exception.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recovery starts with durable operation state
&lt;/h2&gt;

&lt;p&gt;An in-memory retry loop can help the current request, but it does not provide durable ownership of the operation. A process restart or deployment can destroy its state, and request-scoped work may end when the caller disconnects.&lt;/p&gt;

&lt;p&gt;When the result must still be resolved later, the application needs a durable record of the logical operation. That record needs enough information to identify and continue the same work:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a stable operation ID, plus any boundary-specific idempotency identities needed for safe replay&lt;/li&gt;
&lt;li&gt;the command or a durable reference to it&lt;/li&gt;
&lt;li&gt;a command fingerprint used to detect and reject the same key being reused for different work&lt;/li&gt;
&lt;li&gt;the current state and attempt identity&lt;/li&gt;
&lt;li&gt;the next eligible recovery time&lt;/li&gt;
&lt;li&gt;enough outcome information for status lookup and support&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A recovery worker can then process unresolved operations without inventing a new logical request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Pending
    -&amp;gt; submit with the original idempotency key

OutcomeUnknown
    -&amp;gt; query the authoritative status
    -&amp;gt; retry the original command only when the contract makes that safe

PartiallyCompleted
    -&amp;gt; resume the next incomplete step or run a domain-specific compensation

Still unresolved after the recovery deadline
    -&amp;gt; NeedsReview
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Recovery does not always mean forcing the requested state change to succeed. It may confirm a payment decline and close the operation with a &lt;code&gt;Rejected&lt;/code&gt; outcome. If an external reservation succeeded but a later step failed, recovery may cancel the reservation and record &lt;code&gt;Compensated&lt;/code&gt;. If neither action is safe to automate, the operation moves to &lt;code&gt;NeedsReview&lt;/code&gt; and a person owns the next decision.&lt;/p&gt;

&lt;p&gt;Recovery should produce a terminal, explainable business outcome or transfer ownership to manual review. An unresolved operation should not simply disappear when the retry loop stops.&lt;/p&gt;

&lt;h2&gt;
  
  
  Idempotency protects one boundary
&lt;/h2&gt;

&lt;p&gt;An idempotency key tells a boundary that repeated submissions refer to one logical command. The boundary must atomically claim or enforce uniqueness for that command, including when two requests arrive concurrently. It can prevent the second request from applying the side effect again and, once available, return or reference the recorded result.&lt;/p&gt;

&lt;p&gt;While the first request is still running, there is no result to return. The contract may expose the current operation status, wait for completion, or report that the operation is already in progress. A separate check followed by an insert leaves a race in which both requests can execute.&lt;/p&gt;

&lt;p&gt;That protection is local to the boundary that enforces it. A database uniqueness constraint does not deduplicate an email already sent by another service. A job ID does not protect a payment call unless the same identity reaches the payment provider and the provider honors it.&lt;/p&gt;

&lt;p&gt;Each side-effect boundary needs its own guarantee. A transactional outbox can make the handoff from a business transaction to messaging durable, but the relay may publish a message more than once. The receiving boundary still needs an inbox, an idempotent consumer, a provider-supported idempotency key, or a domain-specific deduplication rule.&lt;/p&gt;

&lt;p&gt;Idempotency also does not answer these questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Should the operation still run after the caller has gone away?&lt;/li&gt;
&lt;li&gt;How long should the system keep trying?&lt;/li&gt;
&lt;li&gt;What happens when status lookup remains unavailable?&lt;/li&gt;
&lt;li&gt;Can completed steps be reversed?&lt;/li&gt;
&lt;li&gt;Who owns an operation that never reaches a terminal state?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those are recovery decisions.&lt;/p&gt;

&lt;h2&gt;
  
  
  A retry policy still needs boundaries
&lt;/h2&gt;

&lt;p&gt;Retries are useful when the failure is plausibly transient, repeating the attempt is safe, and another result would still arrive in time to matter.&lt;/p&gt;

&lt;p&gt;For each retryable interaction, define:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Which outcomes are transient. Do not retry validation errors, authorization denials, or other requests that cannot succeed unchanged.&lt;/li&gt;
&lt;li&gt;The scope of the attempt. Retry the failed read or model call, not an entire sequence of completed work.&lt;/li&gt;
&lt;li&gt;The per-attempt timeout and total budget. A new attempt should not start when there is no useful time left.&lt;/li&gt;
&lt;li&gt;The maximum attempts and delay. Respect &lt;code&gt;Retry-After&lt;/code&gt; where available, and use backoff with jitter rather than synchronizing every instance.&lt;/li&gt;
&lt;li&gt;The behavior after the attempts end. Return a specific application outcome, let the circuit-breaker policy update its failure state, degrade explicitly, or hand the operation to durable recovery.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Also inspect retry behavior already present in SDKs and infrastructure. Adding an application policy without accounting for lower layers makes attempt counts and latency hard to predict.&lt;/p&gt;

&lt;p&gt;Record the attempt number, duration, outcome category, and operation ID in telemetry. A request that succeeds on its third attempt still represents two failed dependency calls, more latency, and more cost. Occasional transient failures are normal. A rising retry rate or an increasing share of requests that need several attempts may show that the dependency is degrading. Final success alone hides that change.&lt;/p&gt;

&lt;h2&gt;
  
  
  When retries are enough
&lt;/h2&gt;

&lt;p&gt;A bounded retry policy is often enough for an idempotent read or another side-effect-free call when the failure is clearly transient. The work stays inside the current request, no state is uncertain, and exhausting the attempts can return an honest failure to the caller.&lt;/p&gt;

&lt;p&gt;Examples include a throttled metadata lookup, a connection failure where the client can establish that the request was never submitted, or a model call that has no tools and whose repeated cost is acceptable inside the remaining request budget.&lt;/p&gt;

&lt;h2&gt;
  
  
  When you need a recovery path
&lt;/h2&gt;

&lt;p&gt;Design recovery when an operation changes state, spans several independently committed steps, or must reach a terminal result after the current request ends. You also need it when a response can disappear after a side effect commits, or when an external action requires compensation or human review.&lt;/p&gt;

&lt;p&gt;Recovery costs more to build. It needs persisted state, ownership, deadlines, concurrency control, telemetry, and support procedures. That cost is a reason to keep workflows small and side effects explicit. It is not a reason to call a retry loop recovery.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practical takeaway
&lt;/h2&gt;

&lt;p&gt;Pick one state-changing AI workflow and find every retry around it, including retries inside SDKs and infrastructure.&lt;/p&gt;

&lt;p&gt;For each one, write down what is known after the failed attempt, whether state may have changed, which identity the next attempt uses, and what happens after attempts are exhausted. Returning an error is a valid outcome when no state changed and the application has no promise to continue the work.&lt;/p&gt;

&lt;p&gt;If state may already have changed, the operation is partially complete, or the business contract requires a terminal result after the request ends, "return an error and hope the user tries again" is not enough. That workflow still needs a recovery path.&lt;/p&gt;

&lt;p&gt;Use an explicit &lt;code&gt;OutcomeUnknown&lt;/code&gt; path when the application loses the response after submitting a side effect and cannot determine whether it completed. Resolve that state through authoritative lookup, a safe retry with the original identity, compensation, or manual review.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/posts/failure-modes-before-happy-paths/" rel="noopener noreferrer"&gt;What Happens When Your AI Feature Fails?&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/tips/use-idempotency-keys-for-retryable-writes/" rel="noopener noreferrer"&gt;Use idempotency keys for retryable writes&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/posts/when-ai-request-becomes-background-job/" rel="noopener noreferrer"&gt;When an AI Request Should Become a Background Job&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/azure/architecture/patterns/retry" rel="noopener noreferrer"&gt;Retry pattern&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/azure/architecture/best-practices/transient-faults" rel="noopener noreferrer"&gt;Best practices for transient fault handling&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/azure/architecture/patterns/compensating-transaction" rel="noopener noreferrer"&gt;Compensating Transaction pattern&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/azure/architecture/antipatterns/retry-storm/" rel="noopener noreferrer"&gt;Retry Storm antipattern&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/samples/azure-samples/cosmos-db-design-patterns/transactional-outbox/" rel="noopener noreferrer"&gt;Azure Cosmos DB design pattern: Transactional Outbox&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>dotnet</category>
      <category>ai</category>
      <category>architecture</category>
      <category>systemdesign</category>
    </item>
    <item>
      <title>Embeddings in .NET</title>
      <dc:creator>Lukas Walter </dc:creator>
      <pubDate>Mon, 24 Aug 2026 15:30:00 +0000</pubDate>
      <link>https://dev.to/lukaswalter/embeddings-in-net-imd</link>
      <guid>https://dev.to/lukaswalter/embeddings-in-net-imd</guid>
      <description>&lt;p&gt;Generating one in .NET is a short API call. The part worth designing is the contract around that call: which model or encoder produced each vector, how many values it contains, which comparison metric will be used during search, and whether indexed documents and runtime queries use compatible rules.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;Microsoft.Extensions.AI&lt;/code&gt; gives .NET applications a provider-neutral &lt;code&gt;IEmbeddingGenerator&amp;lt;TInput, TEmbedding&amp;gt;&lt;/code&gt; abstraction. It keeps provider SDK types out of the retrieval code, but provider neutrality does not imply vector-space compatibility.&lt;/p&gt;

&lt;p&gt;Before the first document is indexed, four choices need to be fixed together: the model, vector shape, comparison metric, and input preparation. The examples below use dense text embeddings for semantic retrieval.&lt;/p&gt;

&lt;h2&gt;
  
  
  An embedding is a model output, not application data
&lt;/h2&gt;

&lt;p&gt;An embedding model, together with its configuration, maps an input into a vector of numbers with a defined dimensionality.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"stop a worker during shutdown"
-&amp;gt; embedding model
-&amp;gt; [0.018, -0.042, 0.007, ...]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The vector is a compact representation learned by the model. Inputs the model represents as similar should end up near one another under a comparison metric appropriate for that embedding space.&lt;/p&gt;

&lt;p&gt;The individual coordinates do not have stable application meanings. Position 37 is not "shutdown behavior," and a larger value at that position does not make one passage more relevant. The vector works as a whole.&lt;/p&gt;

&lt;p&gt;This also means an embedding is lossy. You cannot reconstruct the source passage reliably from the vector, and the vector does not retain the document ID, permissions, version, or source URL. Keep the original text and its metadata. The embedding is an index representation of that source, not a replacement for it.&lt;/p&gt;

&lt;p&gt;Similarity scores need context too. Read a cosine score within the embedding model, retrieval setup, and corpus that produced it. Adding candidates does not change the score between two vectors. It can change where a passage ranks and whether it remains in the top results. The score distribution across a corpus also affects whether a fixed threshold is useful. The score is not a universal confidence value, a probability that the passage answers the question, or proof that the passage is correct.&lt;/p&gt;

&lt;p&gt;Embeddings help rank candidates. They do not own truth or authorization.&lt;/p&gt;

&lt;h2&gt;
  
  
  Documents and queries must enter the same vector space
&lt;/h2&gt;

&lt;p&gt;A basic semantic retrieval path has two sides.&lt;/p&gt;

&lt;p&gt;During ingestion:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;document chunk
-&amp;gt; embedding generator
-&amp;gt; document vector
-&amp;gt; vector index
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;During a query:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;user question
-&amp;gt; same embedding contract
-&amp;gt; query vector
-&amp;gt; nearest-neighbor search
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The search works because both vectors are coordinates in the same space. Their text does not need to contain the same words. The configured model or encoders should place semantically related inputs close enough for the search step to find them.&lt;/p&gt;

&lt;p&gt;"Same embedding contract" is more precise than "same number of floats." Two models can return vectors with equal dimensions and still use unrelated coordinate systems. Vectors from independently chosen models generally cannot be compared meaningfully. They are compatible only when the models or encoders are explicitly designed to produce representations in the same retrieval space.&lt;/p&gt;

&lt;p&gt;This is the first boundary to keep clear after deciding that a feature really needs retrieval. &lt;a href="https://www.lukaswalter.dev/posts/why-retrieval-exists/" rel="noopener noreferrer"&gt;Why Retrieval Exists&lt;/a&gt; covers that earlier decision. Once vector search is part of the retrieval path, the embedding contract becomes part of the index schema.&lt;/p&gt;

&lt;h2&gt;
  
  
  Generate an embedding through &lt;code&gt;IEmbeddingGenerator&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;For an OpenAI-backed example, install the &lt;code&gt;Microsoft.Extensions.AI.OpenAI&lt;/code&gt; package:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dotnet add package Microsoft.Extensions.AI.OpenAI &lt;span class="nt"&gt;--version&lt;/span&gt; 10.9.0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I compiled the sample below with .NET 10 and &lt;code&gt;Microsoft.Extensions.AI.OpenAI&lt;/code&gt; 10.9.0 from the public NuGet feed. It uses the &lt;a href="https://learn.microsoft.com/en-us/dotnet/ai/iembeddinggenerator" rel="noopener noreferrer"&gt;&lt;code&gt;IEmbeddingGenerator&lt;/code&gt;&lt;/a&gt; API. Check package versions before copying it into an application.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;Microsoft.Extensions.AI&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="nn"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;apiKey&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Environment&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GetEnvironmentVariable&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"OPENAI_API_KEY"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;??&lt;/span&gt; &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;InvalidOperationException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="s"&gt;"Set OPENAI_API_KEY before running the sample."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="n"&gt;IEmbeddingGenerator&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Embedding&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;float&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;generator&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt;
    &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;OpenAIClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;apiKey&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GetEmbeddingClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"text-embedding-3-small"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;AsIEmbeddingGenerator&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="n"&gt;ReadOnlyMemory&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;float&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;queryVector&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;generator&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GenerateVectorAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="s"&gt;"How do I stop a background service cleanly?"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;CancellationToken&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;None&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="n"&gt;Console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;WriteLine&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;$"Generated &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;queryVector&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Length&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt; values."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;GenerateVectorAsync&lt;/code&gt; is a convenience method for one input. It returns the vector as &lt;code&gt;ReadOnlyMemory&amp;lt;float&amp;gt;&lt;/code&gt;, which can be passed to a vector-store client or copied into a persistence type when the client requires an array.&lt;/p&gt;

&lt;p&gt;The provider-specific client is created at the composition boundary. Application code can depend on &lt;code&gt;IEmbeddingGenerator&amp;lt;string, Embedding&amp;lt;float&amp;gt;&amp;gt;&lt;/code&gt; instead of &lt;code&gt;OpenAIClient&lt;/code&gt;. Another provider can supply the same interface, but its vectors still have to belong to the index's embedding space. Changing to an incompatible model, model revision, or embedding configuration requires reindexing.&lt;/p&gt;

&lt;p&gt;Keep the API key in the environment, a development secret store, or a managed secret service. Do not put it in source code or configuration committed to the repository.&lt;/p&gt;

&lt;h2&gt;
  
  
  Generate document embeddings in batches
&lt;/h2&gt;

&lt;p&gt;Ingestion rarely embeds one passage at a time. Send a supported batch so the provider can process several inputs in one request.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt; &lt;span class="n"&gt;passages&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt;
&lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="s"&gt;"Use CancellationToken to stop long-running work cooperatively."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s"&gt;"Register hosted services with AddHostedService."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s"&gt;"A bounded channel can apply backpressure to producers."&lt;/span&gt;
&lt;span class="p"&gt;];&lt;/span&gt;

&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;Value&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Embedding&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;float&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;Embedding&lt;/span&gt;&lt;span class="p"&gt;)[]&lt;/span&gt; &lt;span class="n"&gt;embeddedPassages&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;generator&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GenerateAndZipAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;passages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;foreach&lt;/span&gt; &lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;passage&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Embedding&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;float&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;embedding&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt;
    &lt;span class="n"&gt;embeddedPassages&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;UpsertAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;passage&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;vector&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;embedding&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Vector&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;GenerateAndZipAsync&lt;/code&gt; keeps each input beside its returned embedding. That pairing matters once the pipeline also carries stable chunk IDs, document versions, and payload metadata.&lt;/p&gt;

&lt;p&gt;Batch size is an operational setting, not a constant copied from this example. Providers impose different input-count and token limits. Start with a conservative size, pass cancellation through the complete ingestion path, and observe latency and throttling before increasing concurrency.&lt;/p&gt;

&lt;p&gt;For a real pipeline, the value sent to &lt;code&gt;UpsertAsync&lt;/code&gt; should also include a stable point ID and the metadata needed for filtering and traceability. Keep the collection schema and indexing details in the vector-store layer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Treat the embedding profile as an index contract
&lt;/h2&gt;

&lt;p&gt;The C# type only says that the vector contains &lt;code&gt;float&lt;/code&gt; values. It does not tell the compiler which embedding space, dimensions, or comparison rules belong to those values.&lt;/p&gt;

&lt;p&gt;Make that missing context explicit as an embedding profile.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;sealed&lt;/span&gt; &lt;span class="k"&gt;record&lt;/span&gt; &lt;span class="nc"&gt;EmbeddingProfile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;Id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;Provider&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;ModelIdentity&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;Dimensions&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;ComparisonMetric&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;InputMode&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;PreprocessingVersion&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;static&lt;/span&gt; &lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;SupportEmbeddings&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;static&lt;/span&gt; &lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="n"&gt;EmbeddingProfile&lt;/span&gt; &lt;span class="n"&gt;V1&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;Id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"support-text-v1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;Provider&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"OpenAI"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;ModelIdentity&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"text-embedding-3-small"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;Dimensions&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="m"&gt;1536&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;ComparisonMetric&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"Cosine"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;InputMode&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"default"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;PreprocessingVersion&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"plain-text-v1"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model and 1,536 dimensions match the default output used above. Cosine is the comparison metric selected for the index and is consistent with &lt;a href="https://platform.openai.com/docs/guides/embeddings" rel="noopener noreferrer"&gt;OpenAI's recommendation&lt;/a&gt;. If you request a different output size or select another model, the profile must change.&lt;/p&gt;

&lt;p&gt;The profile does not need to be this record. It can live in typed options plus deployment metadata, as long as the application can answer which contract produced an index.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;ModelIdentity&lt;/code&gt; should identify the model and, when the provider exposes one, the specific revision or snapshot. A deployment identifier can be tracked separately when it is operationally useful, but it does not define compatibility by itself.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;Provider&lt;/code&gt; is provenance, not a compatibility rule. Keep it because operators may need to trace how an index was produced. Changing that value alone does not prove that the vector space changed.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;InputMode&lt;/code&gt; and &lt;code&gt;PreprocessingVersion&lt;/code&gt; sound similar, but they own different decisions. &lt;code&gt;InputMode&lt;/code&gt; records model-specific behavior such as query and document roles or prefixes. The OpenAI example uses &lt;code&gt;default&lt;/code&gt; because it does not set a model-specific mode. &lt;code&gt;PreprocessingVersion&lt;/code&gt; identifies how the application transforms the source text.&lt;/p&gt;

&lt;p&gt;I would still give the complete profile a stable ID. A model name on its own leaves too much unsaid.&lt;/p&gt;

&lt;p&gt;Keep compatibility decisions and useful provenance together:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Constraint&lt;/th&gt;
&lt;th&gt;What must stay compatible&lt;/th&gt;
&lt;th&gt;Typical failure&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Embedding space&lt;/td&gt;
&lt;td&gt;Document and query vectors must come from the same embedding contract or from encoders explicitly designed to produce compatible representations.&lt;/td&gt;
&lt;td&gt;Search can run successfully while ranking becomes meaningless.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dimensions&lt;/td&gt;
&lt;td&gt;Generated vectors must match the index schema.&lt;/td&gt;
&lt;td&gt;The vector store rejects writes or queries.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Comparison metric&lt;/td&gt;
&lt;td&gt;The index must use a metric appropriate for the embedding space.&lt;/td&gt;
&lt;td&gt;Score semantics can change, and ranking quality can degrade for some models.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Input mode&lt;/td&gt;
&lt;td&gt;Query and document prefixes or task modes must follow the model's instructions.&lt;/td&gt;
&lt;td&gt;Retrieval quality drops even though dimensions match.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Preprocessing version&lt;/td&gt;
&lt;td&gt;Normalization and content preparation must remain reproducible.&lt;/td&gt;
&lt;td&gt;Reindexed and older records represent different inputs.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Qdrant, for example, requires vectors stored under one vector configuration to share a dimensionality and comparison metric. Named vectors can define separate configurations inside the same collection. Its &lt;a href="https://qdrant.tech/documentation/manage-data/collections/" rel="noopener noreferrer"&gt;collection documentation&lt;/a&gt; also notes that the metric should follow how the encoder was trained.&lt;/p&gt;

&lt;p&gt;For the OpenAI model in this example, metric choice needs one more qualification. OpenAI recommends cosine similarity and states that its embeddings are normalized to length 1. Cosine similarity can therefore be calculated with a dot product. Cosine and Euclidean comparison also produce the same ranking for these embeddings, although their score values have different semantics.&lt;/p&gt;

&lt;p&gt;The database can reject a vector with the wrong length. It cannot detect that a 1,536-value query came from the wrong 1,536-dimensional model. That failure is more dangerous because the request succeeds.&lt;/p&gt;

&lt;h2&gt;
  
  
  Validate the contract before indexing
&lt;/h2&gt;

&lt;p&gt;Configuration drift should fail early, not after a retrieval-quality incident.&lt;/p&gt;

&lt;p&gt;One small startup or integration check can verify the output shape:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;static&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="n"&gt;Task&lt;/span&gt; &lt;span class="nf"&gt;VerifyProfileAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;IEmbeddingGenerator&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Embedding&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;float&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;generator&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;EmbeddingProfile&lt;/span&gt; &lt;span class="n"&gt;profile&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;CancellationToken&lt;/span&gt; &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;ReadOnlyMemory&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;float&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;probe&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;generator&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GenerateVectorAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="s"&gt;"embedding-profile-check"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;probe&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Length&lt;/span&gt; &lt;span class="p"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;profile&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Dimensions&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;InvalidOperationException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="s"&gt;$"Embedding profile '&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;profile&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Id&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;' expects "&lt;/span&gt; &lt;span class="p"&gt;+&lt;/span&gt;
            &lt;span class="s"&gt;$"&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;profile&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Dimensions&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt; dimensions but the configured "&lt;/span&gt; &lt;span class="p"&gt;+&lt;/span&gt;
            &lt;span class="s"&gt;$"generator returned &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;probe&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Length&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s"&gt;."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This check catches a dimensional mismatch. It does not prove that the configured model is the intended one because two models may return the same size. Verify the model identity and embedding configuration as well. Track the deployment separately when it matters for operations.&lt;/p&gt;

&lt;p&gt;Also check the vector-store schema against the same profile before starting ingestion. The generator, collection, and query path should not each carry their own independent copy of the expected dimension.&lt;/p&gt;

&lt;h2&gt;
  
  
  Treat an incompatible model change as a data migration
&lt;/h2&gt;

&lt;p&gt;An incompatible embedding model uses a different coordinate system from the existing index.&lt;/p&gt;

&lt;p&gt;If only the query-side generator changes, its vectors cannot be compared meaningfully with the existing document vectors. If only new documents use the new profile, the collection ends up split across two spaces. Neither failure has to produce an exception.&lt;/p&gt;

&lt;p&gt;For this kind of change, I create a new embedding profile and re-embed the corpus. The cost is visible. Mixing vector spaces can fail quietly, which is worse.&lt;/p&gt;

&lt;p&gt;A migration usually needs these stages:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Create storage for the new profile with its dimensions and comparison metric.&lt;/li&gt;
&lt;li&gt;Backfill existing source records through the new generator.&lt;/li&gt;
&lt;li&gt;Write new or changed records to both profiles while the backfill runs, if the application must stay current.&lt;/li&gt;
&lt;li&gt;Evaluate the new index with known queries and expected sources.&lt;/li&gt;
&lt;li&gt;Switch reads to the new profile.&lt;/li&gt;
&lt;li&gt;Retire the old vectors after a rollback window.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Qdrant supports migrations through a new collection and alias swap. Qdrant 1.18 and later can also add a named vector for a second model, which allows backfilling inside the existing collection. Its &lt;a href="https://qdrant.tech/documentation/tutorials-operations/embedding-model-migration/" rel="noopener noreferrer"&gt;embedding-model migration guide&lt;/a&gt; shows the dual-write sequence.&lt;/p&gt;

&lt;p&gt;Do not compare raw similarity scores from the old and new profiles as if they shared a scale. Evaluate whether each profile retrieves the expected sources at the rank positions your application uses.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test retrieval, not attractive coordinates
&lt;/h2&gt;

&lt;p&gt;Printing a few vector values proves that the API returned numbers. A two-dimensional chart can make a demo easier to understand. Neither one tells you whether the model works for your corpus.&lt;/p&gt;

&lt;p&gt;Build a small evaluation set before committing to a profile:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;representative user questions&lt;/li&gt;
&lt;li&gt;the chunks or documents that should be retrieved&lt;/li&gt;
&lt;li&gt;difficult wording differences and domain terms&lt;/li&gt;
&lt;li&gt;cases where lexical search should beat semantic search&lt;/li&gt;
&lt;li&gt;languages the application must support&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then measure whether the expected sources appear in the top results. Compare candidate models, preprocessing rules, and hybrid-search choices against the same set.&lt;/p&gt;

&lt;p&gt;Model descriptions and public benchmarks can narrow the options. Your retrieval questions decide whether the profile is good enough for the application.&lt;/p&gt;

&lt;h2&gt;
  
  
  When embeddings fit
&lt;/h2&gt;

&lt;p&gt;Use embeddings when the application needs semantic ranking across a candidate corpus and exact wording is unreliable. Typical examples include finding a runbook from a problem description, locating related support cases, or selecting documentation passages for RAG.&lt;/p&gt;

&lt;p&gt;Do not use an embedding when the request already contains an exact identifier, when structured filters or SQL express the question directly, or when application code owns the decision. Keep permissions and tenant boundaries in trusted filters. A close vector is not an authorization result.&lt;/p&gt;

&lt;p&gt;Embeddings are also a poor substitute for source metadata. Store the document ID, version, language, permissions, and original text beside the vector or in a system that the indexed point references.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep the contract visible
&lt;/h2&gt;

&lt;p&gt;The minimal .NET call is simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="n"&gt;ReadOnlyMemory&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;float&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;vector&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;generator&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;GenerateVectorAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The production code needs to keep more than the returned values:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a provider-neutral &lt;code&gt;IEmbeddingGenerator&lt;/code&gt; at the application boundary&lt;/li&gt;
&lt;li&gt;an explicit embedding profile for the embedding space and dimensions&lt;/li&gt;
&lt;li&gt;a comparison metric chosen for that embedding space&lt;/li&gt;
&lt;li&gt;reproducible input preparation&lt;/li&gt;
&lt;li&gt;a migration path that re-embeds existing data&lt;/li&gt;
&lt;li&gt;retrieval tests with expected sources&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If those decisions are visible, the next step is mechanical: create a vector-store collection that matches the profile, then index chunks with stable IDs and useful payload data.&lt;/p&gt;

&lt;h2&gt;
  
  
  Further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/ai/iembeddinggenerator" rel="noopener noreferrer"&gt;Use the &lt;code&gt;IEmbeddingGenerator&lt;/code&gt; interface&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.ai.embeddinggeneratorextensions.generatevectorasync?view=net-10.0-pp" rel="noopener noreferrer"&gt;&lt;code&gt;GenerateVectorAsync&lt;/code&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/dotnet/api/microsoft.extensions.ai.embeddinggeneratorextensions.generateandzipasync?view=net-10.0-pp" rel="noopener noreferrer"&gt;&lt;code&gt;GenerateAndZipAsync&lt;/code&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;NuGet: &lt;a href="https://www.nuget.org/packages/Microsoft.Extensions.AI.OpenAI/10.9.0" rel="noopener noreferrer"&gt;&lt;code&gt;Microsoft.Extensions.AI.OpenAI&lt;/code&gt; 10.9.0&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;OpenAI: &lt;a href="https://developers.openai.com/api/docs/guides/embeddings" rel="noopener noreferrer"&gt;Embeddings guide&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Qdrant: &lt;a href="https://qdrant.tech/documentation/manage-data/collections/" rel="noopener noreferrer"&gt;Collections&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Qdrant: &lt;a href="https://qdrant.tech/documentation/tutorials-operations/embedding-model-migration/" rel="noopener noreferrer"&gt;Migrate to a new embedding model&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>dotnet</category>
      <category>csharp</category>
      <category>rag</category>
      <category>ai</category>
    </item>
    <item>
      <title>What Happens When Your AI Feature Fails?</title>
      <dc:creator>Lukas Walter </dc:creator>
      <pubDate>Wed, 19 Aug 2026 15:30:00 +0000</pubDate>
      <link>https://dev.to/lukaswalter/what-happens-when-your-ai-feature-fails-4d5e</link>
      <guid>https://dev.to/lukaswalter/what-happens-when-your-ai-feature-fails-4d5e</guid>
      <description>&lt;p&gt;The first design pass for an AI feature should describe how it fails.&lt;/p&gt;

&lt;p&gt;That can feel backwards when the feature does not work yet. But a timeout is already part of the design. So are a retrieval miss, malformed output, and a write that may have completed after the caller gave up waiting.&lt;/p&gt;

&lt;p&gt;Start with one user-visible flow. At every boundary, ask what can go wrong, what may already have happened, what the application can still promise, and how anyone will know which case occurred.&lt;/p&gt;

&lt;p&gt;Keep the answers in a failure-mode table and use it to shape the API contract, control flow, telemetry, and tests. Writing that table after the first incident is rather late.&lt;/p&gt;

&lt;h2&gt;
  
  
  The happy path hides product decisions
&lt;/h2&gt;

&lt;p&gt;Consider a support feature that prepares a reply and optionally saves it as a draft:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;resolve ticket and evaluate access
    -&amp;gt; retrieve authorized ticket content and policy context
    -&amp;gt; call the model
    -&amp;gt; validate the generated reply
    -&amp;gt; save the draft when requested
    -&amp;gt; return the result
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The sequence is easy to understand. It also leaves almost every difficult question unanswered.&lt;/p&gt;

&lt;p&gt;What does the user see when retrieval returns no policy documents? Does the model get a chance to answer anyway? If output validation fails, can the application repair it? If saving times out, is the draft absent, stored, or still being processed? Can a retry create a second draft? Which failure should wake an operator?&lt;/p&gt;

&lt;p&gt;Those are part of the feature contract. If the team postpones them until implementation, the answers tend to become whatever falls out of an exception handler or SDK default.&lt;/p&gt;

&lt;p&gt;I would rather make the awkward cases visible first. Once those decisions are written down, implementing the successful case tends to be the easy part.&lt;/p&gt;

&lt;h2&gt;
  
  
  Begin with the promise and the unacceptable outcomes
&lt;/h2&gt;

&lt;p&gt;Before listing infrastructure failures, write down what the user believes the operation does.&lt;/p&gt;

&lt;p&gt;For the support feature, the promise might be:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Prepare a reply from the ticket and approved support policy. Save it as a draft only when the user asks.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Now write the outcomes the system must not present as success:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a reply based on another tenant's data&lt;/li&gt;
&lt;li&gt;a policy claim with no approved source&lt;/li&gt;
&lt;li&gt;a valid-looking reply that failed the output contract&lt;/li&gt;
&lt;li&gt;a "saved" result when persistence was never confirmed&lt;/li&gt;
&lt;li&gt;a duplicate draft created by retrying an ambiguous write&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This short list changes the design. Missing retrieval context cannot silently become a normal model request. A model response cannot become a persisted draft before it passes the output contract. Save success requires a confirmed write, not the absence of an exception.&lt;/p&gt;

&lt;p&gt;The unacceptable outcomes matter more than a generic goal such as "handle errors gracefully." They state which properties the application must preserve when part of the flow fails.&lt;/p&gt;

&lt;p&gt;Strictly speaking, not every row in this exercise is a system failure. An authorization denial can be correct behavior. &lt;code&gt;NoEvidence&lt;/code&gt; can be a valid domain outcome. I include those adverse outcomes because they can still prevent the feature from keeping its promise, and the application needs an explicit response for them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start from the interaction map
&lt;/h2&gt;

&lt;p&gt;If you already &lt;a href="https://www.lukaswalter.dev/posts/the-model-is-only-one-dependency/" rel="noopener noreferrer"&gt;mapped the dependencies around the model&lt;/a&gt;, use that interaction map as the input. Otherwise, sketch the important calls and state changes for this one flow. Another inventory of Azure services, SDK clients, and databases will not tell you how the feature should behave.&lt;/p&gt;

&lt;p&gt;A service can fail differently in different interactions. Reading policy context may permit a reduced result. Authorizing access does not. A database read and a draft write may use the same database but have different consequences when their outcomes are unknown.&lt;/p&gt;

&lt;p&gt;Take one interaction at a time and ask:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;How can this interaction be slow, unavailable, wrong, stale, unauthorized, duplicated, or ambiguous?&lt;/li&gt;
&lt;li&gt;What is the effect on the user-visible operation?&lt;/li&gt;
&lt;li&gt;Could data or external state already have changed?&lt;/li&gt;
&lt;li&gt;What response preserves the feature's promise?&lt;/li&gt;
&lt;li&gt;Which signal distinguishes this case in production?&lt;/li&gt;
&lt;li&gt;How will a test force it to happen?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Use the category list to jog your memory; stop when more rows would describe the same effect and response. A model call can return valid JSON with an unsupported claim. A write can succeed while its acknowledgement is lost. Caller cancellation can arrive after work started. These cases are more useful than another row that simply says "dependency unavailable".&lt;/p&gt;

&lt;h2&gt;
  
  
  Build a table that can change the code
&lt;/h2&gt;

&lt;p&gt;For the support flow, a first pass could look like this:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Interaction and adverse condition&lt;/th&gt;
&lt;th&gt;Effect&lt;/th&gt;
&lt;th&gt;Side-effect state uncertainty&lt;/th&gt;
&lt;th&gt;Application response&lt;/th&gt;
&lt;th&gt;Signal&lt;/th&gt;
&lt;th&gt;Test&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Ticket authorization denies access&lt;/td&gt;
&lt;td&gt;Protected ticket content must not be disclosed or used beyond what is necessary for the authorization decision&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;Return a generic not-found or forbidden result according to application policy&lt;/td&gt;
&lt;td&gt;Authorization outcome and controlled resource identifier&lt;/td&gt;
&lt;td&gt;Denied principal&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Policy retrieval returns no approved context&lt;/td&gt;
&lt;td&gt;A grounded policy answer cannot be produced&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;Return &lt;code&gt;NoEvidence&lt;/code&gt;; do not ask the model to fill the gap&lt;/td&gt;
&lt;td&gt;Retrieval outcome, filters, result count&lt;/td&gt;
&lt;td&gt;Empty retrieval result&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Policy retrieval returns stale context&lt;/td&gt;
&lt;td&gt;The reply may use obsolete rules&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;Reject the context or mark the feature temporarily unavailable&lt;/td&gt;
&lt;td&gt;Source version and age, without document contents&lt;/td&gt;
&lt;td&gt;Expired test document&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model call exceeds its time budget&lt;/td&gt;
&lt;td&gt;No reply is available inside the request budget&lt;/td&gt;
&lt;td&gt;The provider may still be processing, but no application state changed&lt;/td&gt;
&lt;td&gt;Stop waiting, request cancellation when supported, and return &lt;code&gt;TimedOut&lt;/code&gt;; retry behavior is decided separately&lt;/td&gt;
&lt;td&gt;Attempt, elapsed time, cancellation reason&lt;/td&gt;
&lt;td&gt;Delayed fake client&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model returns malformed structured output&lt;/td&gt;
&lt;td&gt;The reply cannot be validated&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;Reject it or make one bounded repair attempt when the contract permits&lt;/td&gt;
&lt;td&gt;Schema version and validation category&lt;/td&gt;
&lt;td&gt;Invalid JSON and missing fields&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model cites a source outside the retrieved set&lt;/td&gt;
&lt;td&gt;The reply is unsupported&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;Reject the output&lt;/td&gt;
&lt;td&gt;Returned source IDs and validation outcome&lt;/td&gt;
&lt;td&gt;Unknown source ID&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Draft save times out after submission&lt;/td&gt;
&lt;td&gt;The application cannot confirm whether the draft committed, so it cannot report &lt;code&gt;Saved&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Draft may already exist&lt;/td&gt;
&lt;td&gt;Return &lt;code&gt;Unconfirmed&lt;/code&gt; with the operation ID created before submission; query or reconcile that identity before another write&lt;/td&gt;
&lt;td&gt;Operation ID, attempt, last known state&lt;/td&gt;
&lt;td&gt;Commit succeeds, response is dropped&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best-effort observability export fails&lt;/td&gt;
&lt;td&gt;Diagnosis becomes harder&lt;/td&gt;
&lt;td&gt;Business state is unchanged&lt;/td&gt;
&lt;td&gt;Continue and record the exporter failure locally when possible&lt;/td&gt;
&lt;td&gt;Exporter health and dropped-item count&lt;/td&gt;
&lt;td&gt;Disabled collector&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Required audit or security record cannot be durably written&lt;/td&gt;
&lt;td&gt;The feature cannot satisfy its operating obligation&lt;/td&gt;
&lt;td&gt;A protected action may already have occurred if recording is not atomic&lt;/td&gt;
&lt;td&gt;Apply the audit policy. When the business state and record share a transaction boundary, commit them together. Otherwise prevent the action where possible or reconcile an ambiguous outcome&lt;/td&gt;
&lt;td&gt;Audit-write outcome, policy decision, and operation ID where applicable&lt;/td&gt;
&lt;td&gt;Rejected audit write before and after the protected action&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If the cells do not affect implementation, the table is busywork. "Log the error and retry" leaves open whether retrying is safe, what the caller receives, and when the operation stops.&lt;/p&gt;

&lt;p&gt;An external effect and a local audit record do not share a transaction boundary. Record durable intent before the effect. Afterward, reconcile and record the final outcome. Depending on the operation, that recovery path may also need an idempotency key or compensation.&lt;/p&gt;

&lt;p&gt;Do not try to enumerate every exception type. Group failures when they have the same effect and response. Split them when the system must behave differently.&lt;/p&gt;

&lt;h2&gt;
  
  
  Separate failure, effect, and response
&lt;/h2&gt;

&lt;p&gt;Teams often jump from a technical symptom to a resilience mechanism:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;timeout -&amp;gt; retry
invalid output -&amp;gt; retry
dependency unavailable -&amp;gt; fallback
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That skips the decision that matters: what did the failure do to the operation?&lt;/p&gt;

&lt;p&gt;A timeout on an idempotent policy read is not the same as a timeout after a draft write was submitted. Both may present as a timeout at the application boundary. The read has no side effect and might be attempted again within the remaining budget. The write has an ambiguous outcome. Retrying it with a new identity may create a duplicate.&lt;/p&gt;

&lt;p&gt;Choose the response from the effect and the known state. The exception name is only one input.&lt;/p&gt;

&lt;p&gt;For each row, I use one of a small set of response shapes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;stop and return a specific application outcome&lt;/li&gt;
&lt;li&gt;continue with an explicitly reduced capability&lt;/li&gt;
&lt;li&gt;make another bounded attempt when the operation is safe and time remains&lt;/li&gt;
&lt;li&gt;move unresolved work to a durable recovery path&lt;/li&gt;
&lt;li&gt;require human review before a consequential action&lt;/li&gt;
&lt;li&gt;fail closed because authorization, integrity, or policy cannot be proven&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The exact set belongs to the application. Callers should receive application outcomes without having to interpret exception text.&lt;/p&gt;

&lt;h2&gt;
  
  
  Put the outcomes in the application contract
&lt;/h2&gt;

&lt;p&gt;Once the failure table stabilizes, encode the outcomes that the endpoint or UI needs to handle.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;enum&lt;/span&gt; &lt;span class="n"&gt;ReplyOutcome&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;Prepared&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;NoEvidence&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;InvalidGeneration&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;TimedOut&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;TemporarilyUnavailable&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;enum&lt;/span&gt; &lt;span class="n"&gt;DraftSaveOutcome&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;NotRequested&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;Saved&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;RejectedByPolicy&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;Failed&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;Unconfirmed&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;sealed&lt;/span&gt; &lt;span class="k"&gt;record&lt;/span&gt; &lt;span class="nc"&gt;PrepareReplyResult&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;ReplyOutcome&lt;/span&gt; &lt;span class="n"&gt;ReplyOutcome&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="n"&gt;Reply&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;DraftSaveOutcome&lt;/span&gt; &lt;span class="n"&gt;SaveOutcome&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;Guid&lt;/span&gt;&lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="n"&gt;SaveOperationId&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;RejectedByPolicy&lt;/code&gt; means the application deliberately refused the save before persistence. &lt;code&gt;Failed&lt;/code&gt; means the application knows that persistence did not commit. &lt;code&gt;Unconfirmed&lt;/code&gt; means the write may have committed, but the application has not verified the result.&lt;/p&gt;

&lt;p&gt;Allocate the operation ID before the write:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="n"&gt;Guid&lt;/span&gt; &lt;span class="n"&gt;operationId&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Guid&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;NewGuid&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="kt"&gt;var&lt;/span&gt; &lt;span class="n"&gt;command&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;SaveDraftCommand&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;OperationId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;operationId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;TicketId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;ticketId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;Reply&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;reply&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="n"&gt;DraftSaveResult&lt;/span&gt; &lt;span class="n"&gt;saveResult&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;draftStore&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;SaveAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;command&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;cancellationToken&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The store must persist &lt;code&gt;OperationId&lt;/code&gt; with the draft or an operation record and enforce uniqueness. Reusing the ID for a different command must fail. Persist enough command identity, such as the ticket ID and a canonical command fingerprint, to detect conflicting reuse. If the response disappears after commit, the application queries that identity, and any safe retry reuses it. Otherwise, the ID gives a worker nothing authoritative to reconcile.&lt;/p&gt;

&lt;p&gt;The compact result type still permits invalid combinations. Production code may use factory methods or a result hierarchy to prevent them. Even so, it makes two decisions explicit: reply generation and draft persistence have separate outcomes, and an unconfirmed save is not reported as either success or failure.&lt;/p&gt;

&lt;p&gt;The endpoint can now map those outcomes deliberately. The UI can say that a reply was prepared but its save status is still being checked. A background worker can reconcile the same operation ID. Telemetry can record stable outcome names instead of provider-specific exception messages.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prioritize without pretending the numbers are precise
&lt;/h2&gt;

&lt;p&gt;A full failure-mode analysis can grow quickly. Do not give every row equal attention.&lt;/p&gt;

&lt;p&gt;In a review, I start with three questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;How much harm can this cause to users, data, security, or an operational promise?&lt;/li&gt;
&lt;li&gt;How plausible is the case in this design and environment?&lt;/li&gt;
&lt;li&gt;How likely are we to detect it before the user or an operator has to report it?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I prefer simple priority labels such as critical, high, normal, and low over multiplying guessed scores into a number that looks scientific. Unknown write outcomes, cross-tenant data exposure, silent use of stale policy, and failures that look like success deserve attention before a clean provider error that the application already exposes honestly.&lt;/p&gt;

&lt;p&gt;Keep the treatment decision separate: mitigate, accept, or defer. A high-priority risk can still be accepted when the team understands the consequence and decides that mitigation is not justified. "We do not support draft recovery in the first release" can be a legitimate decision if the UI never claims an unconfirmed write succeeded and the consequence is acceptable. An undocumented gap is not the same thing as an accepted risk.&lt;/p&gt;

&lt;h2&gt;
  
  
  Turn every important row into a forced failure
&lt;/h2&gt;

&lt;p&gt;For every important row, the team should be able to force or simulate the relevant condition and effect. Some provider and infrastructure failures cannot be reproduced exactly, but their application-visible behavior usually can.&lt;/p&gt;

&lt;p&gt;You do not need a production-scale chaos platform for the first pass. Use fake clients and controlled test doubles to return no context, delay a model call, produce malformed output, reject authorization, or lose a write acknowledgement after commit.&lt;/p&gt;

&lt;p&gt;For each high-priority row, verify four things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The application returns the intended outcome.&lt;/li&gt;
&lt;li&gt;It preserves the required data and authorization properties.&lt;/li&gt;
&lt;li&gt;It emits enough context to identify the interaction and failure category without leaking sensitive content.&lt;/li&gt;
&lt;li&gt;Any retry or recovery path preserves side-effect safety.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Keep at least the deterministic cases in the normal test suite. Run slower dependency and recovery drills on a schedule that the team will actually maintain.&lt;/p&gt;

&lt;p&gt;The same rows guide the code and tests. During an incident, they also tell the operator which behavior was intentional.&lt;/p&gt;

&lt;h2&gt;
  
  
  Give retries and runtime budgets their own pass
&lt;/h2&gt;

&lt;p&gt;The analysis should identify where another attempt may be useful. Decide retry behavior holistically across the operation afterward. That does not mean retrying the whole workflow.&lt;/p&gt;

&lt;p&gt;Whether an attempt is safe depends on idempotency, the failure class, remaining time, provider throttling, and the scope of any side effect. A retry can be a valid response to one row and make the next row much worse.&lt;/p&gt;

&lt;p&gt;A five-second model timeout tells you little on its own. It has to fit inside the request budget for retrieval, model calls, validation, tools, persistence, and any synchronous recovery attempts. Durable recovery that outlives the request needs its own deadline or operational objective.&lt;/p&gt;

&lt;p&gt;During the failure pass, note which rows need a retry or budget decision. Resist inventing a local retry count or timeout just to fill the cell.&lt;/p&gt;

&lt;h2&gt;
  
  
  When to use this approach
&lt;/h2&gt;

&lt;p&gt;Use a failure-first pass when an AI feature reads protected data, relies on retrieval, makes authoritative claims, invokes tools, changes state, or creates a meaningful promise about latency and availability. It is also useful before changing a prompt, model, or provider when that change can alter output contracts or runtime behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  When a lighter pass is enough
&lt;/h2&gt;

&lt;p&gt;A disposable local experiment with synthetic data and no side effects does not need a workshop or a large register. Write down the few cases that could invalidate the experiment, keep the error visible, and move on.&lt;/p&gt;

&lt;p&gt;Keep the method proportional to the feature. A small experiment needs a short list, not a reliability program.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practical takeaway
&lt;/h2&gt;

&lt;p&gt;Choose one important AI flow and spend 45 minutes on its unhappy paths before adding more happy-path code.&lt;/p&gt;

&lt;p&gt;Write the user promise and the outcomes that must never appear as success. Walk each interaction, record the effect of plausible failures, and assign an application response, a diagnostic signal, and a way to force the case in a test.&lt;/p&gt;

&lt;p&gt;If a row has no defined response or cannot be tested, it is unfinished design work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/posts/calling-a-model-is-easy-running-an-ai-system-is-not/" rel="noopener noreferrer"&gt;Calling a Model Is Easy. Running an AI System Is Not&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/posts/the-model-is-only-one-dependency/" rel="noopener noreferrer"&gt;The Model Is Only One Dependency. Map the Rest.&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.lukaswalter.dev/posts/when-ai-request-becomes-background-job/" rel="noopener noreferrer"&gt;When an AI Request Should Become a Background Job&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/azure/well-architected/reliability/failure-mode-analysis" rel="noopener noreferrer"&gt;Architecture strategies for performing failure mode analysis&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/azure/architecture/best-practices/transient-faults" rel="noopener noreferrer"&gt;Transient fault handling&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/ef/core/miscellaneous/connection-resiliency" rel="noopener noreferrer"&gt;Connection resiliency in Entity Framework Core&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft Learn: &lt;a href="https://learn.microsoft.com/en-us/azure/well-architected/reliability/reliability-test" rel="noopener noreferrer"&gt;Architecture strategies for designing a reliability testing strategy&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>architecture</category>
    </item>
  </channel>
</rss>
