<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Manish Goud</title>
    <description>The latest articles on DEV Community by Manish Goud (@manish_goud_789).</description>
    <link>https://dev.to/manish_goud_789</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4149936%2F79ed655b-7a8a-4e5e-826e-9956e7f8fdff.jpg</url>
      <title>DEV Community: Manish Goud</title>
      <link>https://dev.to/manish_goud_789</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/manish_goud_789"/>
    <language>en</language>
    <item>
      <title>"Sub-Second Clinical Briefings: FastAPI, Async Recall and Fast Inference "</title>
      <dc:creator>Manish Goud</dc:creator>
      <pubDate>Tue, 29 Sep 2026 14:53:50 +0000</pubDate>
      <link>https://dev.to/manish_goud_789/sub-second-clinical-briefings-fastapi-async-recall-and-fast-inference--2h57</link>
      <guid>https://dev.to/manish_goud_789/sub-second-clinical-briefings-fastapi-async-recall-and-fast-inference--2h57</guid>
      <description>&lt;p&gt;A therapist has a few minutes between sessions. A briefing that takes eight seconds might as well not exist.&lt;/p&gt;

&lt;p&gt;When you pair a memory layer with fast LLM inference, most of the remaining latency is plumbing. Here is how to keep the request path short.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep the request path thin
&lt;/h2&gt;

&lt;p&gt;Ingestion (writing notes) and synthesis (answering a clinician) have different needs. Heavy work like pattern reflection runs in the background, so the query endpoint only does recall plus one generation call.&lt;/p&gt;

&lt;h2&gt;
  
  
  Go fully async
&lt;/h2&gt;

&lt;p&gt;A naive endpoint blocks on each network call. An async client lets the server handle other requests while it waits:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;httpx&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;fastapi&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;FastAPI&lt;/span&gt;

&lt;span class="n"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;FastAPI&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;httpx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;AsyncClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;5.0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nd"&gt;@app.post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/api/copilot-ask&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;copilot_ask&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;CopilotQuery&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;recall&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="nf"&gt;recall_url&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;child_id&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;query&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;HEADERS&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;context&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;recall&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;recall&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;is_success&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;generate_briefing&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Trim the prompt
&lt;/h2&gt;

&lt;p&gt;Latency and cost scale with tokens. Recall only what's relevant, cap the number of memories you inject, and ask for a short output. One or two bullets is plenty for a pre-session briefing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Plan for failure
&lt;/h2&gt;

&lt;p&gt;Set timeouts, and decide what the UI shows if recall or inference is slow. A clear "briefing unavailable, showing last known protocol" beats a spinner.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measure, don't assume
&lt;/h2&gt;

&lt;p&gt;Log recall time and generation time separately, and look at p95, not just the average. Use your own numbers as the target.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Separate the write path from the read path.&lt;/li&gt;
&lt;li&gt;Async I/O and small prompts do most of the work.&lt;/li&gt;
&lt;li&gt;Timeouts and fallbacks are part of the product, not an afterthought.&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>python</category>
      <category>fastapi</category>
      <category>ai</category>
      <category>performance</category>
    </item>
  </channel>
</rss>
