<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ryan Fernandes</title>
    <description>The latest articles on DEV Community by Ryan Fernandes (@ryan_fernandes_fc52e09ac9).</description>
    <link>https://dev.to/ryan_fernandes_fc52e09ac9</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4029084%2F06dc34b4-da18-49c0-b269-7624168611b9.png</url>
      <title>DEV Community: Ryan Fernandes</title>
      <link>https://dev.to/ryan_fernandes_fc52e09ac9</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ryan_fernandes_fc52e09ac9"/>
    <language>en</language>
    <item>
      <title>Understanding Connected and Automated Vehicles (CAV) Part-1</title>
      <dc:creator>Ryan Fernandes</dc:creator>
      <pubDate>Tue, 11 Aug 2026 20:29:50 +0000</pubDate>
      <link>https://dev.to/ryan_fernandes_fc52e09ac9/understanding-connected-and-automated-vehicles-cav-part-1-2n0</link>
      <guid>https://dev.to/ryan_fernandes_fc52e09ac9/understanding-connected-and-automated-vehicles-cav-part-1-2n0</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhazvb7fo7y3nvvkuvm3r.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhazvb7fo7y3nvvkuvm3r.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  So What's This V2V Thing Anyway?
&lt;/h2&gt;

&lt;p&gt;Let's start with the basics. V2V communication (Vehicle-to-Vehicle), for the uninitiated, is exactly what it sounds like: cars talking to each other. But here's the kicker, they're not just chatting about the weather. They're sharing what they perceive through their sensors and ECUs and adjusting their behavior based on what other cars are telling them.&lt;/p&gt;

&lt;p&gt;Sounds simple enough, right? But the implications are absolutely wild.&lt;/p&gt;




&lt;h2&gt;
  
  
  Seven Ways This Technology Could Completely Change Driving
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. The "Phantom Traffic Jam" Problem
&lt;/h3&gt;

&lt;p&gt;You know those traffic jams that seem to appear out of thin air? One person taps their brakes, the person behind taps harder, and two miles back, everyone's at a standstill for absolutely no reason. It's maddening.&lt;/p&gt;

&lt;p&gt;Now imagine this instead: every car shares its exact acceleration and deceleration ten times per second. The cars behind can smooth out their speed before they even reach the brake wave.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Result?&lt;/strong&gt; Forty percent less stop-and-go traffic. No new roads needed. Just smarter cars.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Intersection "Tunnels" (Say Goodbye to Red Lights)
&lt;/h3&gt;

&lt;p&gt;Here's a mind-bender: what if you never had to stop at a red light again?&lt;/p&gt;

&lt;p&gt;The system works like this: a central server calculates the exact speed and arrival time of every car approaching an intersection. It assigns each vehicle a "time slot" to pass through the middle. Cars simply slow down or speed up slightly to hit their slot.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Result?&lt;/strong&gt; Nobody stops. Fuel economy skyrockets. Your commute becomes a smooth, continuous flow rather than a series of frustrating stops.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Emergency Vehicle Preemption (Beyond Sirens)
&lt;/h3&gt;

&lt;p&gt;Right now, ambulances blare sirens and hope people move out of the way. It's chaotic, unpredictable, and frankly, not good enough.&lt;/p&gt;

&lt;p&gt;With V2V, an ambulance tells the server its route. The server tells every car within a mile to "pull over and stop" five minutes before the ambulance arrives.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Result?&lt;/strong&gt; A perfect, empty corridor cleared in advance. No panic, no confusion, just a clear path for emergency vehicles.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. "See-Through" Trucks
&lt;/h3&gt;

&lt;p&gt;We've all been there, stuck behind a massive semi-truck, completely blind to what's ahead. Is there a pedestrian? A stalled car? An accident?&lt;/p&gt;

&lt;p&gt;With V2V, the truck's front camera broadcasts to your car's screen behind it. You effectively "see through" the truck and react to hazards two seconds earlier than you otherwise could.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Result?&lt;/strong&gt; Two seconds doesn't sound like much, but at highway speeds, that's the difference between stopping in time and a collision.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Pothole Mapping &amp;amp; Predictive Suspension
&lt;/h3&gt;

&lt;p&gt;Every car has accelerometers. When fifty cars hit the same pothole and jolt, the server logs the GPS coordinate instantly. It then broadcasts to the next thousand cars: "Slow down five miles per hour" or "Avoid the right lane."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Result?&lt;/strong&gt; You never feel the bump. Your suspension lasts longer. Your wheels stay aligned.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. "Green Wave" for Semi-Trucks
&lt;/h3&gt;

&lt;p&gt;Heavy trucks waste enormous fuel accelerating uphill. But what if the server knew the topography and could tell the truck: "Speed up to sixty-five now, because in two miles there's a steep hill; you'll coast over without downshifting."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Result?&lt;/strong&gt; Fifteen percent fuel savings per truck. That's massive for both operational costs and environmental impact.&lt;/p&gt;

&lt;h3&gt;
  
  
  7. Post-Crash Autonomous Safe-Off
&lt;/h3&gt;

&lt;p&gt;If Car A's airbags deploy, the server immediately tells the five cars behind Car A to automatically steer to the shoulder and stop.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Result?&lt;/strong&gt; No secondary pile-up. No waiting for human reaction time. Just automatic, life-saving prevention.&lt;/p&gt;




&lt;h2&gt;
  
  
  Let's Narrow Our Focus
&lt;/h2&gt;

&lt;p&gt;There are clearly tons of problems this technology could solve. But for our purposes, we'll pick one and build on top of it. The rest will naturally fall into place.&lt;/p&gt;

&lt;p&gt;We're going with &lt;strong&gt;phantom traffic jams&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Researchers at the University of Michigan actually put connected and autonomous vehicles into a convoy and demonstrated that a single autonomous vehicle could dampen traffic waves. The connected vehicle received information about vehicles farther ahead and braked more smoothly. The experiment reported energy savings of up to 19% for the connected vehicle and 7% for following human-driven vehicles.&lt;/p&gt;

&lt;p&gt;That's real-world proof that this works.&lt;/p&gt;




&lt;h2&gt;
  
  
  How Do We Actually Make Cars Talk to Each Other?
&lt;/h2&gt;

&lt;p&gt;Here's a crucial distinction: V2V is &lt;strong&gt;not&lt;/strong&gt; cars talking over the internet. For safety-critical communication, the important mechanism is direct wireless sidelink communication between nearby vehicles. No cell towers. No Wi-Fi routers. Just cars talking directly to each other.&lt;/p&gt;

&lt;h3&gt;
  
  
  What Does the Car Actually Send?
&lt;/h3&gt;

&lt;p&gt;For basic cooperative awareness, a car periodically broadcasts information about itself:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Position&lt;/li&gt;
&lt;li&gt;Speed&lt;/li&gt;
&lt;li&gt;Heading&lt;/li&gt;
&lt;li&gt;Acceleration&lt;/li&gt;
&lt;li&gt;Braking state&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In Europe, this is called a CAM (Cooperative Awareness Message). In the US ecosystem, it's a BSM (Basic Safety Message). Transmission typically happens at 1–10 Hz, depending on the vehicle's state and channel conditions.&lt;/p&gt;

&lt;h3&gt;
  
  
  But What Radio Technology Is Used?
&lt;/h3&gt;

&lt;p&gt;Historically, there have been two major families:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A. Wi-Fi-derived V2X&lt;/strong&gt; (IEEE 802.11p → DSRC/ITS-G5 → IEEE 802.11bd)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;B. Cellular-derived V2X&lt;/strong&gt; (LTE-V2X → 5G NR-V2X)&lt;/p&gt;

&lt;p&gt;Here's the thing: both 802.11bd and 5G NR-V2X send the same safety alerts directly from vehicle to vehicle with ultra-low latency. The only real difference is who built the underlying wireless system: the Wi-Fi committee (IEEE) or the Cellular committee (3GPP).&lt;/p&gt;

&lt;p&gt;The newer generations use &lt;strong&gt;5G NR-V2X&lt;/strong&gt;, which introduces more advanced capabilities.&lt;/p&gt;




&lt;h2&gt;
  
  
  Choosing the Communication Layer for V2V
&lt;/h2&gt;

&lt;p&gt;Before deciding how vehicles should communicate, we need to define what an ideal V2V system should provide. This is fundamentally different from a conventional internet application, where a vehicle may need to react to information from another vehicle within milliseconds, while hundreds of vehicles may be communicating simultaneously in the same area.&lt;/p&gt;

&lt;h3&gt;
  
  
  What Makes an Ideal V2V System?
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;1. Low Latency&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Safety-critical information has a short useful lifetime. If a vehicle suddenly brakes, a warning received hundreds of milliseconds later is significantly less useful than one received immediately.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. High Reliability&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Wireless channels are affected by interference, fading, obstacles, and vehicle density. For safety-critical applications, the goal isn't simply high average throughput, it's a very high probability that important messages are delivered within their required time window.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Predictable Latency&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Average latency alone isn't enough. A system with a 5 ms average but occasional 500 ms delays could be dangerous. Latency variation, reliability, and worst-case behavior matter just as much as the average.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Support for High Vehicle Density&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A highway or urban intersection may contain hundreds of vehicles within communication range. If every vehicle continuously broadcasts information, the wireless channel can become congested. The system needs mechanisms for efficient resource utilization and congestion management.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. High Mobility Support&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Vehicles move at highway speeds and rapidly change their relative position. The system needs to operate reliably under high Doppler shifts, rapidly changing channels, and constantly changing network topology.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;6. Direct Vehicle-to-Vehicle Communication&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Safety-critical communication shouldn't depend on an internet connection or a remote cloud server. The critical path should be&lt;/p&gt;

&lt;p&gt;Vehicle A brakes → Direct V2V → Vehicle B receives warning → React&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Not:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Vehicle A → Cellular network → Cloud → Cellular network → Vehicle B&lt;/p&gt;

&lt;p&gt;The latter introduces dependencies on network coverage, congestion, and infrastructure availability.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;7. Broadcast, Groupcast, and Unicast Capabilities&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If a vehicle detects a crash, it needs to warn every vehicle approaching from behind, not just one specific vehicle. Other applications, like coordinated maneuvers between two vehicles, benefit from more targeted communication.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;8. Sufficient Bandwidth for Future Applications&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Basic V2V messages are relatively small: position, speed, heading, acceleration, and braking state. But future systems will exchange richer information: detected objects, trajectories, and sensor-derived data. We're moving from:&lt;/p&gt;

&lt;p&gt;"I am braking."&lt;/p&gt;

&lt;p&gt;to:&lt;/p&gt;

&lt;p&gt;"I detected a pedestrian behind the truck at this position, moving at this velocity."&lt;/p&gt;

&lt;p&gt;This is cooperative perception, and it places substantially greater demands on the communication system.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;9. Security&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Vehicles must determine whether a message actually originated from a trusted participant. Without authentication and integrity mechanisms, an attacker could inject false information: "Accident ahead." "Emergency vehicle approaching." "The vehicle ahead is braking." Security isn't optional; it's fundamental.&lt;/p&gt;




&lt;h2&gt;
  
  
  Evaluating the Two Major Approaches
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Wi-Fi-derived V2X
&lt;/h3&gt;

&lt;p&gt;The evolution path:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;IEEE 802.11p&lt;/li&gt;
&lt;li&gt;DSRC / ITS-G5&lt;/li&gt;
&lt;li&gt;IEEE 802.11bd&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are derived from the IEEE 802.11 family and were specifically adapted for vehicular communication.&lt;/p&gt;

&lt;h3&gt;
  
  
  Cellular-derived V2X
&lt;/h3&gt;

&lt;p&gt;The evolution path:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;LTE-V2X&lt;/li&gt;
&lt;li&gt;5G NR-V2X&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;5G NR-V2X is based on 3GPP's cellular technology and introduces more advanced capabilities for demanding V2X applications.&lt;/p&gt;

&lt;p&gt;Here's an important point: &lt;strong&gt;5G NR-V2X does not mean every vehicle-to-vehicle message has to travel through a 5G cellular tower.&lt;/strong&gt; NR-V2X supports sidelink communication (through the PC5 interface), allowing vehicles to communicate directly:&lt;/p&gt;

&lt;p&gt;Vehicle A ↔ NR-V2X sidelink ↔ Vehicle B&lt;/p&gt;

&lt;p&gt;This makes it suitable for the low-latency, local communication required by safety-critical V2V applications.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Choose 5G NR-V2X?
&lt;/h2&gt;

&lt;p&gt;The choice isn't simply because "5G is faster." For basic safety messages, both 802.11bd and NR-V2X can provide suitable communication. A message containing position, speed, acceleration, and braking state doesn't require enormous bandwidth.&lt;/p&gt;

&lt;p&gt;The distinction becomes more important as we move toward cooperative driving and cooperative perception.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;NR-V2X was designed to extend V2X beyond basic awareness messages&lt;/strong&gt; toward applications like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Cooperative driving&lt;/li&gt;
&lt;li&gt;Platooning&lt;/li&gt;
&lt;li&gt;Coordinated maneuvers&lt;/li&gt;
&lt;li&gt;Trajectory coordination&lt;/li&gt;
&lt;li&gt;Cooperative perception&lt;/li&gt;
&lt;li&gt;High-density vehicle communication&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It provides mechanisms for different communication patterns and more sophisticated radio-resource management, while supporting direct sidelink communication even when vehicles can't rely on cellular infrastructure.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Resulting Architecture
&lt;/h2&gt;

&lt;p&gt;Rather than designing the system as:&lt;/p&gt;

&lt;p&gt;Vehicle → Central Server → Vehicle&lt;/p&gt;

&lt;p&gt;We use two complementary communication paths:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7gtmye0ap5hlqiiszl85.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7gtmye0ap5hlqiiszl85.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;NR-V2X sidelink&lt;/strong&gt; handles time-critical, local vehicle-to-vehicle cooperation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The cellular network/cloud&lt;/strong&gt; handles information that benefits from a wider view: traffic management, road-condition databases, infrastructure information, and longer-term coordination.&lt;/p&gt;

&lt;p&gt;This separation is crucial: the cloud can provide global intelligence without becoming a single point of failure for every safety-critical V2V interaction.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Decision
&lt;/h2&gt;

&lt;p&gt;Based on these requirements, &lt;strong&gt;5G NR-V2X is the communication technology selected for this project.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This decision is driven by the direction of the system we want to build. If the objective were limited to basic cooperative awareness, sharing speed, position, and braking status, both 802.11bd and NR-V2X would be viable choices.&lt;/p&gt;

&lt;p&gt;But our goal extends beyond basic V2V:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Basic awareness&lt;/li&gt;
&lt;li&gt;Cooperative control&lt;/li&gt;
&lt;li&gt;Coordinated intersections&lt;/li&gt;
&lt;li&gt;Cooperative perception&lt;/li&gt;
&lt;li&gt;Cooperative autonomous driving&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;As the amount and importance of exchanged information increases, the advanced sidelink capabilities, resource-management mechanisms, communication modes, and scalability of 5G NR-V2X become increasingly valuable.&lt;/p&gt;

&lt;p&gt;Therefore, the system will use &lt;strong&gt;NR-V2X PC5 as the primary direct V2V communication layer&lt;/strong&gt;, while cellular connectivity will be used separately for communication with infrastructure and cloud-based services.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Real Challenge
&lt;/h2&gt;

&lt;p&gt;The next engineering challenge is no longer simply:&lt;/p&gt;

&lt;p&gt;"Can two vehicles communicate?"&lt;/p&gt;

&lt;p&gt;It's:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Can hundreds of rapidly moving vehicles exchange time-critical information reliably and predictably in a congested wireless environment, and can that information actually improve safety and traffic efficiency?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's the problem the rest of this project will investigate.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This is the first part of our series on connected and automated vehicles. Next up: how we're building the simulation environment to test these ideas at scale.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>iot</category>
      <category>networking</category>
      <category>automobile</category>
    </item>
    <item>
      <title>Building an Agentic Navigation + Autonomous Execution System</title>
      <dc:creator>Ryan Fernandes</dc:creator>
      <pubDate>Fri, 17 Jul 2026 19:10:30 +0000</pubDate>
      <link>https://dev.to/ryan_fernandes_fc52e09ac9/building-an-agentic-navigation-autonomous-execution-system-12fi</link>
      <guid>https://dev.to/ryan_fernandes_fc52e09ac9/building-an-agentic-navigation-autonomous-execution-system-12fi</guid>
      <description>&lt;p&gt;&lt;strong&gt;Connect with me:&lt;/strong&gt; &lt;a href="https://www.linkedin.com/in/ryan-fernandes-294136231/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; · &lt;a href="https://www.instagram.com/ryan_ferds?igsh=aXVwazYzcWQ1dW43" rel="noopener noreferrer"&gt;Instagram&lt;/a&gt; · &lt;a href="https://github.com/RyanFernandes23" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Hello all, it's nice to have you around.&lt;/p&gt;

&lt;p&gt;In this article, I will be discussing the project that I delivered in 2025.&lt;br&gt;
It is an AI agent that enables users to navigate a web mobile enterprise product and perform CRUD operations using natural language.&lt;/p&gt;

&lt;p&gt;Two partner companies had attempted but failed to develop the same project before I came along. I will walk through the project in stages. I will explain not only what happened but &lt;em&gt;why&lt;/em&gt;, and what alternative paths I could have taken. I will end with changes I would make if I could redesign the project from scratch.&lt;/p&gt;
&lt;h2&gt;
  
  
  Table of Contents
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Background: how this project made its way to me&lt;/li&gt;
&lt;li&gt;System functionality&lt;/li&gt;
&lt;li&gt;High-level architecture&lt;/li&gt;
&lt;li&gt;
Part 1: The navigation sub-agent (RAG-based screen routing)

&lt;ul&gt;
&lt;li&gt;Intent classification – how and why semantic routing&lt;/li&gt;
&lt;li&gt;Why the RAG approach for navigating is useful&lt;/li&gt;
&lt;li&gt;Schema design (why Postgres + pgvector over a separate vector DB)&lt;/li&gt;
&lt;li&gt;Generation embedding – choice of model &amp;amp; why normalization is important&lt;/li&gt;
&lt;li&gt;The reason I eventually switched away from the local model, due to latency&lt;/li&gt;
&lt;li&gt;Ingestion script example&lt;/li&gt;
&lt;li&gt;Retrieval and parameter resolution&lt;/li&gt;
&lt;li&gt;Connection pooling, and why raw connections don't survive production&lt;/li&gt;
&lt;li&gt;Handling expiring and stale connections specifically&lt;/li&gt;
&lt;li&gt;Latency: caching and why 0.8 in particular&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
Part 2: Autonomous Execution Sub-agent (ReAct + MCP)

&lt;ul&gt;
&lt;li&gt;Why ReAct&lt;/li&gt;
&lt;li&gt;Why MCP and not hand-crafted function calling?&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
How to build an MCP tool?

&lt;ul&gt;
&lt;li&gt;What ctx actually is, and how to use it&lt;/li&gt;
&lt;li&gt;Why each tool wraps an existing REST endpoint, and not raw SQL&lt;/li&gt;
&lt;li&gt;Collapsing chained API calls into a single tool&lt;/li&gt;
&lt;li&gt;A backup for wrongly classified queries&lt;/li&gt;
&lt;li&gt;Transient token propagation&lt;/li&gt;
&lt;li&gt;Why streamable HTTP and why co-locate first&lt;/li&gt;
&lt;li&gt;Auth chain end to end, and why each hop verifies independently&lt;/li&gt;
&lt;li&gt;Conversation memory (what a checkpoint actually is)&lt;/li&gt;
&lt;li&gt;Streaming, and why per-node events&lt;/li&gt;
&lt;li&gt;Request/response shape&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;What I'd do differently (retrospective)&lt;/li&gt;
&lt;li&gt;Up next&lt;/li&gt;
&lt;/ul&gt;


&lt;h2&gt;
  
  
  Background: how this project made its way to me
&lt;/h2&gt;

&lt;p&gt;I began working at my startup in 2025. This was mainly due to my past RAG projects:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A Perplexity clone.&lt;/li&gt;
&lt;li&gt;An agent-based text-to-SQL translator.&lt;/li&gt;
&lt;li&gt;A custom vision transformer.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This startup had made several attempts to develop their agent. They employed third parties. It always failed because the agent AI development toolkit wasn't ready yet.&lt;/p&gt;

&lt;p&gt;I received full ownership almost instantly. My first prototype design was too premature. I was provided with someone else's POC document. I was tasked with completing the design right away without discussing the deadline.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Lesson learned:&lt;/strong&gt;&lt;br&gt;
If you're getting a project with an unclear scope, discuss the schedule verbally before starting work. Don't wait until after your first draft fails.&lt;/p&gt;

&lt;p&gt;Everything became easier after I renegotiated this deadline. I connected with a senior developer within the company. He had no background in AI but was proficient in Java Spring Boot. This helped me to get more information about the product. Here's what I've done and for what reason.&lt;/p&gt;


&lt;h2&gt;
  
  
  System functionality
&lt;/h2&gt;

&lt;p&gt;The agent can resolve two different types of user intents:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Navigation (natural language navigation):&lt;/strong&gt;&lt;br&gt;
The user asks for things like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"edit user details"&lt;/li&gt;
&lt;li&gt;"go to main dashboard"&lt;/li&gt;
&lt;li&gt;"onboard a new device"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The agent responds with a widget. This widget contains a button that leads directly to the required screen.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg9a79qhunnqt4db3c4ab.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg9a79qhunnqt4db3c4ab.png" alt="navigation widget example" width="800" height="418"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Action (natural language CRUD):&lt;/strong&gt;&lt;br&gt;
The user asks for things like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"register a new product 'Wesco gas sensor' with [configuration]"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The agent executes it automatically.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd45t8tnapiw5mtn9jv5j.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd45t8tnapiw5mtn9jv5j.png" alt="example of autonomous execution" width="800" height="399"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The previous solution had used a text-to-SQL component for executing tasks. This queried the database directly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The problem with this approach:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Any string input provided by a user and used in constructing a query can be a vulnerability for SQL injection.&lt;/li&gt;
&lt;li&gt;The string can also be used in prompt injection. This can make the LLM generate a query it should not generate.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Text-to-SQL components are truly handy for analytics tools with heavy read workloads. They require human approval of the query beforehand.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;However, using such an approach for automated and unsanctioned writing in a multi-tenant environment is just wrong.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr9ue3r5pk66jqdcw6ytn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr9ue3r5pk66jqdcw6ytn.png" alt=" " width="800" height="428"&gt;&lt;/a&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  High-level architecture
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgkjcoo9oepwelgciqvli.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgkjcoo9oepwelgciqvli.png" alt=" " width="800" height="787"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The Java Spring backend is right in front of everything.&lt;br&gt;
&lt;strong&gt;Its role:&lt;/strong&gt; It serves as a proxy for authentication.&lt;/p&gt;

&lt;p&gt;The FastAPI agent service only serves responses if it receives the request having a service-to-service JWT. This is checked at the JWKS endpoint of the Java Spring backend. Auth for clients does not go through the agent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why use a proxy in between? Why not let clients connect straight to FastAPI?&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It helps in maintaining one place for checking the session and org tenancy.&lt;/li&gt;
&lt;li&gt;It allows the AI layer to be deployed, scaled, or replaced without changing client authentication.&lt;/li&gt;
&lt;/ul&gt;


&lt;h2&gt;
  
  
  Part 1: The navigation sub-agent (RAG-based screen routing)
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Intent classification – how and why semantic routing
&lt;/h3&gt;

&lt;p&gt;Every new query has to go through a classification node. It determines whether it's a navigation intent or an autonomous execution intent.&lt;/p&gt;

&lt;p&gt;There are three practical choices:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;Classify the intent via LLM:&lt;/strong&gt;
Submit the query to a chat model with instructions like "classify into NAVIGATION or AUTONOMOUS_EXECUTION".&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Supervised intent classification:&lt;/strong&gt;
Train a small classifier (logistic regression, tiny transformer fine-tuning, etc.) on annotated examples.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Semantic routing:&lt;/strong&gt;
Calculate the embedding of the input query. Match it to embeddings of reference utterances for each route.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I went with &lt;strong&gt;semantic routing&lt;/strong&gt;. Here is my justification:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;An LLM call introduces around 200 to 800 ms latency. It has a per-request cost for a simple yes/no answer. That's overkill.&lt;/li&gt;
&lt;li&gt;It also introduces nondeterministic behavior at the very beginning of the request path. An LLM sometimes gets confused by its own classifying instructions.&lt;/li&gt;
&lt;li&gt;Classifier training is the correct long-term solution. However, without labeled data at launch time, there is simply nothing to train on.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why semantic routing fills this void:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It's deterministic.&lt;/li&gt;
&lt;li&gt;It requires no training data except a few examples per route.&lt;/li&gt;
&lt;li&gt;It takes single-digit milliseconds to complete because all it does is a vector comparison.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwkjbbyslel4tzlingojz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwkjbbyslel4tzlingojz.png" alt="semantic routing comparison" width="800" height="365"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Machinery-wise:&lt;/strong&gt;&lt;br&gt;
Each route (&lt;code&gt;navigation&lt;/code&gt;, &lt;code&gt;autonomous_execution&lt;/code&gt;) is encoded via a few example utterances. The utterances are encoded just once, during initialization.&lt;br&gt;
At inference time, the query itself is encoded in the same way. The cosine similarity of that encoding to each reference utterance is computed. The route with the maximum (mean) similarity wins. This happens only if it exceeds the minimum required threshold. Otherwise, we consider the query to be ambiguous. I did not implement an ambiguity handler, which is one of the problems I highlight in my retrospective.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;semantic_router&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Route&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;RouteLayer&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;semantic_router.encoders&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;HuggingFaceEncoder&lt;/span&gt;

&lt;span class="n"&gt;navigation_route&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Route&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;navigation&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;utterances&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;take me to the dashboard&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;I want to edit my profile&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;open the onboarding screen&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;auto_execution_route&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Route&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;autonomous_execution&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;utterances&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;register a new product&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;update the user&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s email&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;delete this device&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;encoder&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;HuggingFaceEncoder&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sentence-transformers/all-MiniLM-L6-v2&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;router&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;RouteLayer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;encoder&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;encoder&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;routes&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;navigation_route&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;auto_execution_route&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;

&lt;span class="n"&gt;decision&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;router&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;take me to manage organization page&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="c1"&gt;# decision.name -&amp;gt; "navigation"
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why cosine similarity?&lt;/strong&gt;&lt;br&gt;
Cosine similarity and dot product are interchangeable when working with normalized embeddings. The two metrics are indifferent to the magnitude of the vector. They only care about its direction – that is, what it means.&lt;br&gt;
Euclidean distance is more susceptible to non-semantically significant magnitude differences. It's a less suitable default metric for such comparisons, in my opinion.&lt;/p&gt;

&lt;p&gt;See the &lt;a href="https://github.com/aurelio-labs/semantic-router" rel="noopener noreferrer"&gt;semantic-router documentation&lt;/a&gt; for instructions on how to tune the threshold. This should be done explicitly based on a validation dataset.&lt;/p&gt;
&lt;h3&gt;
  
  
  Why the RAG approach for navigating is useful
&lt;/h3&gt;

&lt;p&gt;The alternatives to RAG in this case:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A massive if/else tree.&lt;/li&gt;
&lt;li&gt;A massive keyword-matching system.&lt;/li&gt;
&lt;li&gt;A one-shot call to the LLM with the entire list of 70+ screens mentioned every single time.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Keyword matching fails when the user makes his request in an unexpected way. That's precisely why you want to use an LLM – to match paraphrases and synonyms like "edit my info", "update user details", "change my profile".&lt;/p&gt;

&lt;p&gt;If you include the list of screens in the prompt every time, it works for 70 screens now, but it does not scale. You burn tokens every time on everything irrelevant. This makes requests longer and increases latency. It also increases the chance that the model gets confused by those screens in the prompt context.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The RAG approach solves both these problems:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Retrieval limits the number of relevant screens to a few &lt;em&gt;before&lt;/em&gt; the LLM even sees them.&lt;/li&gt;
&lt;li&gt;Generation becomes cheap, quick, and reliable.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;
  
  
  Schema design (why Postgres + pgvector over seperate vector DB like chroma, Qdrant, ...)
&lt;/h3&gt;

&lt;p&gt;I chose Postgres with the &lt;code&gt;pgvector&lt;/code&gt; plugin rather than a dedicated vector database like Qdrant or Milvus. It is quite a hard decision, so it's important to be clear on it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;For a dedicated vector DB:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Optimized ANN indexes for searching huge vectors (up to tens of millions) may work better.&lt;/li&gt;
&lt;li&gt;More powerful search primitives are available out-of-the-box.&lt;/li&gt;
&lt;li&gt;Horizontal scalability by design.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;For pgvector:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;One less service to manage and run.&lt;/li&gt;
&lt;li&gt;Consistent transactions between the relational schema and vectors. They reside in the same database and can be updated in one transaction.&lt;/li&gt;
&lt;li&gt;The crucial point: the organization already runs Postgres in production. They did not want to add any new operational complexity.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;With only ~70 screens, none of the dedicated DB advantages made any difference. This is an amount of data small enough that a scan based on cosine distance would be blazingly fast.&lt;/p&gt;

&lt;p&gt;The selection would be different if this were an item search across millions of SKUs. With a limited number of UI screens that grows slowly, pgvector is the safer, cheaper solution.&lt;/p&gt;

&lt;p&gt;Two tables, separated so that relational metadata (infrequently changing) and embeddings (recomputable if you change your embedding model) could be managed separately:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;screens&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;screen_id&lt;/span&gt;       &lt;span class="nb"&gt;SERIAL&lt;/span&gt; &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;screen_name&lt;/span&gt;     &lt;span class="nb"&gt;TEXT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;description&lt;/span&gt;     &lt;span class="nb"&gt;TEXT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;required_params&lt;/span&gt; &lt;span class="n"&gt;JSONB&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;web_url&lt;/span&gt;         &lt;span class="nb"&gt;TEXT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;mobile_path&lt;/span&gt;     &lt;span class="nb"&gt;TEXT&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="n"&gt;EXTENSION&lt;/span&gt; &lt;span class="n"&gt;IF&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;EXISTS&lt;/span&gt; &lt;span class="n"&gt;vector&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;screen_embeddings&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;vector_id&lt;/span&gt;   &lt;span class="nb"&gt;SERIAL&lt;/span&gt; &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;screen_id&lt;/span&gt;   &lt;span class="nb"&gt;INTEGER&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;screens&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;screen_id&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;description&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;embedding&lt;/span&gt;   &lt;span class="n"&gt;VECTOR&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;384&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Below is what the two tables actually look like once populated:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;screens&lt;/code&gt; holding the relational metadata.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;screen_embeddings&lt;/code&gt; holding the vectors tied back to it by &lt;code&gt;screen_id&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkqf3phum0be61wq7jc9h.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkqf3phum0be61wq7jc9h.png" alt="screens and screen_embeddings tables populated after ingestion" width="516" height="733"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;VECTOR(384)&lt;/code&gt; is relevant because it has to be compatible with the dimension of the embedding that your model generates.&lt;br&gt;
&lt;strong&gt;Important note:&lt;/strong&gt; This is a very common reason for a hidden bug. Switching embedding models without changing the vector column causes errors, as a 384 and a 768 dimensional vector cannot be compared.&lt;br&gt;
  &lt;iframe src="https://www.youtube.com/embed/lPTcTh5sRug"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;Data used was stored in a &lt;code&gt;navigation.json&lt;/code&gt; seed file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"screen_name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"profile"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"description"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"User's personal profile page showing account details."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"required_parameters"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"user_id"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"hostname"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"web_url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://{hostname}/profile/{user_id}"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"mobile_path"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"/profile"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;{hostname}&lt;/code&gt; and &lt;code&gt;{user_id}&lt;/code&gt; are important placeholders. This particular product is white-labeled; each client organization will have their own &lt;code&gt;hostname&lt;/code&gt;. Screen URLs need to be generated per request, not during ingest.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why use a template string?&lt;/strong&gt;&lt;br&gt;
Given 70 URLs and a limited number of placeholders, substitution using &lt;code&gt;.format()&lt;/code&gt; is easy to implement, test, and debug. A more structured solution would be required if the logic behind the placeholders was more complex.&lt;/p&gt;
&lt;h3&gt;
  
  
  Generation embedding – choice of model &amp;amp; why normalization is important
&lt;/h3&gt;

&lt;p&gt;The &lt;code&gt;all-MiniLM-L6-v2&lt;/code&gt; model (384-dimensional embeddings) from &lt;code&gt;sentence-transformers&lt;/code&gt; was used.&lt;/p&gt;

&lt;p&gt;In terms of architecture, the model space can be divided into:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Small &amp;amp; fast sentence embedding models&lt;/strong&gt; (MiniLMs, 384 dimensions): Produce good semantic quality on short texts. Run on CPU without problems. Encode in a fraction of a second.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Larger &amp;amp; more accurate models&lt;/strong&gt; (e.g., &lt;code&gt;bge-large&lt;/code&gt;, &lt;code&gt;text-embedding-3-large&lt;/code&gt; from OpenAI): Better semantic understanding of longer or ambiguous texts. Take longer to execute. Are paid per use when using a hosted API. &lt;a href="https://huggingface.co/sentence-transformers" rel="noopener noreferrer"&gt;sentence transformer&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Descriptions of screen elements are short and clear, generated by us. That is exactly the kind of data where a small model will give good semantic quality. Running it locally will save time and money.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;sentence_transformers&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;SentenceTransformer&lt;/span&gt;

&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;SentenceTransformer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;all-MiniLM-L6-v2&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;embed_text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;float&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;normalize_embeddings&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;tolist&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Setting &lt;code&gt;normalize_embeddings=True&lt;/code&gt; is not merely a stylistic preference. It is necessary to make cosine similarity work as expected when using the &lt;code&gt;&amp;lt;=&amp;gt;&lt;/code&gt; operator from the pgvector extension.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why is this important?&lt;/strong&gt;&lt;br&gt;
This process rescales all vectors to unit length. It ensures that the comparison depends only on direction/meaning. It does not depend on how "large" the descriptions are. This is a frequent but subtle mistake since retrieval "works well enough," but sorting deteriorates.&lt;/p&gt;
&lt;h3&gt;
  
  
  The reason I eventually switched away from the local model, due to latency
&lt;/h3&gt;

&lt;p&gt;Using the &lt;code&gt;all-MiniLM-L6-v2&lt;/code&gt; model locally proved to be a significant part of the 5-second latency during navigation. The encoding of one query took about &lt;strong&gt;500ms&lt;/strong&gt;. This is quite a lot for one forward pass through a fairly small model. This does not seem to be an intrinsic property of the model. It is a consequence of running it on standard application-level infrastructure using CPU, without any batching or GPU.&lt;/p&gt;

&lt;p&gt;Rather than tuning the inference pipeline, I moved the embeddings extraction to &lt;strong&gt;Amazon Titan Text Embeddings&lt;/strong&gt; via the Bedrock API. The configured dimension of the embeddings is either 512 or 1024.&lt;/p&gt;

&lt;p&gt;This introduces an extra network hop for the Bedrock API compared to local compute. At first glance, it looks like it will make the whole thing slower. However, a managed embedding service provisioned to serve requests ended up faster than an unoptimized local model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The takeaway:&lt;/strong&gt; "Local" does not necessarily mean "fast" if your local environment is not sufficiently tuned. It's better to measure the actual numbers.&lt;/p&gt;

&lt;p&gt;This was not a plug-and-play replacement. It entailed re-embedding all the rows in &lt;code&gt;screen_embeddings&lt;/code&gt;. I also had to convert the &lt;code&gt;VECTOR(n)&lt;/code&gt; column to conform to the Titan vector size. A 384-dimensional vector from MiniLM and a 512- or 1024-dimensional vector from Titan cannot be compared.&lt;/p&gt;

&lt;p&gt;For the generation step (the LLM call), I chose &lt;strong&gt;MiniMax 2.5&lt;/strong&gt;. The embedding model should be good at generating comparable vectors. The generation model has to be good at instruction-following and structured tool-calling.&lt;/p&gt;
&lt;h3&gt;
  
  
  Ingestion script example
&lt;/h3&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;psycopg2&lt;/span&gt;

&lt;span class="n"&gt;conn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;psycopg2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;dsn&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;DATABASE_URL&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;cur&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;cursor&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;navigation.json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;screens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;screen&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;screens&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;cur&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
        INSERT INTO screens (screen_name, description, required_params, web_url, mobile_path)
        VALUES (%s, %s, %s, %s, %s) RETURNING screen_id
        &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;screen&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;screen_name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;screen&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;description&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
         &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;screen&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;required_parameters&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]),&lt;/span&gt;
         &lt;span class="n"&gt;screen&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;web_url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;screen&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;mobile_path&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]),&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;screen_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cur&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fetchone&lt;/span&gt;&lt;span class="p"&gt;()[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

    &lt;span class="n"&gt;vector&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;embed_text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;screen&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;description&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="n"&gt;cur&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
        INSERT INTO screen_embeddings (screen_id, description, embedding)
        VALUES (%s, %s, %s)
        &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;screen_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;screen&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;description&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;vector&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;commit&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;I ensured that each description was concise enough to fit as a standalone unit. This is significant. Chunking is a technique used when dealing with documents longer than the context size of a model. Chunking a screen description that is already one or two sentences long will serve no purpose. It will only break down the meaning across several vectors, making retrieval more complex.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Screen: Edit User Details

This screen allows administrators to view and modify a user's profile information,
including full name, email address, phone number, assigned role, account status,
and department. The page also provides options to save changes, reset the user's
password, or deactivate the account.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Retrieval and parameter resolution
&lt;/h3&gt;

&lt;p&gt;During querying, the steps are:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; Embed the user query.&lt;/li&gt;
&lt;li&gt; Find the top-k closest screens based on cosine similarity.&lt;/li&gt;
&lt;li&gt; Pass the metadata associated with them to a generation prompt.&lt;/li&gt;
&lt;li&gt; Resolve any dynamic path parameters (&lt;code&gt;{hostname}&lt;/code&gt;, &lt;code&gt;{org_id}&lt;/code&gt;, &lt;code&gt;{user_id}&lt;/code&gt;).
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;retrieve_screens&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;query_vec&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;embed_text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;cur&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
        SELECT s.screen_name, s.web_url, s.mobile_path, s.required_params
        FROM screen_embeddings e
        JOIN screens s ON s.screen_id = e.screen_id
        ORDER BY e.embedding &amp;lt;=&amp;gt; %s
        LIMIT %s
        &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query_vec&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;cur&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fetchall&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;resolve_params&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url_template&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;url_template&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;format&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why k=3 and not k=1?&lt;/strong&gt;&lt;br&gt;
When we retrieve the only nearest screen, a poor embedding retrieval will silently mislead the user. There is nothing we can do about it. Retrieving the top 3 screens gives the LLM the opportunity to mitigate retrieval errors. It can distinguish the meaning of "edit my details" between "personal profile" and "account settings" screens using finer prompts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why two steps instead of one?&lt;/strong&gt;&lt;br&gt;
When split into two steps, the retrieval step becomes cheaper and cacheable regardless of generation. This way we can easily update any of these components, the embeddings, or the prompts.&lt;/p&gt;

&lt;p&gt;The navigation sub-agent is &lt;strong&gt;stateless&lt;/strong&gt; on purpose. It does not require any conversational context to answer "take me to X" requests. Statelessness is a design choice. Introducing memory would only introduce unnecessary overhead and potential bugs for an operation that never requires memory.&lt;/p&gt;
&lt;h3&gt;
  
  
  Connection Pooling, and Why Raw Connection Doesn't Survive Production Requests
&lt;/h3&gt;

&lt;p&gt;This ingestion script creates one instance of &lt;code&gt;psycopg2.connect()&lt;/code&gt; for a one-off batch job. This is fine for a script run once and exit.&lt;/p&gt;

&lt;p&gt;This approach would be unacceptable for the retrieval code path. There, the code is executed with each request from each concurrently active user. Creating a new TCP connection to the database on each request is relatively expensive (TCP handshake, TLS handshake, spinning up a new backend Postgres process).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The solution: a connection pool.&lt;/strong&gt;&lt;br&gt;
A connection pool is a set of already established connections that can be borrowed by requests and reused. I used &lt;code&gt;psycopg_pool&lt;/code&gt; for this. I configured the connection pool with min/max limits so the service cannot create more connections than Postgres is configured to accept.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;psycopg_pool&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;ConnectionPool&lt;/span&gt;

&lt;span class="n"&gt;pool&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;ConnectionPool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;conninfo&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;DATABASE_URL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;min_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;       &lt;span class="c1"&gt;# kept warm even when idle
&lt;/span&gt;    &lt;span class="n"&gt;max_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;       &lt;span class="c1"&gt;# hard ceiling, must stay under Postgres's max_connections
&lt;/span&gt;    &lt;span class="n"&gt;max_idle&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;300&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;      &lt;span class="c1"&gt;# close connections idle longer than 5 minutes
&lt;/span&gt;    &lt;span class="n"&gt;max_lifetime&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1800&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;# recycle every connection after 30 minutes
&lt;/span&gt;    &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;         &lt;span class="c1"&gt;# how long a request will wait for a free connection
&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;retrieve_screens&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;query_vec&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;embed_text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;pool&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;connection&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;cursor&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;cur&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;cur&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
                SELECT s.screen_name, s.web_url, s.mobile_path, s.required_params
                FROM screen_embeddings e
                JOIN screens s ON s.screen_id = e.screen_id
                ORDER BY e.embedding &amp;lt;=&amp;gt; %s
                LIMIT %s
                &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query_vec&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
            &lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;cur&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fetchall&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;with pool.connection() as conn:&lt;/code&gt; construct is doing something besides mere syntactic sugar.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;On completion, it gives the connection back to the pool, rather than closing it.&lt;/li&gt;
&lt;li&gt;If the block fails due to an exception, the pool determines if that connection should be returned or destroyed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is what separates pooling from simply having a globally-accessible connection object. There is no mechanism within a global connection to protect one request's broken transaction from another request. Pooling allows you to replace a connection that ends up in an invalid state.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwzdufesp5d1nyohm7sxs.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwzdufesp5d1nyohm7sxs.png" alt=" " width="800" height="271"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Handling expiring and stale connections specifically
&lt;/h3&gt;

&lt;p&gt;Database connections that live for a long time do not remain valid indefinitely. This is due to factors beyond your control:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Idle timeout kills on the server side:&lt;/strong&gt; Many managed PostgreSQL providers will quietly terminate connections left idle beyond a certain threshold.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Network drops:&lt;/strong&gt; Load balancers, NAT gateways, and cloud networking layers tend to drop idle TCP connections.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;PostgreSQL restarts or failovers:&lt;/strong&gt; A failover of a managed database will invalidate all current connections.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Such a pool will give you a perfectly valid-looking connection that will fail as soon as any query is executed on it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The proactive part of the solution:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;max_lifetime&lt;/code&gt; ensures that each connection must expire and be re-established within a certain period. No connection can get so old that it becomes stale on the server side.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;max_idle&lt;/code&gt; expires connections that have been sitting idle for too long. This circumvents most idle-timeout expirations.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The reactive part of the solution:&lt;/strong&gt;&lt;br&gt;
A health check during checkout. This detects any connection that &lt;em&gt;has&lt;/em&gt; become stale between uses.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;psycopg_pool&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;ConnectionPool&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;psycopg&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OperationalError&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;check_connection&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SELECT 1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;pool&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;ConnectionPool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;conninfo&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;DATABASE_URL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;min_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_idle&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;300&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_lifetime&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1800&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;check&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;check_connection&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;   &lt;span class="c1"&gt;# run before a connection is handed out
&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This went together with a thin retry wrapper around the execution of the query. It covers the situation where the connection fails &lt;em&gt;during a transaction&lt;/em&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;retrieve_screens_with_retry&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;attempts&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;attempt&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;attempts&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;retrieve_screens&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="n"&gt;OperationalError&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;attempt&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;attempts&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="k"&gt;raise&lt;/span&gt;
            &lt;span class="c1"&gt;# connection died mid-use; loop and let the pool hand out a fresh one
&lt;/span&gt;            &lt;span class="k"&gt;continue&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Summary of what enabled this solution to withstand production traffic:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A bounded pool size to protect Postgres.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;max_lifetime&lt;/code&gt; / &lt;code&gt;max_idle&lt;/code&gt; parameters for rotating connections proactively.&lt;/li&gt;
&lt;li&gt;A health check at checkout.&lt;/li&gt;
&lt;li&gt;An extremely specific retry strategy on &lt;code&gt;OperationalError&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Output shape:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"navigation"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"screen_name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"manage_organization"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://client.example.com/org/482"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Latency: caching and why the 0.8 latency in particular
&lt;/h3&gt;

&lt;p&gt;The first p50 latency for navigation was about 5 seconds per query. This is far too long for something so fundamentally a redirect.&lt;/p&gt;

&lt;p&gt;Those 5 seconds consisted of:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The embedding of the query.&lt;/li&gt;
&lt;li&gt;The round trip from the database to retrieve it.&lt;/li&gt;
&lt;li&gt;The dominant latency: the call to the LLM for generation of the reply.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Navigation queries are frequently duplicates of questions asked before.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Embedding-based response caching was added.&lt;/strong&gt;&lt;br&gt;
If the embedding of a new query was similar to a cached query by a cosine similarity greater than 0.8, I reused the cached result. The parameter resolution was always done fresh.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why 0.8?&lt;/strong&gt;&lt;br&gt;
0.8 is a threshold balancing precision against recall, specific to this particular task.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Below 0.8 (e.g., 0.6): Queries will be semantically distinct but still have the same answer cached. This causes potential errors.&lt;/li&gt;
&lt;li&gt;Above 0.8 (e.g., 0.95): The cache will almost never return anything, as paraphrasing will rarely produce semantically identical embeddings.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It was determined by testing some examples. This is why my retrospective calls for a proper evaluation on a labeled dataset.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Clarification on the caching process:&lt;/strong&gt;&lt;br&gt;
The key cannot simply be the raw query itself. That would be useless. It also cannot ignore the individual user information. The same query string from two different users will have two different URLs. What is cached is the &lt;em&gt;result of the retrieval and generation&lt;/em&gt;. Parameter resolution is always done fresh.&lt;/p&gt;


&lt;h2&gt;
  
  
  Part 2: Autonomous Execution Sub-agent (ReAct + MCP)
&lt;/h2&gt;

&lt;p&gt;This was the more challenging half of the system. It was the reason why the attempts by the previous vendors failed.&lt;/p&gt;
&lt;h3&gt;
  
  
  Why ReAct
&lt;/h3&gt;

&lt;p&gt;Autonomous execution is totally different from navigation. It doesn't involve "finding the one correct solution out of a given set of choices." It involves figuring out &lt;em&gt;how&lt;/em&gt; to achieve a certain goal.&lt;/p&gt;

&lt;p&gt;"Register a new product Wescco gas sensor with [config]" may involve several steps. Examples include finding the manufacturer id first, then the category id, and finally registering the product. All of these were not stated by the user.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft2ta12u6nmssoljt7r6o.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft2ta12u6nmssoljt7r6o.png" alt="react agent graph" width="642" height="363"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://docs.langchain.com/oss/python/langchain/agents" rel="noopener noreferrer"&gt;ReAct&lt;/a&gt; (Reasoning + Acting) is a pattern where the model alternates between reasoning and tool use. It observes the outcome of each tool before deciding the next step.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Example of reasoning:&lt;/strong&gt; "I need the manufacturer ID before I can create the product."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Example of tool use:&lt;/strong&gt; Looking up the manufacturer ID.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A one-shot prompt that plans the entire sequence up front fails when the output of a step determines the next step (e.g., the manufacturer does not exist yet and must be created first). The loop-until-done architecture of ReAct solves this problem naturally.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The real trade-off:&lt;/strong&gt; Latency and expense. ReAct makes multiple LLM invocations per task instead of just one. This is why autonomous execution latency becomes variable and, on average, much longer than navigation.&lt;/p&gt;
&lt;h3&gt;
  
  
  Why MCP and not hand-crafted function calling?
&lt;/h3&gt;

&lt;p&gt;All major LLM vendors provide the ability to perform "function calling" or "tool calling" directly from the prompt. You could do this by hand-crafting your own Python functions.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://modelcontextprotocol.io/" rel="noopener noreferrer"&gt;Model Context Protocol (MCP)&lt;/a&gt; standardizes the transport &amp;amp; discovery layer. Tools can be called over stdio or HTTP at run-time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why was this relevant practically?&lt;/strong&gt;&lt;br&gt;
The tools used to do tasks were shared infrastructure. They could theoretically be used by other internal tools/agents. The tools required running as an independently deployable and scalable piece of infrastructure. MCP does this for free. The tool server doesn't care which agent is hitting it, as long as it follows the protocol.&lt;/p&gt;


&lt;h2&gt;
  
  
  How to build an MCP tool?
&lt;/h2&gt;

&lt;p&gt;Building an MCP tool with FastMCP is remarkably straightforward. It feels just like writing standard Python functions. By wrapping your code in the &lt;code&gt;@mcp.tool&lt;/code&gt; decorator, FastMCP handles the complex underlying protocol. It automatically generates JSON schemas based on your Python type hints and docstrings. It manages the JSON-RPC communication transport.&lt;/p&gt;

&lt;p&gt;You just write the logic and define the inputs. The framework instantly exposes your tools to any MCP-compatible LLM or agent.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;fastmcp&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;FastMCP&lt;/span&gt;

&lt;span class="c1"&gt;# Initialize the MCP server
&lt;/span&gt;&lt;span class="n"&gt;mcp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;FastMCP&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CalculatorServer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Use the decorator to expose this function as an MCP tool
&lt;/span&gt;&lt;span class="nd"&gt;@mcp.tool&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;calculate_sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Adds two integer numbers together. This description tells the LLM when to use it.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="c1"&gt;# Starts the server using the standard input/output transport (default)
&lt;/span&gt;    &lt;span class="n"&gt;mcp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;p&gt;If you want a deeper visual walkthrough, check out &lt;a href="https://www.youtube.com/watch?v=_mUuhOwv9PY" rel="noopener noreferrer"&gt;Building Python MCP Servers with FastMCP&lt;/a&gt;. This video provides a comprehensive guide on building, testing, and connecting MCP servers using the framework.&lt;/p&gt;

&lt;h3&gt;
  
  
  What ctx actually is, and how to use it
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;ctx&lt;/code&gt; is the context object FastMCP injects into your tool or middleware function. It's your handle into everything about the current call that isn't a declared argument. This includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Request headers.&lt;/li&gt;
&lt;li&gt;State set by earlier middleware.&lt;/li&gt;
&lt;li&gt;Session data.&lt;/li&gt;
&lt;li&gt;Utilities like logging or progress reporting.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Full reference here: &lt;a href="https://gofastmcp.com/servers/context" rel="noopener noreferrer"&gt;gofastmcp.com/servers/context&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;In this project, &lt;code&gt;ctx&lt;/code&gt; is where two crucial things lived:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; The user's bearer token, read off the request headers.&lt;/li&gt;
&lt;li&gt; The &lt;code&gt;user_email&lt;/code&gt; decoded from the JWT by the auth middleware.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Both live on &lt;code&gt;ctx&lt;/code&gt; rather than being passed as regular tool parameters. They are set by code, not by the model.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why each and every one of our tools wraps an existing REST endpoint, and not raw SQL
&lt;/h3&gt;

&lt;p&gt;This is the straightforward solution to the problem of generating SQL code from natural language. Each of our MCP tools wraps an already-authorized REST endpoint at the Java Spring backend.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The implication:&lt;/strong&gt; The LLM cannot initiate anything other than what the same user could have done using the existing UI. It is limited precisely by the same backend validation and authorization checks. The LLM does not get to create a query. It only provides parameters for some predefined, reviewed action.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;fastmcp&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;FastMCP&lt;/span&gt;

&lt;span class="n"&gt;mcp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;FastMCP&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;custom-tools&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nd"&gt;@mcp.tool&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;create_user&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;email&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;org_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Register a new user in the given organization.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;token&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;request_context&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Authorization&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;http_client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;JAVA_BACKEND_URL&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/api/users&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;email&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;email&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;org_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;org_id&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Authorization&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;ctx&lt;/code&gt; parameter is where the caller's bearer token comes from. See what ctx actually is above.&lt;/p&gt;

&lt;h3&gt;
  
  
  Reasons for the collapsing of tools created from chained API calls into a single tool
&lt;/h3&gt;

&lt;p&gt;After reverse engineering the internal API calls, I understood sequences like: manufacturer look-up, then user-id look-up, then creation. In each case, the chain of API calls was collapsed into a &lt;em&gt;single&lt;/em&gt; MCP tool. They were not separate tools for each API call.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Justification:&lt;/strong&gt;&lt;br&gt;
Every extra tool added to the agent's set of tools adds to the reasoning load on the model. It now has to reason about the order of calls and how to pass outputs as inputs. Reducing a known sequence of calls down to one tool removes this sequencing logic. It is done via deterministic Python code, which will always be correct. It leaves the LLM with the simple decision of which tool to call and what parameters to use.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;In general:&lt;/strong&gt; If it's deterministic and you can calculate it at build time, then do so. Don't let the LLM figure it out again on every request.&lt;/p&gt;

&lt;p&gt;Where the underlying APIs provided for paging and good default values, I let that pass through unaltered. This lowers the number of arguments the LLM has to juggle. It lowers the probability of an ill-formed API call.&lt;/p&gt;
&lt;h3&gt;
  
  
  A backup for wrongly classified queries: navigating within the ReAct loop
&lt;/h3&gt;

&lt;p&gt;There are cases where the semantic router doesn't quite get it right. A navigational request gets classified as an API call. This gets sent to the ReAct agent, which has no way to fulfill "take me to the dashboard" using &lt;code&gt;create_user&lt;/code&gt; or &lt;code&gt;update_form&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Instead of making the classifier perfect, I put in a safety net on the execution part. I added a navigation tool to the list of tools provided to the ReAct agent. This consists of the exact retrieval and resolution mechanism used for navigation, embedded within a callable function.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;langchain_core.tools&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;tool&lt;/span&gt;

&lt;span class="nd"&gt;@tool&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;find_screen&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;hostname&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;org_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Use this when the user&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s request is actually about navigating to a
    screen in the product, rather than creating, updating, or deleting data.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;screens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;retrieve_screens&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;top&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;screens&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;url&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;resolve_params&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;top&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;web_url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;hostname&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;hostname&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;org_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;org_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;screen_name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;top&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;screen_name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="n"&gt;react_tools&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;mcp_tools&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;find_screen&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Two things to clarify:&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; This tool lives inside FastAPI, not the MCP server. MCP is worth its overhead only when a tool needs to be independently deployable or shareable.&lt;/li&gt;
&lt;li&gt; The docstring is important. It is the only clue the ReAct agent has regarding &lt;em&gt;when&lt;/em&gt; to use it. It must explicitly indicate the situation.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This addresses only one side of the problem. When a real autonomous execution query is mistakenly directed to the navigation sub-agent, there's no safety mechanism. The navigation sub-agent's RAG pipeline doesn't have any means of executing the write request. This is one of the reasons why the below retrospective calls attention to the problem of confidence-based fallback.&lt;/p&gt;

&lt;h3&gt;
  
  
  Transient token propagation - the harder bit, explained in depth
&lt;/h3&gt;

&lt;p&gt;Typically, MCP servers work with long-lived credentials. You may keep one token per client in your database and use it many times. In our case, tokens expired in &lt;strong&gt;2 hours&lt;/strong&gt;. Each incoming request from a user could be performed with &lt;em&gt;another&lt;/em&gt; token.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How this breaks the naive approach:&lt;/strong&gt;&lt;br&gt;
An agent, according to LangChain, should establish a connection to its MCP server once per agent initialization. That means all tool objects are initialized using some credentials available at that time. To use transient per-user tokens naively, you would need to create a new agent object for each incoming request.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The solution:&lt;/strong&gt;&lt;br&gt;
Separate the moment of &lt;strong&gt;binding a tool&lt;/strong&gt; from &lt;strong&gt;supplying the credentials of the tool&lt;/strong&gt;. The MCP adapter in LangChain allows for &lt;strong&gt;tool interceptors&lt;/strong&gt;. These are call-time callbacks that allow the insertion of request-specific information (like a bearer token) into a tool that was bound at startup time.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://docs.langchain.com/oss/python/langchain/mcp#tool-interceptors" rel="noopener noreferrer"&gt;Tool interceptors&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://docs.langchain.com/oss/python/langchain/mcp#custom-interceptors" rel="noopener noreferrer"&gt;Custom interceptors&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;langchain_mcp_adapters.client&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;MultiServerMCPClient&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;token_interceptor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;user_token&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Authorization&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bearer &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;user_token&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;MultiServerMCPClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;nimbly&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;transport&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;streamable_http&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;http://localhost:9001/mcp&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;tools&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_tools&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;interceptor&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;lambda&lt;/span&gt; &lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;token_interceptor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;current_user_token&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;On the MCP server-side, the middleware takes the token out of the request &lt;code&gt;ctx&lt;/code&gt; &lt;em&gt;on a per-call basis&lt;/em&gt;. It then uses it to authenticate the downstream API request.&lt;/p&gt;

&lt;p&gt;Wherever an API had to know who was behind the change (audit purposes), I took the email claim from the JWT itself. I did not request the LLM to provide the email as a parameter. This is the trust boundary. The JWT claim is signed and belongs to an authorized session. The tool parameter provided by the LLM is arbitrary text.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@mcp.middleware&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;auth_middleware&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;call_next&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;token&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Authorization&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user_email&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;decode_jwt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;email&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;call_next&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Why use streamable HTTP and why co-locate first
&lt;/h3&gt;

&lt;p&gt;Two MCP transports are relevant here:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;stdio:&lt;/strong&gt; Used by tools operating as subprocesses of the agent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Streamable HTTP:&lt;/strong&gt; Tools that can be found over the network in a tool server form.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Streamable HTTP was the best choice. This tool server had to be independently addressable. stdio is tightly coupled to a particular agent process lifecycle.&lt;/p&gt;

&lt;p&gt;The initial setup had both the MCP server and the FastAPI agent running inside the same container. They were listening on &lt;code&gt;localhost&lt;/code&gt;. This was to avoid the additional latency of an extra hop while the system was still relatively small. The two components were eventually separated into individually deployable entities when there was sufficient demand.&lt;/p&gt;

&lt;p&gt;The basic idea is to colocate based on latency and simplicity. Only separate them when there is a real, measurable requirement.&lt;/p&gt;

&lt;h3&gt;
  
  
  Auth chain end to end, and why each hop verifies independently
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4085ds4vgb9ha7g1o5c5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4085ds4vgb9ha7g1o5c5.png" alt=" " width="799" height="374"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;In this scenario, two separate tokens are involved. Mixing them up would be erroneous.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The service JWT:&lt;/strong&gt; Used to authenticate Java Spring itself whenever it sends a request to FastAPI. It shows "this request is genuinely sent by our backend." FastAPI verifies this using Java Spring's &lt;a href="https://auth0.com/docs/secure/tokens/json-web-tokens/json-web-key-sets" rel="noopener noreferrer"&gt;JWKS&lt;/a&gt; endpoint.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The user token:&lt;/strong&gt; Authenticates the end user for whom the specific write downstream is done. FastAPI does not validate this token. It gets passed through the interceptor directly to the MCP server. It is ultimately authenticated by the downstream Java REST API.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is by design. FastAPI does not need to know how to validate a user session. Validating the same thing again in another place could lead to inconsistency.&lt;/p&gt;

&lt;h3&gt;
  
  
  Conversation memory (what a checkpoint actually is)
&lt;/h3&gt;

&lt;p&gt;The agent state, in LangGraph, tracks the user query and a typed final-answer object:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pydantic&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;BaseModel&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;typing&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Optional&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;NavigationResult&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;screen_name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;AutoExecutionResult&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;summary&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;tool_calls&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;FinalAnswer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;navigation&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Optional&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;NavigationResult&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="n"&gt;auto_execution&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Optional&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;AutoExecutionResult&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;AgentState&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;user_query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;final_answer&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;FinalAnswer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;FinalAnswer&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;final_answer&lt;/code&gt; is a struct that contains &lt;em&gt;both&lt;/em&gt; fields. Only the relevant one is set for a particular turn. This was done deliberately. A client interpreting the answer does not have to do a case analysis. They can always look at &lt;code&gt;final_answer.navigation&lt;/code&gt; or &lt;code&gt;final_answer.autonomous_execution&lt;/code&gt;. The other field remains &lt;code&gt;None&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;A new UUID &lt;code&gt;thread_id&lt;/code&gt; is assigned to each chat thread. A &lt;strong&gt;checkpointer&lt;/strong&gt; saves the whole state of the graph at the end of each step. It is identified by &lt;code&gt;thread_id&lt;/code&gt;. Next time input comes for the same &lt;code&gt;thread_id&lt;/code&gt;, LangGraph will restore the last checkpoint state and continue. Hence, the agent will have a memory between steps.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;langgraph.checkpoint.postgres.aio&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;AsyncPostgresSaver&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;AsyncPostgresSaver&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_conn_string&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;DATABASE_URL&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;checkpointer&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;graph&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;builder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;compile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;checkpointer&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;checkpointer&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;graph&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ainvoke&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user_query&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;configurable&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;thread_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;thread_id&lt;/span&gt;&lt;span class="p"&gt;}},&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Docs: &lt;a href="https://docs.langchain.com/oss/python/langgraph/persistence" rel="noopener noreferrer"&gt;LangGraph persistence&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why use Postgres for the checkpointer?&lt;/strong&gt;&lt;br&gt;
It comes down to durability and consistency with the rest of the system. An in-memory checkpointer loses all conversation information whenever a process is restarted. It is unusable for production.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A point of confusion worth clarifying:&lt;/strong&gt;&lt;br&gt;
A checkpoint snapshot includes &lt;em&gt;all&lt;/em&gt; knowledge the graph accumulated. This includes reasoning steps and sub-agent calls. This is necessary to resume correctly. However, this is not what should be included in the "history" displayed to the user.&lt;/p&gt;

&lt;p&gt;I included a &lt;strong&gt;separate &lt;code&gt;chat_messages&lt;/code&gt; table&lt;/strong&gt;. This was specifically logged by the application code when a turn had completed. It is used for rendering clean, de-duplicated conversation history in the UI. The checkpointer's job is to save the working memory of the agent. &lt;code&gt;chat_messages&lt;/code&gt; saves the user's view of that memory.&lt;/p&gt;
&lt;h3&gt;
  
  
  Streaming – Why Per-Node Events Rather Than Waiting for the Entire Response
&lt;/h3&gt;

&lt;p&gt;Responses are streamed from FastAPI to the Java layer using Server Sent Events (SSE). LangGraph's &lt;code&gt;astream&lt;/code&gt; creates an event each time a node completes its execution. This is as opposed to the client waiting for the entire multi-step ReAct process to complete.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;stream_agent_response&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;thread_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;graph&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;astream&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user_query&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;configurable&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;thread_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;thread_id&lt;/span&gt;&lt;span class="p"&gt;}},&lt;/span&gt;
        &lt;span class="n"&gt;stream_mode&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;values&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;data: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is especially important for task processing. The latency there is unpredictable. Without streaming, the interface would be sitting on a loading screen with no indication of what is being processed. Streaming individual node-level events (e.g., "checking manufacturer") makes it possible for the frontend to see progress.&lt;/p&gt;

&lt;p&gt;SSE is useful because the data stream goes in one direction only (server to client) for a single request-response interaction. It is the simpler protocol that does not require a connection to go through HTTP.&lt;/p&gt;

&lt;h3&gt;
  
  
  Request/response shape
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"user_query"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"register a new product called Wescco gas sensor"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"thread_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"b3f1..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"hostname"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"client.example.com"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"platform"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"web"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"metadata"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"user_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"u_123"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"org_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"org_42"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"header"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"user_token"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;org_id&lt;/code&gt; variable is important. The tool serves several organizations with entirely separate datasets. Every query and tool function has an implicit scope based on the &lt;code&gt;org_id&lt;/code&gt;. This scoping is what prevents the multi-tenant architecture from exposing one organization's data to another. This scoping must take place at the query/function call level, not just in the prompt.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I'd do differently (retrospective)
&lt;/h2&gt;

&lt;p&gt;Shipping software that works does not equal shipping software that works &lt;em&gt;well&lt;/em&gt;. There are a few areas where I would change my approach in retrospect:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No official eval set:&lt;/strong&gt; The similarity thresholds and the 0.8 cache hit threshold were tuned empirically through spot checks. They were not tested against an annotated set of queries. This is an easy problem to solve with even a small gold set of a few hundred annotated query/route pairs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Absence of a fail-safe alternative in case of low confidence:&lt;/strong&gt; When the semantic router score just exceeds the threshold, the design follows the same route. There is no safeguard for an execution query mistakenly routed to navigation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Token expiration isn't done gracefully mid-job:&lt;/strong&gt; A ReAct job spanning longer than a 2-hour token lifetime requires a refresh-and-retry pattern. Currently, a job crossing this line will simply fail halfway with a partially completed transaction.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Insufficient auditing of tool calls:&lt;/strong&gt; Every tool invocation should potentially be included in a structured audit log. I developed the code for extracting emails but not yet aggregated in an audit log.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Did not take into account the expenses of designing on the same day:&lt;/strong&gt; In itself, not a technical mistake, but an organizational one that caused me more stress. When you get an unclear, critical scope on day one, discuss the time frame before writing, not after your first attempt fails.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Up next
&lt;/h2&gt;

&lt;p&gt;In the next post, I'll cover adding LangSmith for tracing, evaluation, and catching regressions in the classifier and retrieval thresholds before they hit production.&lt;/p&gt;

&lt;p&gt;Thanks for reading. If you have any doubts, suggestions, or want to go deeper, I am happy to answer.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Connect with me:&lt;/strong&gt; &lt;a href="https://www.linkedin.com/in/ryan-fernandes-294136231/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; · &lt;a href="https://www.instagram.com/ryan_ferds?igsh=aXVwazYzcWQ1dW43" rel="noopener noreferrer"&gt;Instagram&lt;/a&gt; · &lt;a href="https://github.com/RyanFernandes23" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>agents</category>
      <category>softwaredevelopment</category>
    </item>
  </channel>
</rss>
