<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: B Gopinath</title>
    <description>The latest articles on DEV Community by B Gopinath (@bgopinath0825).</description>
    <link>https://dev.to/bgopinath0825</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150369%2F42956608-7eb5-4780-b154-b28022463f11.png</url>
      <title>DEV Community: B Gopinath</title>
      <link>https://dev.to/bgopinath0825</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/bgopinath0825"/>
    <language>en</language>
    <item>
      <title>Semantic Similarity Isn't Relevance: Stopping an Incident Agent From Hallucinating Across Services</title>
      <dc:creator>B Gopinath</dc:creator>
      <pubDate>Tue, 29 Sep 2026 17:49:05 +0000</pubDate>
      <link>https://dev.to/bgopinath0825/semantic-similarity-isnt-relevance-stopping-an-incident-agent-from-hallucinating-across-services-37b</link>
      <guid>https://dev.to/bgopinath0825/semantic-similarity-isnt-relevance-stopping-an-incident-agent-from-hallucinating-across-services-37b</guid>
      <description>&lt;p&gt;The first bad answer my incident agent gave me was also the most convincing one.&lt;/p&gt;

&lt;p&gt;I had submitted a pool exhaustion incident for a service that had never had one. The agent replied with a crisp diagnosis: a connection leak in &lt;code&gt;OrderPaymentProcessor&lt;/code&gt;. Correct format, plausible cause, confident tone. The only problem was that the service in question had no class by that name. The memory it drew from belonged to a different service entirely.&lt;/p&gt;

&lt;p&gt;Vector similarity had done its job. The two incidents &lt;em&gt;were&lt;/em&gt; similar. That's exactly why it was dangerous.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why incident memory is different
&lt;/h2&gt;

&lt;p&gt;In a general-purpose RAG system, a semantically close chunk is usually a good chunk. Incidents break that assumption, because the useful unit isn't just "what happened" but "what happened &lt;em&gt;to this system&lt;/em&gt;." Two services can both exhaust connection pools for entirely different reasons. A fix for one might be meaningless, or harmful, for the other.&lt;/p&gt;

&lt;p&gt;So the design question became: how do I let the agent use cross-service knowledge as a hint without letting it treat that knowledge as fact?&lt;/p&gt;

&lt;h2&gt;
  
  
  Step one: make service a parseable field
&lt;/h2&gt;

&lt;p&gt;Every memory in the bank follows one format:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[&amp;lt;TIMESTAMP&amp;gt;] INCIDENT: &amp;lt;TITLE&amp;gt; | SERVICE: &amp;lt;SERVICE_NAME&amp;gt; | ROOT CAUSE: &amp;lt;ROOT_CAUSE&amp;gt; | FIX (&amp;lt;OUTCOME&amp;gt;): &amp;lt;FIX_DESCRIPTION&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The SERVICE: segment is what everything else hangs on. Recall is done through Hindsight, which blends temporal, entity, semantic, and recency strategies (TEMPR), so it already does better than raw nearest-neighbor. Service names are naturally entities, which helps. But I didn't want to rely on the retrieval layer alone to enforce a correctness boundary. Retrieval finds candidates; my code decides what they're allowed to mean.&lt;/p&gt;

&lt;p&gt;That decision lives in relevance.py. It parses each recalled memory back into structured fields and compares the memory's service to the incident's service.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step two: three tiers, not a threshold&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I first tried a score cutoff: below 0.7, drop it. It didn't work, because score measures similarity and I needed to measure applicability. A 0.94 cross-service memory is still cross-service.&lt;/p&gt;

&lt;p&gt;So each memory gets a grade:&lt;/p&gt;

&lt;p&gt;direct_match: the memory is from the same service (is_same_service == True). This is the only tier that can ground a high-confidence recommendation.&lt;br&gt;
analogy: a different service with a similar failure pattern. The agent can say "this resembles a pool exhaustion we saw elsewhere," but it is prohibited from claiming that cross-service internals exist in the target. Confidence is capped at medium.&lt;br&gt;
low: spurious matches, filtered out or pushed down the ranking.&lt;/p&gt;

&lt;p&gt;The grade travels with the memory all the way to the client:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;  "id": "mem_01",
  "score": 0.94,
  "service": "order-service",
  "is_same_service": true,
  "outcome": "WORKED",
  "relevance_grade": "direct_match"
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The client shows the grade next to each recalled memory. When an engineer looks at a recommendation, they can see whether it rests on a direct precedent or on an analogy, and calibrate accordingly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step three: give the model rules, not vibes&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The grades reach the LLM as constraints. The call goes to Groq's openai/gpt-oss-120b at temperature 0.2 with a JSON response format. The model must produce:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;interface StructuredRecommendation {
  likely_root_cause: string;
  confidence: "high" | "medium" | "low";
  historical_evidence: string[];
  recommended_fix: string;
  citing: string;
  runbook: string[];
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The constraint on analogies is stated explicitly in the prompt: analogies may inform the pattern but may not import specific component names, configuration, or code paths from another service. And citing must name the memory the recommendation depends on. If the model says "high" but cites an analogy, that inconsistency is visible right there in the response.&lt;/p&gt;

&lt;p&gt;This split matters. I don't ask the model to judge whether a memory applies to a given service. That's a boolean I can compute exactly. I ask it to write clearly within limits that code has already established. Models are good at the second job and unreliable at the first.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What happens in each case&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Direct match. An order-service pool incident with a recalled order-service leak memory returns high confidence, the leak diagnosis, a circuit breaker fix, and a citing line referencing the 14 May precedent.&lt;/p&gt;

&lt;p&gt;No history. A service with no prior incidents gets a recommendation that stays generic: symptom-driven diagnostic steps, medium or low confidence, no invented class names. This is a scenario in the test suite (test_scenario_a_and_c_no_matching_memory_and_unrelated_service), and it asserts that components from other services don't leak in.&lt;/p&gt;

&lt;p&gt;Analogy. A different service with similar symptoms gets something like: "this pattern has appeared on another service, so check for leaked connections around external calls," with confidence capped at medium and no claim about that service's internals.&lt;/p&gt;

&lt;p&gt;That third case is the one I'm proudest of. The agent is still useful when it doesn't have a direct precedent. It just says less, and says it honestly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The cost of doing it this way&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It's not free. A few things I'd flag for anyone copying this approach:&lt;/p&gt;

&lt;p&gt;It depends on service naming discipline. If order-service and orders-svc are the same thing in practice, exact matching will miss it. Canonical service names are a real requirement, and you'll need a normalization step or a registry.&lt;br&gt;
Analogies are conservative by design. Sometimes a cross-service memory really does apply, and the agent will under-claim. I accept that. In an incident, under-claiming is cheaper than a confident wrong fix.&lt;br&gt;
Parsing is a contract. Because relevance.py parses the memory line format, any change to that format has to be made in the writer and the parser together. There's a dedicated unit test file for parsing, service matching, and ranking, because this is the part I least want to break silently.&lt;br&gt;
Lessons learned&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Similarity measures resemblance, not applicability. If your domain has a scope boundary (service, tenant, environment, region), enforce it in code rather than hoping the embedding captures it.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Grade before you generate. Compute what each piece of context is allowed to justify, and pass that to the model as a constraint.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Cap confidence structurally. I don't ask the model to "be careful." The tier determines the maximum confidence, and the schema carries it.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Surface the grade to humans. Showing "direct match" or "analogy" next to each memory lets a tired engineer decide how much to trust it in about a second.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Test the negative cases. The most valuable test I have checks that something doesn't appear in the output.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you're building on a memory layer, the Hindsight docs are worth reading for what recall gives you out of the box, and Vectorize's overview of agent memory covers the broader design space.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://hindsight.vectorize.io/guides/2026/04/23/guide-why-your-ai-agent-needs-memory?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Hindsight — Why Your AI Agent Needs Memory&lt;/a&gt;&lt;br&gt;
&lt;a href="https://vectorize.io/what-is-agent-memory?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Vectorize — What Is Agent Memory?&lt;/a&gt;&lt;br&gt;
&lt;a href="https://hindsight.vectorize.io/blog/2026/07/31/evaluate-agent-memory-system?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Hindsight — Agent Memory Evaluation Guide&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>llm</category>
      <category>softwareengineering</category>
    </item>
  </channel>
</rss>
