<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: ALUVALA REVANTH</title>
    <description>The latest articles on DEV Community by ALUVALA REVANTH (@aluvala_revanth_65bf5157f).</description>
    <link>https://dev.to/aluvala_revanth_65bf5157f</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4147885%2F404e21ae-1ec9-49cc-956f-3a18bd53b455.png</url>
      <title>DEV Community: ALUVALA REVANTH</title>
      <link>https://dev.to/aluvala_revanth_65bf5157f</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/aluvala_revanth_65bf5157f"/>
    <language>en</language>
    <item>
      <title>How I used Hindsight to trigger strict support escalations</title>
      <dc:creator>ALUVALA REVANTH</dc:creator>
      <pubDate>Mon, 28 Sep 2026 18:38:24 +0000</pubDate>
      <link>https://dev.to/aluvala_revanth_65bf5157f/how-i-used-hindsight-to-trigger-strict-support-escalations-4c42</link>
      <guid>https://dev.to/aluvala_revanth_65bf5157f/how-i-used-hindsight-to-trigger-strict-support-escalations-4c42</guid>
      <description>&lt;h1&gt;
  
  
  How I used Hindsight to trigger strict support escalations
&lt;/h1&gt;

&lt;p&gt;Every support team has a version of the same argument. One person says we escalate too late and customers are furious. Another says we escalate too often and the senior queue is drowning. Both are right, because "escalate when it seems bad" is not a rule, it is a mood.&lt;/p&gt;

&lt;p&gt;I wanted a rule I could state in one sentence, defend in a review, and test in CI. This is how I built one with &lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight&lt;/a&gt; and FastAPI.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rule
&lt;/h2&gt;

&lt;p&gt;A case escalates when a customer has contacted us three or more times about the same unresolved issue.&lt;/p&gt;

&lt;p&gt;That sentence hides three separate questions, and each one needs a different kind of answer:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Who is the customer?&lt;/strong&gt; Identity, across chat, email, and phone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Which contacts are about the same issue?&lt;/strong&gt; A judgement call.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Is it still unresolved, and how many times?&lt;/strong&gt; Arithmetic.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Most of the trouble I had came from mixing those together. The rest of this post is about pulling them apart.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the system looks like
&lt;/h2&gt;

&lt;p&gt;The backend is a small FastAPI service. &lt;code&gt;hindsight_client.py&lt;/code&gt; wraps the memory layer, &lt;code&gt;llm_client.py&lt;/code&gt; wraps Groq running &lt;code&gt;qwen/qwen3-32b&lt;/code&gt;, &lt;code&gt;memory_schema.py&lt;/code&gt; holds the Pydantic models, and &lt;code&gt;app.py&lt;/code&gt; exposes endpoints such as &lt;code&gt;POST /escalate/{email}&lt;/code&gt; and &lt;code&gt;POST /summarise/{email}&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Customers are keyed on their email address, because it is the one identifier that all three channels reliably carry. Every interaction is stored against that email in &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;Hindsight&lt;/a&gt;, and every escalation check starts by recalling them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why "strict" matters
&lt;/h2&gt;

&lt;p&gt;A model asked "should we escalate this customer?" will answer fluently and mostly sensibly. Mostly is the problem. Run the same history twice and the answer can move. Change the wording of one message and it can move again. I cannot write a test for a decision that drifts, and I cannot explain to a support lead why customer A was escalated and customer B, with a nearly identical history, was not.&lt;/p&gt;

&lt;p&gt;Strict, for me, means three properties:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Deterministic.&lt;/strong&gt; The same history produces the same decision, every time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Explainable.&lt;/strong&gt; I can point at the exact records that triggered it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Non-overridable by the model.&lt;/strong&gt; The language model can describe a decision. It cannot make one.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Memory is what makes those properties possible, because a rule can only be strict about facts that were recorded. That is the job &lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;agent memory&lt;/a&gt; does here: it turns "what happened" from something a model reconstructs into something the system retrieves.&lt;/p&gt;

&lt;h2&gt;
  
  
  The facts I store
&lt;/h2&gt;

&lt;p&gt;Each interaction is recorded with just enough structure for the rule to use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pydantic&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;BaseModel&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;typing&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Literal&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Interaction&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;email&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;channel&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Literal&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;chat&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;email&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;phone&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;issue_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;        &lt;span class="c1"&gt;# which issue this contact belongs to
&lt;/span&gt;    &lt;span class="n"&gt;summary&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;resolved&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;issue_id&lt;/code&gt; field is where the judgement call from question two lives. I decide it once, when the interaction is stored, and after that it is a fact like any other. Deciding "same issue or not" at write time means I am not re-litigating it on every read.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rule as code
&lt;/h2&gt;

&lt;p&gt;With facts in hand, the check is short enough to read in one glance:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;collections&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Counter&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;needs_escalation&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;interactions&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;Interaction&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;threshold&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;open_counts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Counter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;issue_id&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;interactions&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;resolved&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;any&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;threshold&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;open_counts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;values&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two details carry a lot of weight:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Resolved contacts do not count.&lt;/strong&gt; A customer who had a problem, got it fixed, and later reports something new should not inherit the first problem's history.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The count is per issue, not per customer.&lt;/strong&gt; Three contacts about three different things is a busy customer, not a failing case.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The endpoint then keeps the model on the communication side of the line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@app.post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/escalate/{email}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;escalate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;email&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;interactions&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;recall_interactions&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;email&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;decision&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;needs_escalation&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;interactions&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;note&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;llm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write_handoff_note&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;escalate&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;history&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;interactions&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;escalate&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;note&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;note&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The decision is computed before the model is called, and the model receives it as input. There is no path where the model's output changes &lt;code&gt;decision&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it behaves
&lt;/h2&gt;

&lt;p&gt;My test data has five synthetic customers. The stress case has four interactions across chat, email, and phone, all about a billing problem that never got resolved. It escalates, and the handoff note summarises what was tried on each channel. A customer whose bug was reported and resolved with a workaround does not escalate, because nothing is open. A customer with a smooth plan upgrade does not escalate either.&lt;/p&gt;

&lt;p&gt;The cases I care about most are the ones that should stay quiet. A strict rule is only useful if it has a false-positive story as good as its true-positive one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lessons learned
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Write the rule as a sentence first.&lt;/strong&gt; If you cannot say it in one sentence, you cannot enforce it. Mine took three tries.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Separate judgement from arithmetic.&lt;/strong&gt; Let the model do the fuzzy part once, store the result, and count in code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Choose your error deliberately.&lt;/strong&gt; A threshold of three means a customer can be frustrated twice before anyone senior sees the case. That was a conscious trade against flooding the senior queue, and I would rather state it than pretend the threshold is neutral.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Make the model unable to override the decision.&lt;/strong&gt; Not "unlikely to". Unable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Know where the fuzziness lives.&lt;/strong&gt; Assigning &lt;code&gt;issue_id&lt;/code&gt; is still a model call, and it can still be wrong. A rephrased complaint that gets a new &lt;code&gt;issue_id&lt;/code&gt; will slip past the threshold. Tightening that is the next thing I want to fix.&lt;/p&gt;

&lt;p&gt;If you want a memory layer to hold the facts while you build rules like this, the &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;Hindsight docs&lt;/a&gt; are the place to start.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agentskills</category>
      <category>agents</category>
      <category>python</category>
    </item>
  </channel>
</rss>
