<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: AKSHITH REDDY</title>
    <description>The latest articles on DEV Community by AKSHITH REDDY (@akshith_reddy_453d62a9a9a).</description>
    <link>https://dev.to/akshith_reddy_453d62a9a9a</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150712%2Fc1210c4d-0c8d-4a0f-b11f-8bb12cb29b5c.png</url>
      <title>DEV Community: AKSHITH REDDY</title>
      <link>https://dev.to/akshith_reddy_453d62a9a9a</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/akshith_reddy_453d62a9a9a"/>
    <language>en</language>
    <item>
      <title>Why My Reorder Model Now Asks Hindsight Before Ordering</title>
      <dc:creator>AKSHITH REDDY</dc:creator>
      <pubDate>Tue, 29 Sep 2026 18:07:28 +0000</pubDate>
      <link>https://dev.to/akshith_reddy_453d62a9a9a/why-my-reorder-model-now-asks-hindsightbefore-ordering-2k9o</link>
      <guid>https://dev.to/akshith_reddy_453d62a9a9a/why-my-reorder-model-now-asks-hindsightbefore-ordering-2k9o</guid>
      <description>&lt;p&gt;Last quarter my forecasting model told a shopkeeper to order 50 cartons of noodles from a supplier&lt;br&gt;
whose minimum order quantity was 50. The owner had told my system, weeks earlier, that he never&lt;br&gt;
wanted more than 35 units of that product on the shelf. The model was right about demand and wrong&lt;br&gt;
about the shop.&lt;br&gt;
That gap between what the numbers say and what the owner has already decided is the problem I spent most of this project&lt;br&gt;
on. This is how I closed it with Hindsight, an open-source agent memory system, and what I would do differently.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8iox3qktvadele9srdr6.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8iox3qktvadele9srdr6.jpeg" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What the system does
&lt;/h2&gt;

&lt;p&gt;DukaanPulse is an operations console for small Indian retail shops (kirana and general stores). The owner gets a dashboard&lt;br&gt;
with today's sales, profit, order count and low-stock items; billing by keyboard, voice or receipt scan; a khata (udhaar) ledger&lt;br&gt;
for customer credit; an expenses ledger; and an advisor they can ask questions in Hinglish.&lt;br&gt;
The backend is a FastAPI service that acts as an orchestration layer. It has four stores of knowledge behind it, and I was strict&lt;br&gt;
about what belongs where:&lt;br&gt;
PostgreSQL holds structured facts: products, inventory, sales and purchases, suppliers, customers and khata balances,&lt;br&gt;
expenses, orders, audit logs.&lt;br&gt;
The ML layer (LightGBM/XGBoost) produces demand forecasts, adapts them online to recent sales, flags anomalies, and&lt;br&gt;
folds in weather, holidays and price changes&lt;br&gt;
Hindsight holds long-term business memory: owner preferences, supplier conditions, customer patterns, business events,&lt;br&gt;
and past decisions with their outcomes.&lt;br&gt;
Gemini does the reasoning and the conversation, using everything above.&lt;br&gt;
The rule I wrote on the whiteboard, and mostly kept: Postgres stores what happened, the model predicts what will happen,&lt;br&gt;
Hindsight remembers what the owner knows and what we tried, and the LLM never does arithmetic.&lt;/p&gt;

&lt;h2&gt;
  
  
  The through-line: a forecast is not a decision
&lt;/h2&gt;

&lt;p&gt;A demand forecast is a number. A reorder decision is a number filtered through constraints that never appear in your sales&lt;br&gt;
table:&lt;br&gt;
"Don't keep more than 35 Maggi units in stock." (the owner's shelf, cash and habits)&lt;br&gt;
"Sharma Distributors: good prices, MOQ 50, two-day delivery." (a supplier relationship)&lt;br&gt;
"This product sells faster during local festivals." (something learned by watching)&lt;br&gt;
"A supermarket opened nearby in September 2026." (an event that changes everything after it)&lt;br&gt;
My first version put all of this into Postgres columns. max_stock_units on products, moq on suppliers. That worked for exactly&lt;br&gt;
the constraints I had thought of in advance. Owners do not talk in schema. They say "Sharma is fine but his delivery has been&lt;br&gt;
late" in the middle of a voice command, and that sentence has nowhere to go in a relational model.&lt;br&gt;
So I stopped trying to anticipate the fields and started treating the owner's conversation as a source of memory.&lt;/p&gt;

&lt;h2&gt;
  
  
  Retaining what the owner says
&lt;/h2&gt;

&lt;p&gt;Every owner conversation, and every event I can detect, goes through Hindsight's retain . I use one bank per store. Banks are&lt;br&gt;
strictly isolated, which matters when one profile manages several stores.&lt;br&gt;
from hindsight_client import Hindsight&lt;br&gt;
hs = Hindsight(base_url=settings.HINDSIGHT_URL)&lt;br&gt;
def remember_owner_statement(store_id: str, text: str, source: str):&lt;br&gt;
hs.retain(&lt;br&gt;
bank_id=f"store-{store_id}",&lt;br&gt;
content=text,&lt;br&gt;
context=f"owner {source}", # "voice command", "advisor chat", ...&lt;br&gt;
timestamp=utc_now_iso(),&lt;br&gt;
)&lt;/p&gt;

&lt;h1&gt;
  
  
  Ask AI Advisor, typed or spoken through the voice pipeline:
&lt;/h1&gt;

&lt;p&gt;remember_owner_statement(&lt;br&gt;
"sharma-general",&lt;br&gt;
"Don't keep more than 35 Maggi units in stock. Sharma Distributors "&lt;br&gt;
"has good prices but MOQ is 50 and delivery takes two days.",&lt;br&gt;
source="advisor chat",&lt;br&gt;
)&lt;br&gt;
Retain runs an LLM extraction pass over that sentence and turns it into entities (Maggi, Sharma Distributors), facts (a stock&lt;br&gt;
ceiling, an MOQ, a lead time) and time-stamped records. I did not have to decide in advance that "stock ceiling" was a concept.&lt;br&gt;
The part I did not expect to matter: language handling. Owners mix Hindi and English inside a single sentence, and the&lt;br&gt;
transcription layer hands me Hinglish text. I feared I would need my own normalization step. In practice I pass the text&lt;br&gt;
through as spoken and let retention and recall deal with it, since Hindsight preserves the input language rather than&lt;br&gt;
translating it.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Recall before every recommendation&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The reorder path is where memory earns its keep. Before Gemini sees anything, the orchestrator gathers the numbers from&lt;br&gt;
Postgres and the forecaster, then asks Hindsight what it knows that is relevant to this product.&lt;br&gt;
2 / 6&lt;br&gt;
async def build_reorder_context(store_id: str, product: Product) -&amp;gt; ReorderContext:&lt;br&gt;
stock = await inventory.current_stock(store_id, product.id)&lt;br&gt;
forecast = await forecaster.predict(store_id, product.id, horizon_days=7)&lt;br&gt;
memories = hs.recall(&lt;br&gt;
bank_id=f"store-{store_id}",&lt;br&gt;
query=f"What should I know before reordering {product.name}? "&lt;br&gt;
f"Owner limits, supplier terms, past reorders, upcoming events.",&lt;br&gt;
)&lt;br&gt;
return ReorderContext(&lt;br&gt;
stock=stock,&lt;br&gt;
forecast=forecast, # numbers: from the ML layer only&lt;br&gt;
memories=[m.text for m in memories.results], # context: from Hindsight&lt;br&gt;
product=product,&lt;br&gt;
)&lt;br&gt;
Recall runs several retrieval strategies in parallel (semantic, keyword, graph and temporal) and merges them, which is the&lt;br&gt;
reason I stopped hand-tuning a vector search over my own notes table. The temporal part turned out to matter a lot. "Festival&lt;br&gt;
in five days" and "supermarket opened in September" are different kinds of memory, and a query about reordering wants both.&lt;br&gt;
Here is one real path through the system, using the numbers from my test store:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Stock is 18 units. Recent sales run at 25 a day. A festival is five days out.&lt;/li&gt;
&lt;li&gt;The forecaster says 32 a day at 0.78 confidence, trending up.&lt;/li&gt;
&lt;li&gt;Hindsight returns the owner's 35-unit ceiling, Sharma's MOQ of 50, a preference for Sharma on price, and a note about a
previous overstock.&lt;/li&gt;
&lt;li&gt;Gemini reasons over all of that and explains, in plain Hinglish, that ordering 50 from Sharma would blow through the
owner's limit, and suggests 25 from an alternate supplier and watching demand next week.&lt;/li&gt;
&lt;li&gt;The owner orders 25.
Nothing in step 4 required the model to compute a single number. It read a forecast, read some memories, and wrote an
explanation.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo6ilqz45o414so6pxsh7.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo6ilqz45o414so6pxsh7.jpeg" alt=" " width="800" height="433"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Keeping the LLM away from the math
&lt;/h2&gt;

&lt;p&gt;The failure I feared most was a fluent, confident recommendation with a wrong number inside it. So the numbers Gemini is&lt;br&gt;
allowed to propose get clamped by ordinary code before they reach the owner.&lt;br&gt;
3 / 6&lt;br&gt;
def clamp_order(proposed: int, ctx: ReorderContext) -&amp;gt; tuple[int, list[str]]:&lt;br&gt;
notes = []&lt;br&gt;
ceiling = ctx.owner_max_stock # parsed from memories or set by the owner in settings&lt;br&gt;
if ceiling is not None:&lt;br&gt;
room = max(ceiling - ctx.stock, 0)&lt;br&gt;
if proposed &amp;gt; room:&lt;br&gt;
notes.append(f"capped at {room}: owner ceiling is {ceiling}")&lt;br&gt;
proposed = room&lt;br&gt;
if ctx.supplier and proposed and proposed &amp;lt; ctx.supplier.moq:&lt;br&gt;
notes.append(f"below {ctx.supplier.name} MOQ of {ctx.supplier.moq}")&lt;br&gt;
return proposed, notes&lt;br&gt;
This is deliberately boring. Hindsight supplies the qualitative constraint; deterministic code enforces it. I use the same split&lt;br&gt;
everywhere: the memory layer does not replace the forecast model, does not hold transactions, and does not calculate. It stores&lt;br&gt;
context, experience and qualitative knowledge so the reasoning step can make a decision that fits this particular shop.&lt;br&gt;
I resisted a lot of temptation to blur that boundary. Early on I let the LLM say "you'll run out in about three days" from memory&lt;br&gt;
of past sales. It was sometimes right. A stock-days-remaining figure now comes from the inventory table, always, and the&lt;br&gt;
dashboard shows it (for example, 3.2 days remaining under the Maggi card).&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnwmwqcv8srfuq2tqvnza.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnwmwqcv8srfuq2tqvnza.jpeg" alt=" " width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing the loop: decisions and outcomes
&lt;/h2&gt;

&lt;p&gt;Recall gets you context. The part that made the assistant feel like it was learning about one shop was writing the outcome&lt;br&gt;
back.&lt;br&gt;
def record_decision(store_id, product, recommended, ordered, reason):&lt;br&gt;
hs.retain(&lt;br&gt;
bank_id=f"store-{store_id}",&lt;br&gt;
content=(f"Reorder decision for {product.name}: recommended {recommended}, "&lt;br&gt;
f"owner ordered {ordered}. Reason: {reason}."),&lt;br&gt;
context="reorder decision",&lt;br&gt;
timestamp=utc_now_iso(),&lt;br&gt;
)&lt;br&gt;
def record_outcome(store_id, product, week_result):&lt;br&gt;
hs.retain(&lt;br&gt;
bank_id=f"store-{store_id}",&lt;br&gt;
content=(f"Outcome for {product.name} reorder: {week_result.summary}"),&lt;br&gt;
context="reorder outcome",&lt;br&gt;
timestamp=utc_now_iso(),&lt;br&gt;
)&lt;br&gt;
4 / 6&lt;br&gt;
A week after the owner orders 25, a job compares stock and sales against what we expected and retains a line like "sales met&lt;br&gt;
demand, no excess stock." The next time the same product comes up, that experience is among the things recalled: last time&lt;br&gt;
we ordered less than the supplier's MOQ and it worked.&lt;br&gt;
Because Hindsight consolidates related facts into observations, repeated evidence strengthens a belief rather than piling up as&lt;br&gt;
duplicates. Three uneventful reorders of the same size should read as one supported belief, not three noisy notes. I did not&lt;br&gt;
build that consolidation, and I would not have wanted to.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The same pattern in the khata ledger&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The udhaar screen looks unrelated, but it runs on the same idea. The ledger shows dues per customer, days overdue and a&lt;br&gt;
"Send Reminder" action that goes out over WhatsApp. The numbers are Postgres. What Hindsight adds is how to use them:&lt;br&gt;
"Ramesh usually buys groceries at the beginning of the month," "call in the evening," "regular buyer." A reminder for a&lt;br&gt;
customer twelve days overdue with a ₹5,000 credit limit reads differently from one for a customer eighteen days overdue&lt;br&gt;
whom you should call rather than message. I still let the owner press the button. The system just drafts it better.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Felh74xjg4fi0zc6ubzwh.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Felh74xjg4fi0zc6ubzwh.jpeg" alt=" " width="800" height="423"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would tell another engineer
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Split by what kind of truth it is. Transactions and counts belong in a database. Predictions belong to a model. Context and
history belong to memory. Every time I let one layer do another's job, I got a subtle bug, usually a plausible one.&lt;/li&gt;
&lt;li&gt;Do not design a schema for things owners say. I lost weeks adding columns for constraints I only half understood.
Retaining the raw statement and recalling by question was less code and covered cases I never listed.&lt;/li&gt;
&lt;li&gt;Enforce constraints outside the LLM, explain them inside it. Memory tells the model that the ceiling is 35. Code makes
sure a recommendation cannot exceed it. The model's job is to say why in language the owner trusts.&lt;/li&gt;
&lt;li&gt;Write decisions and outcomes, not only facts. The value of long-term memory grew once the bank held what we
recommended, what the owner did, and what happened. Facts alone gave me a better chatbot. Facts plus outcomes gave me
something closer to a colleague who remembers last month.&lt;/li&gt;
&lt;li&gt;Scope memory tightly. One bank per store, no cross-talk. For anything that might carry secrets or personal data, look at
Hindsight's memory defense policy before you retain, since owners do say things like account numbers out loud.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  What is still hard
&lt;/h2&gt;

&lt;p&gt;Recall quality depends on how you phrase the query, and I rewrote my reorder query several times. Owners contradict&lt;br&gt;
themselves, and an old preference ("keep 50 in stock") can outlive its usefulness; I want the advisor to ask before assuming the&lt;br&gt;
newer statement wins. And I still want a clearer view into which memories actually changed a recommendation, so the owner&lt;br&gt;
can see and correct them.&lt;br&gt;
If you are building an agent where the interesting knowledge lives in a person's head rather than your database, start with the&lt;br&gt;
Hindsight documentation and the overview of agent memory. The retain, recall, reflect trio maps onto more of a real business&lt;br&gt;
than I expected, and it kept my forecast honest.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>machinelearning</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
