Last quarter my forecasting model told a shopkeeper to order 50 cartons of noodles from a supplier
whose minimum order quantity was 50. The owner had told my system, weeks earlier, that he never
wanted more than 35 units of that product on the shelf. The model was right about demand and wrong
about the shop.
That gap between what the numbers say and what the owner has already decided is the problem I spent most of this project
on. This is how I closed it with Hindsight, an open-source agent memory system, and what I would do differently.
What the system does
DukaanPulse is an operations console for small Indian retail shops (kirana and general stores). The owner gets a dashboard
with today's sales, profit, order count and low-stock items; billing by keyboard, voice or receipt scan; a khata (udhaar) ledger
for customer credit; an expenses ledger; and an advisor they can ask questions in Hinglish.
The backend is a FastAPI service that acts as an orchestration layer. It has four stores of knowledge behind it, and I was strict
about what belongs where:
PostgreSQL holds structured facts: products, inventory, sales and purchases, suppliers, customers and khata balances,
expenses, orders, audit logs.
The ML layer (LightGBM/XGBoost) produces demand forecasts, adapts them online to recent sales, flags anomalies, and
folds in weather, holidays and price changes
Hindsight holds long-term business memory: owner preferences, supplier conditions, customer patterns, business events,
and past decisions with their outcomes.
Gemini does the reasoning and the conversation, using everything above.
The rule I wrote on the whiteboard, and mostly kept: Postgres stores what happened, the model predicts what will happen,
Hindsight remembers what the owner knows and what we tried, and the LLM never does arithmetic.
The through-line: a forecast is not a decision
A demand forecast is a number. A reorder decision is a number filtered through constraints that never appear in your sales
table:
"Don't keep more than 35 Maggi units in stock." (the owner's shelf, cash and habits)
"Sharma Distributors: good prices, MOQ 50, two-day delivery." (a supplier relationship)
"This product sells faster during local festivals." (something learned by watching)
"A supermarket opened nearby in September 2026." (an event that changes everything after it)
My first version put all of this into Postgres columns. max_stock_units on products, moq on suppliers. That worked for exactly
the constraints I had thought of in advance. Owners do not talk in schema. They say "Sharma is fine but his delivery has been
late" in the middle of a voice command, and that sentence has nowhere to go in a relational model.
So I stopped trying to anticipate the fields and started treating the owner's conversation as a source of memory.
Retaining what the owner says
Every owner conversation, and every event I can detect, goes through Hindsight's retain . I use one bank per store. Banks are
strictly isolated, which matters when one profile manages several stores.
from hindsight_client import Hindsight
hs = Hindsight(base_url=settings.HINDSIGHT_URL)
def remember_owner_statement(store_id: str, text: str, source: str):
hs.retain(
bank_id=f"store-{store_id}",
content=text,
context=f"owner {source}", # "voice command", "advisor chat", ...
timestamp=utc_now_iso(),
)
Ask AI Advisor, typed or spoken through the voice pipeline:
remember_owner_statement(
"sharma-general",
"Don't keep more than 35 Maggi units in stock. Sharma Distributors "
"has good prices but MOQ is 50 and delivery takes two days.",
source="advisor chat",
)
Retain runs an LLM extraction pass over that sentence and turns it into entities (Maggi, Sharma Distributors), facts (a stock
ceiling, an MOQ, a lead time) and time-stamped records. I did not have to decide in advance that "stock ceiling" was a concept.
The part I did not expect to matter: language handling. Owners mix Hindi and English inside a single sentence, and the
transcription layer hands me Hinglish text. I feared I would need my own normalization step. In practice I pass the text
through as spoken and let retention and recall deal with it, since Hindsight preserves the input language rather than
translating it.
Recall before every recommendation
The reorder path is where memory earns its keep. Before Gemini sees anything, the orchestrator gathers the numbers from
Postgres and the forecaster, then asks Hindsight what it knows that is relevant to this product.
2 / 6
async def build_reorder_context(store_id: str, product: Product) -> ReorderContext:
stock = await inventory.current_stock(store_id, product.id)
forecast = await forecaster.predict(store_id, product.id, horizon_days=7)
memories = hs.recall(
bank_id=f"store-{store_id}",
query=f"What should I know before reordering {product.name}? "
f"Owner limits, supplier terms, past reorders, upcoming events.",
)
return ReorderContext(
stock=stock,
forecast=forecast, # numbers: from the ML layer only
memories=[m.text for m in memories.results], # context: from Hindsight
product=product,
)
Recall runs several retrieval strategies in parallel (semantic, keyword, graph and temporal) and merges them, which is the
reason I stopped hand-tuning a vector search over my own notes table. The temporal part turned out to matter a lot. "Festival
in five days" and "supermarket opened in September" are different kinds of memory, and a query about reordering wants both.
Here is one real path through the system, using the numbers from my test store:
- Stock is 18 units. Recent sales run at 25 a day. A festival is five days out.
- The forecaster says 32 a day at 0.78 confidence, trending up.
- Hindsight returns the owner's 35-unit ceiling, Sharma's MOQ of 50, a preference for Sharma on price, and a note about a previous overstock.
- Gemini reasons over all of that and explains, in plain Hinglish, that ordering 50 from Sharma would blow through the owner's limit, and suggests 25 from an alternate supplier and watching demand next week.
- The owner orders 25. Nothing in step 4 required the model to compute a single number. It read a forecast, read some memories, and wrote an explanation.
Keeping the LLM away from the math
The failure I feared most was a fluent, confident recommendation with a wrong number inside it. So the numbers Gemini is
allowed to propose get clamped by ordinary code before they reach the owner.
3 / 6
def clamp_order(proposed: int, ctx: ReorderContext) -> tuple[int, list[str]]:
notes = []
ceiling = ctx.owner_max_stock # parsed from memories or set by the owner in settings
if ceiling is not None:
room = max(ceiling - ctx.stock, 0)
if proposed > room:
notes.append(f"capped at {room}: owner ceiling is {ceiling}")
proposed = room
if ctx.supplier and proposed and proposed < ctx.supplier.moq:
notes.append(f"below {ctx.supplier.name} MOQ of {ctx.supplier.moq}")
return proposed, notes
This is deliberately boring. Hindsight supplies the qualitative constraint; deterministic code enforces it. I use the same split
everywhere: the memory layer does not replace the forecast model, does not hold transactions, and does not calculate. It stores
context, experience and qualitative knowledge so the reasoning step can make a decision that fits this particular shop.
I resisted a lot of temptation to blur that boundary. Early on I let the LLM say "you'll run out in about three days" from memory
of past sales. It was sometimes right. A stock-days-remaining figure now comes from the inventory table, always, and the
dashboard shows it (for example, 3.2 days remaining under the Maggi card).
Closing the loop: decisions and outcomes
Recall gets you context. The part that made the assistant feel like it was learning about one shop was writing the outcome
back.
def record_decision(store_id, product, recommended, ordered, reason):
hs.retain(
bank_id=f"store-{store_id}",
content=(f"Reorder decision for {product.name}: recommended {recommended}, "
f"owner ordered {ordered}. Reason: {reason}."),
context="reorder decision",
timestamp=utc_now_iso(),
)
def record_outcome(store_id, product, week_result):
hs.retain(
bank_id=f"store-{store_id}",
content=(f"Outcome for {product.name} reorder: {week_result.summary}"),
context="reorder outcome",
timestamp=utc_now_iso(),
)
4 / 6
A week after the owner orders 25, a job compares stock and sales against what we expected and retains a line like "sales met
demand, no excess stock." The next time the same product comes up, that experience is among the things recalled: last time
we ordered less than the supplier's MOQ and it worked.
Because Hindsight consolidates related facts into observations, repeated evidence strengthens a belief rather than piling up as
duplicates. Three uneventful reorders of the same size should read as one supported belief, not three noisy notes. I did not
build that consolidation, and I would not have wanted to.
The same pattern in the khata ledger
The udhaar screen looks unrelated, but it runs on the same idea. The ledger shows dues per customer, days overdue and a
"Send Reminder" action that goes out over WhatsApp. The numbers are Postgres. What Hindsight adds is how to use them:
"Ramesh usually buys groceries at the beginning of the month," "call in the evening," "regular buyer." A reminder for a
customer twelve days overdue with a ₹5,000 credit limit reads differently from one for a customer eighteen days overdue
whom you should call rather than message. I still let the owner press the button. The system just drafts it better.
What I would tell another engineer
- Split by what kind of truth it is. Transactions and counts belong in a database. Predictions belong to a model. Context and history belong to memory. Every time I let one layer do another's job, I got a subtle bug, usually a plausible one.
- Do not design a schema for things owners say. I lost weeks adding columns for constraints I only half understood. Retaining the raw statement and recalling by question was less code and covered cases I never listed.
- Enforce constraints outside the LLM, explain them inside it. Memory tells the model that the ceiling is 35. Code makes sure a recommendation cannot exceed it. The model's job is to say why in language the owner trusts.
- Write decisions and outcomes, not only facts. The value of long-term memory grew once the bank held what we recommended, what the owner did, and what happened. Facts alone gave me a better chatbot. Facts plus outcomes gave me something closer to a colleague who remembers last month.
- Scope memory tightly. One bank per store, no cross-talk. For anything that might carry secrets or personal data, look at Hindsight's memory defense policy before you retain, since owners do say things like account numbers out loud.
What is still hard
Recall quality depends on how you phrase the query, and I rewrote my reorder query several times. Owners contradict
themselves, and an old preference ("keep 50 in stock") can outlive its usefulness; I want the advisor to ask before assuming the
newer statement wins. And I still want a clearer view into which memories actually changed a recommendation, so the owner
can see and correct them.
If you are building an agent where the interesting knowledge lives in a person's head rather than your database, start with the
Hindsight documentation and the overview of agent memory. The retain, recall, reflect trio maps onto more of a real business
than I expected, and it kept my forecast honest.




Top comments (0)