On the 12th of July the sentiment score of our export feature decreased by 41 points in one day. No one on the engineering team realized what was happening until 3 days later, when a support lead manually pasted 19 similar tickets in a slack channel. By that point 47 customers had already been impacted by the same crash.
This number of 47 tickets, the three day span, is what caused us to rethink customer feedback as a living memory system.
The Concrete Problem
Prior to this project, our feedback triage process looked roughly like this:
- Support tickets and tweets get dumped in Zendesk + shared slack channel
- PM or support lead eventually observes a pattern and writes a summary
- Engineering gets a vague Jira ticket (e.g., users are complaining about export)
- Hours spent re-reading the underlying tickets to understand severity and reproduction steps
The lag between when a customer had pain and when an engineer could take actionable steps on it was routinely 2-5 days. Classic RAG over tickets helped a little, but still required a human to ask the right questions, to surface an actionable insight based on the tickets. We wanted something that could continuously learn from the stream, form higher level observations, and then surface a mix of visual trends and specific evidence for human inspection.
Where and How We Integrated Hindsight
We leveraged Hindsight as the memory for this whole system. Every piece of incoming feedback goes through a light ingestion job that performs a retain() operation on Hindsight, which extracts facts, entities (feature name, version numbers, etc.), and begins forming observations. When the dashboard needs to answer a question or detect a cause, we use the recall() and reflect() operations, where the latter synthesizes a reasoned answer from the evidence.
It sits between the raw signal sources and the interactive dashboard like so:
Raw signals → retain() into Hindsight bank
↓
Hindsight (World + Experience + Observation networks)
↓
Nightly synthesizer + on-demand reflect()
↓
Structured Insights API → Interactive Dashboard
The dashboard never talks to the LLM directly for every user click, but instead consumes the observations and reflect results that Hindsight has synthesized.
Real Code from the Project
Here's a snippet of the code we use to perform the retain + reflect loop, which runs after every batch of new feedback:
from hindsight import HindsightClient
client = HindsightClient(bank_id="product-feedback-prod")
# After normalizing a batch of tickets/tweets
for item in new_feedback:
client.retain(
content=item["text"],
metadata={
"source": item["source"],
"timestamp": item["created_at"],
"user_id": item.get("user_id"),
"version": item.get("app_version"),
},
tags=["feedback", item["source"]]
)
# Later, when the synthesizer runs or a user asks a question
insight = client.reflect(
query="What caused the sharp drop in export-related sentiment after July 12?",
budget="high",
include_evidence=True
)
print(insight.answer)
# → "Export feature X began failing after the v2.1 release on July 11.
# 47 supporting reports mention crash or timeout. Highest confidence
# observation formed on July 13."
The same reflect() call is also used to power the clickable cause nodes on the sentiment chart, as well as the conversational "Query Your Feedback" UI.
Concrete Before / After
Before
Support lead notices pattern → writes Slack summary → creates vague Jira ticket → engineer spends 45-90 minutes hunting through original tickets for quotes and reproduction steps. Avg. time from first customer report to actionable engineering issue: 3.2 days.
After
Hindsight forms an observation overnight. The dashboard shows a red cause node on the sentiment chart the next morning. An engineer clicks it, sees the top quotes with source links, and hits “Create Draft GitHub Issue.” The issue is created with title, description, reproduction steps, and linked evidence in under 90 seconds. Avg. time from first report to actionable issue: under 18 hours (often same-day).
Honest Lesson and Limitation
The amount of the quality of the initial retain() metadata made a huge difference. When we first launched the product, we only retained the raw text. Hindsight still formed observations, but they were less precise. Adding more metadata (app version, source, etc.) improved the observation quality.
A big limitation that still stands today is that Hindsight's reflect() is good at synthesizing what has already been retained, but it can't invent new context. If customers never mentioned the version number, we can't magically know that the regression started with v2.1. We still have to enrich the signals in the ingestion layer, or have an explicit step to do so. Simply throwing raw text at Hindsight isn't enough for really high-stakes root-cause detection.


Top comments (0)