DEV Community

kondapalli Saipranay
kondapalli Saipranay

Posted on

From Raw Chat Logs to Customer Stories using Hindsight.

Introduction

When building AI tools for customer support, one of the primary technical challenges is handling long-term memory.

Customer interactions occur over days, weeks, or months across different channels. A customer might report a shipping delay on Monday, express frustration over product setup on Wednesday, and request a refund by Friday.

Standard large language model (LLM) implementations struggle with this pattern. Passing an ever-expanding array of raw chat transcripts into every prompt consumes excessive tokens and makes context extraction noisy. Conversely, stateless LLM calls forget past customer issues entirely.

To solve this, I built CustomerStory AI — an application that captures unstructured support notes, retains them in a persistent episodic memory layer using Hindsight, and synthesizes them into actionable customer background stories using FastAPI, Groq, and React.

"CustomerStory AI Dashboard" (https://dev-to-uploads.s3.us-east-2.amazonaws.com/uploads/articles/k4z18o7jl3wrbzkkm36h.png)

The application allows customer interactions to be stored as memories and later recalled to generate a concise customer story.

In this article, I will walk through the architecture, memory retention pipeline, prompt grounding strategy, and real-world frontend adjustments required to integrate Hindsight into a full-stack local setup.


Technical Stack and Architecture Overview

The system is structured as a decoupled full-stack application:

  • Backend: FastAPI (Python 3.11+) handling REST API endpoints, memory operations, and LLM orchestration.
  • Memory Engine: Hindsight running locally at "http://localhost:8888", accessed through the "hindsight-client" Python SDK.
  • LLM Inference: Groq API using the "openai/gpt-oss-20b" model for prompt completion.
  • Frontend: React application built with TypeScript and Vite, featuring a dashboard that highlights customer stories, timelines, and analytical insights.

System Architecture

"CustomerStory AI System Architecture" (https://dev-to-uploads.s3.us-east-2.amazonaws.com/uploads/articles/61kroc5qfz690akrnoo0.png)

The overall request flow is:

React Frontend
↓
FastAPI Backend
↙ ↘
Hindsight Groq

The React frontend communicates with FastAPI, which coordinates both long-term memory operations through Hindsight and LLM generation through Groq.


Retaining and Recalling Customer Memories

Rather than managing custom vector embeddings or manually querying a relational database for past notes, I integrated Hindsight as a long-term memory engine.

In my backend service layer ("services/memory.py"), I instantiated the Hindsight client targeting a local instance:

from hindsight_client import Hindsight

client = Hindsight(
base_url="http://localhost:8888"
)

BANK_ID = "customer-story"

def remember_customer(interaction: str):
return client.retain(
bank_id=BANK_ID,
content=interaction,
context="customer support interaction"
)

def recall_customer(query: str):
return client.recall(
bank_id=BANK_ID,
query=query
)

When a user submits a customer interaction note in the application, FastAPI processes the payload in "/customer/message" and passes the text directly to "remember_customer()":

@app.post("/customer/message")
def save_customer_message(data: CustomerMessage):
interaction = f"Customer {data.customer_name} said: {data.message}"
remember_customer(interaction)

return {
    "message": "Customer interaction saved successfully!",
    "customer": data.customer_name
}
Enter fullscreen mode Exit fullscreen mode

By storing interaction snippets inside the ""customer-story"" memory bank with explicit context (""customer support interaction""), Hindsight can retain the semantic content for later retrieval.

When I need to review a customer's history, calling "recall_customer()" returns relevant memories based on natural-language queries without requiring the backend to reload and parse historical chat logs manually.


Grounding LLM Summaries

Generating customer background stories requires turning raw recalled factual fragments into a readable paragraph.

A common risk with LLM summarization is hallucination, where the model introduces assumed facts or invents details that are not present in the customer's history.

To reduce this risk in CustomerStory AI, I structured the "/customer/{customer_name}/summary" endpoint to retrieve relevant memories from Hindsight first, then pass those memories to Groq ("openai/gpt-oss-20b") with explicit prompt constraints and a low temperature ("0.2"):

@app.get("/customer/{customer_name}/summary")
def get_customer_summary(customer_name: str):
query = f"What is the complete story of customer {customer_name}?"
results = recall_customer(query)

memories = [memory.text for memory in results.results]
memory_text = "\n".join(memories)

prompt = f"""
Enter fullscreen mode Exit fullscreen mode

You are a customer support assistant.

Create a short, clear customer story from the information below.

Customer: {customer_name}

Memory:
{memory_text}

Include:

  • important past problems
  • what made the customer unhappy
  • preferences
  • recent issues

Do not invent any information.
Keep the answer to 3-4 sentences.
"""

response = groq_client.chat.completions.create(
    model="openai/gpt-oss-20b",
    messages=[
        {
            "role": "user",
            "content": prompt
        }
    ],
    temperature=0.2
)

summary = response.choices[0].message.content

return {
    "customer": customer_name,
    "summary": summary
}
Enter fullscreen mode Exit fullscreen mode

The important idea is that the LLM is not asked to reconstruct the customer's history from an entire raw conversation archive. Instead, Hindsight first retrieves relevant memories, and those memories become the grounding context for the LLM.


Frontend Integration

The frontend dashboard is built in React with TypeScript and Vite.

It includes components such as:

  • "CustomerStoryCard" — displays the AI-generated customer story.
  • "CustomerInsights" — displays important customer insights.
  • "InteractionHistory" — displays recalled interactions chronologically.

"CustomerStory AI Dashboard — Customer Insights" (https://dev-to-uploads.s3.us-east-2.amazonaws.com/uploads/articles/n4b5ckg5p39eu4lsushj.png)

The dashboard provides a visual layer over the memory and summarization system, allowing support information to be understood without manually reading every historical interaction.


Handling Local Hindsight Concurrency

During local testing, I encountered an operational challenge.

When selecting a customer profile, the frontend initially issued concurrent requests to:

/customer/{name}/history
/customer/{name}/summary

using "Promise.all()".

Because both endpoints queried the underlying local Hindsight service at nearly the same time, concurrent requests occasionally caused database lock contention and 500 errors from the local Hindsight instance running on "localhost:8888".

To make the requests more reliable during local development, I changed the frontend logic to execute them sequentially:

const fetchCustomerData = useCallback(async (customerName: string) => {
if (!customerName.trim()) return;

setApiError(null);
setIsHistoryLoading(true);
setIsSummaryLoading(true);

// 1. Fetch History FIRST
try {
const historyData = await api.getHistory(customerName);
setHistory(historyData.history || []);
} catch (err) {
console.error('Failed to fetch customer history:', err);
setHistory([]);
} finally {
setIsHistoryLoading(false);
}

// 2. Fetch Summary SECOND
try {
const summaryData = await api.getSummary(customerName);
setSummary(summaryData.summary || '');
} catch (err) {
console.error('Failed to fetch customer summary:', err);
setSummary('');
} finally {
setIsSummaryLoading(false);
}
}, []);

The sequential approach avoided the concurrent requests that were triggering local Hindsight database lock conflicts during development.


Interaction History

One of the useful parts of the application is the ability to visualize the customer's recalled interactions as a timeline.

"Customer Interaction History Timeline" (images/interaction_timeline.png)

Instead of treating every support conversation as an isolated event, the system can bring relevant historical interactions together and provide a continuous view of the customer's journey.


Current Setup and Practical Takeaways

In its current state, CustomerStory AI serves as a local proof-of-concept.

All customer memories are stored in a single Hindsight bank:

BANK_ID = "customer-story"

The backend routes communicate with a locally running Hindsight instance, while Groq provides the LLM inference layer.

Building this project highlighted several practical lessons:

  1. Decoupled Memory Operations

Storing unstructured notes through "client.retain()" keeps memory ingestion separate from the summarization process.

  1. Grounded LLM Generation

Retrieving relevant memories before generating the summary provides the LLM with a focused context instead of the entire conversation history.

Explicit prompt constraints such as:

Do not invent any information.

also help keep the generated customer story tied to the available memory.

  1. Operational Awareness

Local infrastructure can introduce problems that are easy to overlook.

In this case, concurrent requests to the local Hindsight service caused database lock contention during development. Sequential frontend requests provided a simple solution for the local setup.

  1. Episodic Memory for AI Applications

The project demonstrates a practical pattern for applications that need to remember information across multiple interactions:

Customer Interaction
↓
FastAPI
↓
Hindsight
↓
Relevant Memories
↓
Groq LLM
↓
Customer Story
↓
React Dashboard


Conclusion

CustomerStory AI demonstrates how long-term episodic memory can be integrated into a modern AI application without continuously passing entire conversation histories to an LLM.

By combining Hindsight for persistent memory, FastAPI for backend orchestration, Groq for LLM inference, and React for the user interface, the application transforms individual customer support notes into a more continuous customer story.

The project also showed that building an AI application is not only about the model itself. Memory retrieval, prompt grounding, API orchestration, frontend behavior, and local infrastructure all play important roles in making the system reliable.

Tags:

ai
aidev
python
react

Top comments (0)