DEV Community

Pranavi
Pranavi

Posted on

I Made Hindustan Memory Survive a New Interaction

I Kept Hindsight and Groq Behind One Server Route
When a browser page needs both persistent memory and generated text, it is tempting to call both services from the client. That makes the UI code look direct, but it also puts secrets and orchestration decisions in the wrong place. I put the Hindsight and Groq calls behind one Next.js API route and let the browser send only a customer name and issue.
This is a small system, but the route is where its important responsibilities meet: input validation, customer-scoped recall, response generation, retention, and error handling. Keeping them together makes the sequence inspectable. It also makes the order consequential, because this endpoint crosses two services that can fail independently.
The browser sends a small request
The client owns transient interaction state. It validates that both fields have some text, marks the request as loading, and sends a JSON body to /api/memory:

const result = await fetch("/api/memory", {
  method: "POST",
  headers: {
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    customer: customer.trim(),
    issue: issue.trim(),
  }),
});
Enter fullscreen mode Exit fullscreen mode

The browser does not choose a Hindsight bank, construct customer tags, or build the Groq prompt. Those choices belong with the service credentials and the policy that decides what memory can be used. The server reads Hindsight and Groq credentials from environment variables, never from client-visible props.
The route validates JSON structure and field limits before making external calls. It rejects malformed bodies, non-object payloads, missing strings, and values above the configured limits of 100 characters for the customer name and 2,000 for the issue. That does not solve every abuse case, but it gives the API a well-defined input contract before it spends work on recall or generation.
The limits are deliberately ordinary. They stop obviously oversized values from flowing through every later stage, and they make the API behavior easier to explain to the client. Validation is not a substitute for rate limits, authentication, or authorization. Those concerns belong at the service boundary too, especially because a customer field controls which memory partition gets queried. The route is a reasonable place to validate shape; a deployed support tool also has to establish who is making the request and which customer they may access.
One orchestration point makes the sequence clear
After validation, the server reads the configured bank ID and recalls memory for the customer and issue. Hindsight receives a strict customer tag. The selected memory and current issue then go to Groq. Finally, the issue and generated response are retained with the same tag.

const memory = await recallCustomerMemory(bankId, customer, issue);

const previousMemory =
  memory?.results?.length > 0
    ? memory.results[0].text
    : `No previous memory found for ${customer}.`;

const completion = await groq.chat.completions.create({
  model: "openai/gpt-oss-120b",
  messages: [
    {
      role: "system",
      content: `
You are a professional customer support agent.

Use previous customer memory when available.

Rules:
- Be friendly and concise.
- Never invent customer history.
- Do not mention information belonging to another customer.
- If there is no previous memory, simply handle the current issue.
- Keep the response under 120 words.
`,
    },
    {
      role: "user",
      content: `
Customer: ${customer}

Current issue:
${issue}

Previous memory:
${previousMemory}

Write a personalized support response.
`,
    },
  ],
});
Enter fullscreen mode Exit fullscreen mode

The key point is the trust boundary: the browser cannot substitute its own “previous memory” field. The route retrieves the context itself and then builds the prompt.
Hindsight is more than a dependency called from a helper here. It is the persistent side of the route's contract. It lets the endpoint answer a new request with context from a completed request, and the retention call establishes the next request's starting point. The Hindsight open-source repository, Hindsight documentation, and Vectorize agent memory overview describe the memory model behind that role.
Architecture diagram showing Hindsight and Groq behind the Next.js API route
The browser sends the issue to the route; the route owns recall, generation, and retention.
The subtle failure is between generation and retention
The route waits for Groq to generate a response, then awaits Hindsight retention before returning success. That ordering preserves a useful invariant: a successful response from this API means the new interaction was retained for future recall.
It also creates a partial-failure case. Groq may generate a useful answer, then Hindsight retention may fail. The catch block returns an error, so the browser does not receive the generated answer even though inference already happened. A client retry could generate another answer and attempt another retention. There is no idempotency key in this route, so I would not assume retries are exactly-once.
For a production service I would make that contract explicit. One option is to treat response delivery and memory durability as separate outcomes, returning the answer while marking retention pending. Another is to persist a request record first and make memory retention retryable by idempotency key. The right answer depends on whether a support agent can safely reply without future memory being guaranteed. What I would not do is hide this partial-failure case behind the word “success.”
The inverse failure is possible too: retention may succeed, but the client may lose the response while the connection closes. A retry can then create another model call and another retained interaction. The route does not pass an idempotency key to Hindsight, so I would not assume retries are exactly-once. Before enabling automatic retries, I would pair them with a request identifier that survives the client retry and a clear policy for whether a repeated interaction should append, replace, or be rejected. The SDK exposes an operation identifier for idempotent asynchronous retries; the application still has to decide when to use it.
The UI reinforces the distinction between request state and durable state. loading is cleared in a finally block, so a network or API error does not leave the button stuck. On an API error, the message is shown in the response area. On success, the response and recalled memory are shown separately. Those states are kept in React; the customer history itself lives in Hindsight.

try {
  const result = await fetch("/api/memory", {
    method: "POST",
    headers: { "Content-Type": "application/json" },
    body: JSON.stringify({ customer: customer.trim(), issue: issue.trim() }),
  });
  const data = (await result.json()) as MemoryApiResponse;

  if (result.ok && data.success && data.memory && data.response) {
    setMemory(data.memory);
    setResponse(data.response);
  } else {
    setResponse(data.message || "Could not complete the support interaction.");
  }
} catch (error) {
  console.error(error);
  setResponse("Something went wrong. Please try again.");
} finally {
  setLoading(false);
}
Enter fullscreen mode Exit fullscreen mode

The important separation is that a visible answer is not the same as a retained memory, and each boundary deserves explicit failure behavior.
Keeping the calls on the server also gives the application one place to control timeouts, logging, and correlation IDs. The current route logs an error and returns a generic message, which avoids exposing service details to the browser. Operationally, I would attach a request ID and record which stage failed without logging API keys or unnecessary customer content. That makes a slow Hindsight recall distinguishable from a Groq error while leaving the client contract small.
What I would preserve as the system grows
I would keep the browser contract small even as the service grows. A UI that sends customer, issue, and a supplied “memory” string is not just convenient; it gives the caller control over the context the model trusts. The server should remain responsible for selecting memory and constructing the prompt.
I would also keep recall and retention as observable stages. If a response is wrong, I want to know whether no memory was found, the wrong item was retrieved, the model ignored a useful item, or the write failed. Those are different faults and should produce different operational signals.
Three lessons came out of this route:
Keep secrets and memory policy server-side. The browser submits the issue; the API chooses what Hindsight and Groq receive.
Make success semantics match the durable side effect. If retention is required for a successful interaction, say so and design for the partial-failure case.
Do not imply exactly-once behavior without idempotency. A timeout after generation can lead to a retry; external calls need a deliberate retry contract.
The route is the seam between a transient support interaction and persistent customer history. Hindsight gives that seam somewhere durable to write and a scoped way to read later. Keeping the orchestration on the server makes the memory boundary visible—and makes the consequences of a failed write something I can reason about instead of discovering from a browser console.

Top comments (0)