DEV Community

Cover image for Architecting Context-Aware LLMs: Decoupling State from Cognition
Nitesh Gupta
Nitesh Gupta

Posted on

Architecting Context-Aware LLMs: Decoupling State from Cognition

If you've spent any time building LLM-backed applications in production, you quickly run into the "institutional amnesia" problem. You build a beautifully orchestrated agent, it performs a complex reasoning task perfectly, and then... it completely forgets everything the moment the context window clears.

When you're building financial due diligence systems for institutional investors, "forgetting" isn't just an annoyance; it's a critical failure. You can spend thirty minutes meticulously correcting an LLM's assumptions about your data, only for it to revert to its baseline generic responses the very next time you open a session. If an analyst resolves a discrepancy on Monday ("The March ARR figure was an annualized projection, not live revenue"), the agent cannot be asking about that exact same discrepancy when a new ledger is uploaded on Friday.

I recently architected the backend for Chrimata, a financial intelligence platform, and curing this amnesia was our biggest roadblock. I want to share exactly how I solved this statefulness problem using Vectorize agent memory architecture and implementation via an embedded daemon, and how it fundamentally shifted our agent's behavior from generic to hyper-personalized.

The Limitation of Stateless LLMs

Most developers try to solve context limits by aggressively stuffing the prompt with recent conversation history or wiring up a naive RAG implementation that does keyword searches against a vector database.

This fails in complex workflows. A standard database stores what happened (e.g., ticket status = closed). Naive RAG stores documents. But neither stores the semantic reasoning behind why a decision was made.

To fix this, I integrated Hindsight—an open-source persistent memory layer for AI agents. We needed the system to learn from its past outcomes, acting as a persistent, vector-backed semantic brain that captures the context of past interactions.

Wiring Up Hindsight in Production

I wanted absolute cryptographic data privacy. Institutional diligence data cannot be shipped off to a third-party memory cloud. So, instead of using a cloud API, I initialized Hindsight as an embedded daemon directly within our infrastructure.

Here's a look at the adapter I built to interface with it:

from hindsight import HindsightEmbedded

class HindsightAdapter:
    def __init__(self):
        # We enforce cryptographic privacy by running the embedded daemon locally
        self.provider, self.api_key, self.base_url, self.model = _resolve_hindsight_provider()
        if self.api_key:
            self.client = HindsightEmbedded(
                profile="chrimata", 
                llm_provider=self.provider, 
                llm_api_key=self.api_key,
                llm_base_url=self.base_url,
                llm_model=self.model
            )

    async def retain_review(self, bank_id: str, document_id: str, text: str, meta: dict) -> bool:
        """Saves the raw text and metadata of an analyst's decision."""
        if not self.is_available():
            return False
        try:
            await self.client.aretain(bank_id=bank_id, document_id=document_id, content=text, metadata=meta)
            return True
        except Exception as e:
            print(f"Hindsight retain failed: {e}")
            return False
Enter fullscreen mode Exit fullscreen mode

Understanding the aretain vs retain methods

Notice the use of aretain instead of the synchronous retain. This was a hard-learned lesson. FastAPI runs synchronous path operations in a threadpool. Calling synchronous client methods was tripping cross-task timeout crashes. Switching to the async SDK ensures the memory layer runs safely on the primary event loop.

Decoupling State from Cognition

In our architecture, PostgreSQL acts as the State Machine (tracking what is open, closed, or uploaded), but Hindsight acts as the Cognitive Engine.

When an analyst makes a decision on the UI, it hits our API, updates Postgres, and triggers our memory service to retain the context.

    async def retain_decision(
        self, bank_id: str, document_id: str, text: str,
        entity: str, issue_id: str, metrics: List[str],
        period: str, decision: str, review_id: str,
    ) -> bool:
        """Retains an analyst decision with enough context for future recall."""
        meta = {
            "memory_type": "analyst_decision",
            "entity": entity,
            "issue_id": issue_id,
            "review_id": review_id,
            "metrics": metrics,
            "period": period,
            "decision": decision,
            "retained_at": self._now(),
        }
        return await self.adapter.retain_review(bank_id, document_id=document_id, text=text, meta=meta)
Enter fullscreen mode Exit fullscreen mode

Retaining the "Why"

Similarly, when an automated document request fails or succeeds, we explicitly retain that semantic outcome.

    async def update_request_status(
        self, db: Session, bank_id: str, request_id: str,
        status: str, outcome_note: Optional[str] = None,
    ) -> EvidenceRequestSchema:
        # ... standard DB state update logic ...

        # Retain outcome in Hindsight for future learning
        if status == "insufficient" and outcome_note:
            text = f"Evidence request for {req.issue_id} was insufficient. Reason: {outcome_note}. Original request: {req.request_text}"
            await self.memory_service.retain_evidence_request_outcome(
                bank_id, document_id=req.id, text=text,
                issue_type=req.requested_document_type or "unknown",
                request_summary=req.request_text,
                outcome=status, reason=outcome_note,
            )
Enter fullscreen mode Exit fullscreen mode

This snippet shows the shift in mindset. We aren't just updating a boolean flag in Postgres. We are saving a semantic explanation of what went wrong.

The Before and After: Evolving Responses

The behavioral shift happens when we query this memory before making the next request. For example, when a new document is uploaded (a "Change Review"), the backend pulls the relevant context via arecall() using a highly targeted query structure:

    async def recall_for_change_review(
        self, bank_id: str, entity: str, metrics: List[str],
        issue_types: List[str], period: Optional[str] = None,
    ) -> List[Any]:
        """Builds a rich memory retrieval query from structured investigation context."""
        query_parts = [entity] + metrics + issue_types
        if period:
            query_parts.append(period)
        query_parts.extend(["analyst decisions", "prior resolution", "previous discrepancy"])
        query = " ".join(query_parts)
        return await self.adapter.recall(bank_id, query)
Enter fullscreen mode Exit fullscreen mode

Contextual Inference Generation

By doing this, the analyzer knows exactly why a document was updated based on past analyst reviews, heavily reducing false-positive alerts on trivial updates.

When generating the next evidence request, the same principle applies. We inject the recalled memory straight into the system prompt:

        # Recall prior request outcomes from Hindsight
        prior_outcomes = await self.memory_service.recall_for_evidence_request(
            bank_id, issue_type=metric, metric=metric, period=period,
        )

        prompt = (
            f"Issue: {issue.question}\n"
            f"Missing evidence type: {doc_type}\n"
            f"Missing fields: {fields_str}\n"
        )
        if prior_outcomes:
            prompt += f"\nPrevious evidence request outcomes (learn from these):\n"
            for po in prior_outcomes[:3]:
                prompt += f"- {po}\n"
            prompt += "\nAvoid repeating mistakes from prior insufficient requests.\n"
Enter fullscreen mode Exit fullscreen mode

Interaction 1 (No Memory):
System: "Please provide the Q1 subscription ledger showing active subscriptions and activation dates."
(Result: The client uploads a ledger, but it excludes the EMEA region).

Interaction 5 (With Memory):
System: "Please provide the Q1 subscription ledger. Note: The previous ledger submitted was insufficient because it excluded EMEA data. Please ensure all global regions are included to verify the ARR."

The difference is staggering. The agent is no longer an isolated script; it feels like a seasoned analyst who remembers the friction points of this specific deal.

Synthesizing vs. Stuffing

For internal automated pipelines, we use recall to grab raw structured objects and parse them into specific prompt variables. But we also expose an interactive chat agent. Injecting 50 raw JSON memories into the context window would destroy token limits and confuse the model.

Instead, I used Hindsight's reflect() method:

    async def reflect(self, bank_id: str, query: str) -> Optional[str]:
        """Synthesizes a markdown answer from consolidated memory."""
        try:
            res = await self.client.areflect(bank_id=bank_id, query=query, budget="low")
            return getattr(res, "text", None)
        except Exception:
            return None
Enter fullscreen mode Exit fullscreen mode

Synthesizing memories with areflect

reflect invokes Hindsight's internal LLM to dynamically synthesize a clean, consolidated Markdown summary of the memory bank before it reaches our primary chat model. When the user asks, "Why did we flag the churn rate last month?", the agent instantly pulls from the synthesized memory of the exact UI interactions and decisions the human analyst made.

Technical Takeaways

  1. State != Memory: Don't confuse your transactional database with your agent's cognitive layer. Postgres tracks state; Hindsight tracks context and reasoning. Retain the "Why", not just the "What".
  2. Local Daemons Win for Privacy: If you're building in fintech, healthcare, or legal, an embedded daemon keeps your cryptographic integrity intact while giving you the power of vector memory.
  3. Async is Mandatory: When wiring up I/O bound memory operations in Python (like Hindsight's arecall), always use the async SDKs to prevent blocking your API's event loop.
  4. Pre-prompt Injection is Magic: Injecting recalled memories directly into the system prompt transforms an LLM from a generic responder into a highly contextual participant.
  5. Synthesize Context: Use reflection/synthesis tools to pre-digest memory before dumping it into a Chat LLM's context window.

By giving our agent persistent object permanence, the entire user experience shifted from "managing an AI tool" to "collaborating with a colleague." If you are wrestling with context limits or building stateful agent workflows, I highly recommend checking out the comprehensive Hindsight documentation for persistent agent memory. It completely changed how I architect LLM backends.

Top comments (0)