DEV Community

Pragnav Rao
Pragnav Rao

Posted on

Memory Without Authority: Building Safer AI-Powered Treasury Reviews with Hindsight

 That sounds obvious, but it becomes critical when the workflow involves sensitive financial decisions. I built Project Arbitrage, a treasury review service, to explore this exact problem. I wanted to see how memory (via Hindsight) could give an agent useful historical context while keeping the actual authority with application logic and a human reviewer.

The main architectural rule behind the project: Let memory help the system understand what happened before, but never let memory decide what happens now.

Memory can connect the current situation with similar events from the past. But it should not replace a real market rate, a configured policy, or an explicit human approval. Here is how I built it.

A Treasury Review, End to End

The application takes current FX (Foreign Exchange) reference data, user-entered exposures, and policy thresholds, and turns them into a review that a human can inspect and act on.

The backend runs on a FastAPI service composed of:

  • Frankfurter: For dated reference-rate series and market inputs.
  • SQLite: For exposures, policies, assessment snapshots, and audit events.
  • Hindsight Cloud: For memory recall and retention.
  • Groq: To generate the written assessment.
  • Browser Frontend: For viewing the workflow and recording decisions.

The assessment follows a fixed sequence:

  1. Check if required exposure and policy information is available.
  2. Fetch market data, calculate risk locally, and evaluate policy rules.
  3. Compare the current FX movement with recent historical daily movements.
  4. Recall relevant memories from Hindsight and send the combined context to the reasoning step.

The final assessment is stored with its own ID, original inputs, and provenance, creating a clear audit trail.

When a reviewer makes a decision, the endpoint takes the stored assessment ID, rechecks policies and guardrails, and records the approval, escalation, or rejection. Only then is the human decision retained in Hindsight as another event connected to that assessment.

The system also features a simulation endpoint for protective actions. It records a proposed action but does not execute a real trade, hedge, payment, or transfer. The route explicitly marks the result as simulation_only, leaving the protected_amount unset. The system records what a person authorized without pretending it actually moved money.

Memory is Context, Not Authority

I gave Hindsight a very specific job: to remember historical market assessments and human decisions, along with enough metadata to understand where those memories came from.

Every retained event connects back to the project and the assessment that produced it. The Hindsight integration is isolated in one service rather than spread throughout the API routes.

The recall method looks roughly like this:

return await client.arecall(
    bank_id=self.bank_id,
    query=query,
    max_tokens=2000,
    budget="mid",
)

Enter fullscreen mode Exit fullscreen mode

The recall query focuses on the currency pair and current scenario. I intentionally kept the returned memory count low. The goal is a few pieces of highly relevant context, not an information dump.

The application also explicitly excludes memories marked as synthetic:

tagged_synthetic = any(str(tag).lower() == "synthetic" for tag in result_tags)
metadata_synthetic = str(result_metadata.get("synthetic", "")).lower() == "true"

if tagged_synthetic or metadata_synthetic:
    continue

Enter fullscreen mode Exit fullscreen mode

This is an application-level rule defining what records the workflow considers. The market snapshot and portfolio info still come from designated sourcesβ€”Hindsight only provides history.

The Guardrail Does Not Depend on Memory

One of the most important design choices was keeping the main numerical guardrail separate from semantic memory.

The current FX movement is compared directly against the worst move in the dated historical series returned by the market-data provider. This calculation happens in standard application code. Even if Hindsight returns no memories, the core numerical guardrail works flawlessly.

If historical series data is unavailable, the application doesn't guess. It flags the missing data and forces a manual review:

if not historical_events:
    return {
        "guardrail_status": "INSUFFICIENT_REFERENCE_DATA",
        "manual_override_required": True,
        "live_move_pct": round(live_move, 6),
        "worst_historical_move_pct": None,
        "precedent_variance_index": None,
        "safety_message": "No live historical FX series is available; human review is required without a precedent comparison.",
    }

Enter fullscreen mode Exit fullscreen mode

This separation highlights two distinct meanings of "history" in the system:

  • Numerical History: "How large was the movement?" (Handled by market data)
  • Contextual History: "What happened in a similar situation before, and how did the reviewer respond?" (Handled by semantic memory)

Durable Decisions Need a Local Record

Memory shouldn't be the only place where important decisions live. Hindsight is a network dependency. The local audit record shouldn't disappear just because the memory service experiences downtime.

For human decisions, the application writes the decision and audit information to SQLite first. Only then does it try to retain the event in Hindsight.

saved_decision, created = portfolio_data.record_decision(
    record.assessment_id, decision_payload
)

if created:
    hindsight_status = "retained"

    try:
        retention_result = hindsight_memory_service.retain_memory(
            decision_payload,
            context="Human decision on live Project Arbitrage assessment",
            tags=[
                "project-arbitrage",
                "human-decision",
                record.decision.lower(),
                "treasury",
            ],
            metadata={
                "assessment_id": record.assessment_id,
                "decision": record.decision,
            },
            live_mode=True,
        )
    except Exception:
        hindsight_status = "unavailable"

Enter fullscreen mode Exit fullscreen mode

If memory recall fails early on, the reasoning step just misses that context. But if memory retention fails after a decision is stored locally, the decision still safely exists. Idempotency is crucial here to ensure retries don't create duplicate audit records or memories.

What a Review Looks Like in Practice

Let's say the latest USD/INR reference rate moves 0.25% from the previous business day:

  1. A user has a saved $100,000 exposure, along with a hedge percentage and an FX policy threshold.
  2. The service calculates the risk, records the market source, and compares the 0.25% move against recent historical series.
  3. It asks Hindsight for relevant treasury memories and passes them to the reasoning step.
  4. The system produces an assessment linked to a specific ID.
  5. The reviewer inspects the data, memories, and generated analysis, then makes a decision.
  6. The decision is stored locally, then sent to Hindsight as a tagged event linked to that assessment ID.

If Hindsight doesn't find a matching memory, the system simply states there is no precedent. It doesn't invent one. Memory makes past context available, but application logic decides which facts matter.

5 Lessons I Would Reuse

Building Project Arbitrage crystallized a few core principles for AI agents:

  1. Give memory a narrow job. It retrieves and retains contextual events. It does not calculate exposure, enforce policy, or approve decisions.
  2. Keep hard rules in code. If something can be expressed as a deterministic rule (like a numerical threshold), keep it explicit and inspectable.
  3. Store provenance with the memory. A memory is infinitely more useful when you can answer: Where did this come from?
  4. Plan for provider failures. Recall, retention, and local persistence have different failure consequences. Treat them separately.
  5. Make retries safe. Idempotent handling prevents duplicate audit records and corrupted memory sets.

Top comments (0)