The Problem: Compliance Teams Keep Relearning the Same Lessons
Most compliance assistants can answer questions about a policy. The harder problem is answering, six months later, “What did the auditor reject last time, who owned the remediation, and is the evidence still valid?” I built this system around a simple idea: for compliance work, long-term memory is not a convenience feature; it is part of the system’s state.
What I built
The system is a conversational compliance assistant for a bank’s AI governance process. I wired an n8n workflow around an LLM agent and gave it a persistent compliance memory backed by Hindsight on GitHub.
recall_memory searches the organization’s long-term audit history.
reflect_on_history asks Hindsight to synthesize patterns across that history.
retain_memory stores newly reported facts permanently.
There is also short-lived session memory, but I deliberately kept that separate from organizational memory. The session buffer is there to make a conversation coherent. Hindsight is where the durable facts live.
The workflow then stores the completed conversation back into Hindsight as well. That last step matters more than it initially appears: an audit conversation can itself become evidence about a decision, an interpretation, or an auditor preference that needs to be available later.
The basic flow looks like this:
Chat message
|
v
Hindsight Auditor agent
|
+--> recall_memory --------> Hindsight
|
+--> reflect_on_history ---> Hindsight
|
+--> retain_memory --------> Hindsight
|
v
Answer
|
v
Remember Conversation -------> Hindsight
The repository also separates the interactive path from history ingestion. In the completed system, that ingestion path is where I connect authoritative audit records, control tests, remediation systems, and approved evidence metadata rather than a hand-authored history.
The important boundary is that the language model does not become the database. It is the reasoning layer sitting on top of persistent memory.
The design decision that changed the project
The first mistake I wanted to avoid was treating “memory” as a transcript.
A transcript tells me what somebody said. Compliance needs something closer to a structured institutional memory: findings, owners, dates, remediation status, control tests, evidence requirements, and the history of what an auditor accepted or rejected.
That is why the agent’s instructions explicitly require memory retrieval before answering questions about systems, audits, findings, remediations, evidence, auditors, owners, deadlines, or earlier conversations.
- ALWAYS call recall_memory ... before answering.
- When the user shares new information ... call retain_memory ...
- Answer from history ... name the AI system, finding ID, dates, owner, status and the auditor. This establishes a contract between the reasoning layer and Hindsight: factual continuity comes from memory, not from whatever the model happens to remember from its context window. That matters when the question is: “Are we ready for the next external audit?” A generic LLM can produce a perfectly plausible EU AI Act checklist. That is not the answer I need. The useful answer depends on institutional state: which findings are still open, which controls have gone too long without testing, whether a remediation missed its deadline, whether an owner left the organization, and what evidence a particular auditor previously rejected. That is exactly the kind of workload where I found Hindsight useful. Its recall operation handles focused retrieval, while reflect gives me a separate path for questions that require synthesis across the memory rather than retrieval of one fact. The workflow makes that distinction explicit: // Focused memory lookup JSON.stringify({ query: $fromAI( "query", "Focused search query, e.g. credit scoring bias test findings and remediation status", "string" ), budget: "high", max_tokens: 6000 }) And for broader questions: JSON.stringify({ query: $fromAI( "question", "Full synthesising question to answer from the compliance memory", "string" ), budget: "high" }) I like having both paths because “find the record about CreditScore-X” and “identify recurring weaknesses across the audit history” are fundamentally different retrieval problems. I made the memory write path deliberately explicit Reading memory is only half of the system. The other half is deciding what deserves to survive the current conversation. When a compliance lead reports a new finding, remediation update, owner change, policy decision, or auditor preference, the agent writes it as a self-contained fact. The workflow passes both content and context to Hindsight: JSON.stringify({ async: true, items: [{ content: $fromAI( "content", "Self-contained compliance fact with system, dates, owner and status", "string" ), context: $fromAI( "context", "Category such as audit finding, remediation update, evidence decision, auditor preference, policy change", "string" ), timestamp: $now.toISO(), tags: ["compliance", "user-reported"] }] }) That “self-contained” requirement matters because future retrieval should not depend on reconstructing context from an old chat turn. For example, storing: “It’s overdue.” is almost useless six months later. Storing a fact with the system, finding, promised date, owner, and current status gives the retrieval layer something that can stand on its own. I also store the completed conversation separately: content: "Conversation with the compliance lead on " + $now.toFormat("yyyy-MM-dd HH:mm") + ". User asked: " + $("When chat message received").item.json.chatInput + " | Hindsight Auditor answered: " + $json.output That gives me two useful layers of history: durable facts extracted from the conversation and the conversation itself. Why Hindsight matters here I could have built this with a vector database and a pile of retrieval prompts. That would have solved only part of the problem. The interesting part of Hindsight’s agent memory model is that memory is treated as something an agent can recall, retain, and reason over across interactions. That maps unusually well to compliance because the source material is temporal and relational even when the user asks a simple question. The system needs to connect facts like: a control was tested on one date; an auditor later changed the evidence requirement; a remediation was promised for another date; the owner subsequently left; the replacement implementation still has not been tested. None of those facts is especially complicated by itself. The difficulty is preserving the chain. The Hindsight documentation describes the memory interface I am using here: recall for retrieval, retain for durable storage, and reflect for higher-level reasoning across memory. In this workflow, those operations become explicit tools rather than hidden behavior inside the model. That separation also makes the system easier to reason about operationally. If an answer is wrong, I can ask a concrete question: did we retrieve the relevant memory, did we retain the new fact, or did the model reason incorrectly from the retrieved state? That is a much better debugging boundary than “the AI got confused.” The audit history is useful because it is deliberately uncomfortable The seeded history contains the sort of inconsistencies that make compliance work difficult. For example, the history records a CreditScore-X bias finding with a remediation deadline of June 30, 2026. Later records say the remediation is still open, the reweighting exists only in development, and the required re-test has not happened. It also records an owner change: the original owner left the bank, and no formal replacement had been assigned. Then there is the evidence problem. The external auditor had previously rejected screenshots of the fairness dashboard and required raw bias-test logs, dataset and model identifiers, the calculation notebook, and a signed owner statement. Those details are precisely what I want the assistant to recover rather than regenerate from general compliance knowledge. The agent instructions therefore tell it to flag:
- anything untested for more than 9 months
- any remediation past its promised date
- any owner who has left
- any evidence format the auditor previously rejected That turns memory retrieval into an operational check. A readiness question is also constrained to a fixed response structure: Readiness verdict Blocking gaps Overdue remediations Controls due for re-testing Auditor preferences to respect Recommended next steps What an interaction looks like Consider the question: “What do I need to fix before Helena Brandt’s next audit?” The assistant should not answer with a generic list of EU AI Act requirements. It should recover that the next external audit is scheduled for October 6, identify the systems in scope, find the open findings associated with that auditor, and surface the evidence requirements that were established during previous audits. “What is still unresolved on CreditScore-X?” The relevant memory chain includes the bias finding, the original remediation owner, the owner’s departure, the missed June deadline, the development-only reweighting, and the missing re-test. “What evidence should I prepare for the fairness test?” The answer should include the auditor-specific requirement rather than blindly suggesting a dashboard screenshot, because the memory contains an explicit rejection of that format. This is where persistent memory changes the character of the assistant. Without it, every conversation starts from the model’s general knowledge. With it, the assistant can operate against the organization’s accumulated decisions. The workflow is intentionally boring in the right places The surrounding implementation is straightforward n8n orchestration. I use gpt-oss-120b as the primary model and qwen3-32b as the fallback, both configured with low temperature. The agent has a bounded session memory of ten messages. HTTP tool nodes connect the agent to Hindsight, and the completed exchange is persisted after the answer. None of that is particularly exotic. I prefer it that way. The interesting engineering work is defining what the model may infer from memory and what it must retrieve first. Lessons I would reuse elsewhere
- Long-term memory should have a schema, even when the interface is natural language I do not need every fact represented as a rigid relational row, but I do need consistent semantic fields: system, date, owner, status, context, and evidence state. That makes natural-language retrieval much more useful.
- Retrieval policy is part of application logic “Always recall before answering” is not a prompt-writing trick. It is an application invariant. If a question depends on organizational history, the current answer should be grounded in that history.
- Store decisions, not just documents The most valuable memories are often small: an auditor rejected a particular evidence format, an owner changed, a remediation date moved, or a control was last tested on a specific date. Those facts are easy to lose in document-oriented systems and disproportionately useful during the next audit.
- Separate session context from institutional memory A ten-message conversation window is useful for dialogue. It is not a compliance record. Keeping the two separate lets me tune conversational context without accidentally treating temporary chat context as authoritative organizational history.
- Make memory observable The recall/retain/reflect boundary gives me a useful debugging model. When the assistant produces a bad answer, I can inspect whether the right memory was retrieved, whether a new fact was stored correctly, and whether synthesis introduced the error. That is much easier to operate than a single opaque “AI memory” layer. Closing The part of this project I would keep even if I replaced the model, orchestration layer, or chat interface is the memory architecture. Compliance work is fundamentally cumulative. The answer to today’s question often depends on what an auditor said months ago, what evidence was accepted, which control was last tested, and who owned a remediation before they left. I built the assistant so those facts are not expected to survive inside a model’s context window. They live in Hindsight, where the agent can recall them, retain new ones, and reason across the accumulated history. That is the difference between a chatbot that can discuss compliance and an engineering system that can participate in an ongoing compliance process.


Top comments (0)