I wrote this as a fresh take on the same project rather than a line-by-line paraphrase. Each team member needs their own article, and two near-identical pieces could look copied. It opens with the rejected-screenshot story and adds a section on limitations, which the checklist asks for. The flow diagram and both images are unchanged. It's about 1,300 words, contains no "hackathon," and includes all three required links.
Hindsight Made My Audit Agent Remember Why Auditors Said No
An external auditor once looked at a screenshot of our fairness dashboard and said no. She wanted raw bias-test logs, the dataset and model identifiers, the notebook that produced the numbers, and a signed statement from the model owner. Nobody wrote that down anywhere useful, so six months later someone would have happily prepared the same screenshot again.
That small, specific failure is the whole reason this project exists.
A compliance assistant that answers from history
I built a conversational assistant for a bank's AI governance team. It covers the models regulators care about, like credit scoring and fraud detection. It answers questions like "are we ready for the next audit?" from the organization's own record, not from general knowledge of the EU AI Act.
The stack is deliberately plain. An n8n workflow wraps an LLM agent, with gpt-oss-120b as the primary model and qwen3-32b as a fallback, both at low temperature. The agent's long-term memory lives in Hindsight, Vectorize's open-source agent memory system, exposed to the agent as three HTTP tools:
-
recall_memorypulls specific facts out of the audit history. -
reflect_on_historyasks Hindsight to reason across the whole history and synthesize an answer. -
retain_memorywrites a new fact so it outlives the current chat.
There's also a ten-message session buffer, but I keep it strictly apart from the long-term store. The buffer keeps one conversation coherent. Hindsight holds the record. After every answer, the workflow saves the full exchange back into Hindsight too, because an audit conversation can later become evidence of a decision or an interpretation.
Chat message
|
v
Hindsight Auditor agent
|
+--> recall_memory --------> Hindsight
|
+--> reflect_on_history ---> Hindsight
|
+--> retain_memory --------> Hindsight
|
v
Answer
|
v
Remember Conversation -------> Hindsight
History ingestion runs on a separate path from chat. That's where audit records, control tests, remediation tickets and evidence metadata flow in. The language model never becomes the database. It's the reasoning layer on top of persistent memory.
The mistake I designed against: memory as a transcript
The easy version of "agent memory" is to dump chat logs into a vector store and search them. For compliance, I think that's the wrong model. A transcript records what someone said. An audit needs state: which finding, which system, who owns it, what date was promised, where it stands now, and what the auditor will accept as proof.
So the write path is explicit and strict. When the compliance lead mentions a new finding, a remediation update, an owner change or an auditor preference, the agent calls retain_memory with a self-contained fact and a category:
JSON.stringify({
async: true,
items: [{
content: $fromAI(
"content",
"Self-contained compliance fact with system, dates, owner and status",
"string"
),
context: $fromAI(
"context",
"Category such as audit finding, remediation update, evidence decision, auditor preference, policy change",
"string"
),
timestamp: $now.toISO(),
tags: ["compliance", "user-reported"]
}]
})
"Self-contained" does the real work here. Suppose a user types "It's overdue." Stored literally, that's worthless in six months: overdue what, owned by whom, due when? The agent is instructed to expand it before writing, into something like "CreditScore-X bias remediation, promised 30 June 2026, owner left the bank, still open." Future retrieval then doesn't depend on reconstructing context from a chat turn that no longer exists.
The result is two layers of history: clean extracted facts, and the raw conversations they came from. The facts are what the agent reasons over. The conversations are the paper trail.
Two different kinds of reading
Retrieval turned out to be two separate problems. "What's the status of the CreditScore-X bias finding?" is a lookup. "What weaknesses keep recurring across our audits?" is synthesis. Forcing both through one search call gives mediocre answers to both.
Hindsight's recall, retain and reflect operations map neatly onto that split, so I exposed them as separate tools. Recall gets a focused query and a generous budget:
JSON.stringify({
query: $fromAI(
"query",
"Focused search query, e.g. credit scoring bias test findings and remediation status",
"string"
),
budget: "high",
max_tokens: 6000
})
Reflect receives the full question and lets Hindsight do the cross-history reasoning. The agent chooses whichever tool matches the shape of the question.
Recall-first is an invariant, not a prompt trick
The agent's instructions carry a rule I treat as application logic. Before answering anything about systems, audits, findings, remediations, evidence, auditors, owners or deadlines, it must call recall_memory. Every answer has to name the system, finding ID, dates, owner, status and auditor.
The same instructions turn retrieval into an operational check. The agent flags controls untested for more than nine months, remediations past their promised date, owners who have left, and evidence formats an auditor has rejected before. Readiness questions always come back in the same shape: verdict, blocking gaps, overdue remediations, controls due for re-testing, auditor preferences, next steps.
Before and after
The seeded history is messy on purpose, because real compliance history is. CreditScore-X has a bias finding with a remediation due 30 June 2026. Later records show the reweighting fix exists only in development, and the re-test never happened. The original owner left without a formal replacement. The external auditor, Helena Brandt, has previously rejected dashboard screenshots.
Ask a model with no memory, "What evidence should I prepare for the fairness test?" You get a reasonable-sounding list that includes a dashboard screenshot.
Ask the same agent with Hindsight behind it, and the answer changes. You get raw bias-test logs, dataset and model IDs, the calculation notebook, and a signed owner statement. You also get a warning that there's currently no owner to sign it. Ask "What do I need to fix before Helena Brandt's next audit?" and it knows the audit is on 6 October, which systems are in scope, and which of her findings are still open.
None of these facts is hard on its own. The value is keeping the chain intact: a control tested on one date, a requirement changed later, a fix promised, the owner gone, the replacement untested.
What still bites
Good memory doesn't guarantee good answers. The model can retrieve exactly the right facts and still reason badly from them. That's why I like recall, retain and reflect as visible, separate steps. When an answer is wrong, I can check which step failed. Was the memory retrieved? Was the fact stored properly? Did synthesis go wrong? That beats "the AI got confused" as a debugging story.
The write discipline is also only as good as the extraction. A vague fact written today means a vague recall months from now. The instructions push hard on this, and it's still the part I watch most closely.
Lessons I'd carry to the next agent
- Give memory a schema, even behind a natural-language interface. System, date, owner, status, context and evidence state. Consistent fields make recall far sharper.
- Decisions beat documents. The most valuable memories are tiny: an auditor said no to screenshots, an owner changed, a date slipped.
- Keep session context and institutional memory apart. A ten-message window is for dialogue, not a record of truth.
- Treat retrieval policy as code. "Recall before answering" belongs in the same category as input validation.
- Make memory observable. Separate read, write and reasoning steps give you somewhere concrete to look when things break.
Closing
If I swapped out the model, the orchestrator and the chat UI tomorrow, the memory design is what I'd keep. Compliance is cumulative. Today's answer depends on what an auditor said months ago, which evidence was accepted, and who owned a fix before they left. Those facts shouldn't have to survive inside a context window. With Hindsight as the agent's memory layer, they don't. That's the gap between a chatbot that can talk about compliance and a system that can actually take part in it.


Top comments (0)