How Hindsight's retain/recall/reflect Split Shaped SprintMind's Architecture
Every team I've worked on has the same problem: the standup says NW-231 is in review, GitHub says the PR just merged, and the customer call from Tuesday said the deadline is Friday. Three sources, three tools, zero shared state. Somebody always gets surprised.
SprintMind is my answer to that problem. It sits on top of Jira, GitHub, and your meeting tool and connects what was promised, what engineering actually did, and what the team collectively remembers. The core insight is that the last piece—team memory—is the hardest to get right and the one most systems ignore entirely. For that, I built on top of Hindsight, a purpose-built agent memory layer, and the three primitives it exposes—retain(), recall(), and reflect()—ended up shaping nearly every architectural decision I made.
What the System Does
SprintMind ingests four kinds of input: meeting transcripts (recorded, uploaded, or pasted), standup and Jira digest documents, GitHub webhook events, and direct corrections from team members. From those it maintains two synchronized stores:
- A SQLite ledger that holds the current operational state: tickets, customer commitments, dependencies, open blockers, risk findings, and a full audit trail of every field change including the ones that were rejected and why.
- A Hindsight memory bank that holds the narrative history: what was said, by whom, from which source, and when.
The ledger answers "what is true right now." Hindsight answers "what did we say, learn, and commit to over the past two weeks." Both are updated in lockstep through a single write pipeline, and the query side combines evidence from both before generating any answer.
The risk engine is entirely deterministic—nine rules that read from SQLite and produce findings like "CI is failing on NW-231's PR and the Acme sandbox commitment is due in two business days." The LLM is involved only for explaining those findings in plain English, not for discovering them. That separation was intentional and I'll come back to it.
The Architecture Decision I Keep Coming Back To
When I started, I assumed I'd build a retrieval-augmented generation (RAG) pipeline: embed documents, store vectors, retrieve on query, pass to the LLM. That's what most people reach for. I abandoned it after sketching the first few query shapes I needed to handle.
An employee asking "what's blocking me today?" needs personal context scoped to their tickets, plus team-wide context, plus any corrections that have been entered since they last asked. A manager asking for a sprint health summary needs to synthesize two weeks of standups, GitHub activity, customer calls, and SOP guidance into a structured narrative. A simple top-k vector search doesn't compose across those cases cleanly, and building the retrieval logic for each one from scratch would have been a lot of plumbing.
Hindsight gives you three operations with different semantics: retain() stores memories with tags and metadata; recall() does semantic retrieval scoped by those tags; and reflect() synthesizes a structured narrative over the entire bank with a configurable schema. That map almost exactly onto the three things SprintMind needs to do—store facts, answer scoped questions, and produce manager-level summaries—without me having to build the retrieval infrastructure.
The bank itself is configured upfront with a mission, directives, and mental models:
# app/memory.py
BANK_MISSION = (
"I am SprintMind, the scrum and delivery memory of a software engineering team. "
"I track who owns which task, the status of every ticket across the sprint, blockers and "
"what they are waiting on, commitments made to customers in meetings (with owner and due "
"date), and the team's SOPs. I care about how task status evolves over time, and I always "
"prefer the most recent update when statuses conflict."
)
DIRECTIVES = [
("cite-sources", "When stating a task status, owner, or customer commitment, say where it "
"came from (e.g. which standup or meeting and its date)."),
("no-invented-status", "Never invent a task status, owner, or deadline. If memory does not "
"contain it, say it is unknown."),
("flag-blockers", "Always call out blocked work explicitly, including what it is waiting on "
"and who can unblock it."),
("latest-wins", "If two memories disagree about a task, prefer the most recent one and "
"mention that the status changed."),
]
Setting disposition_skepticism=4 and disposition_literalism=4 tells Hindsight to behave as a precise, literal tracker that doesn't speculate. Those aren't magic numbers—they represent the operational persona the bank should have when it synthesizes answers. For a system that tracks legal commitments to external customers, speculation is a bug.
Two Stores, One Write Path
The most painful part of building this was making sure the ledger and Hindsight stayed consistent without either becoming the source of truth for the wrong things. My solution was a single write pipeline in DeliveryService that every source—meetings, standups, GitHub webhooks, dashboard edits, corrections—must go through:
source event
→ ledger.add_event() [deduplicated by deterministic key]
→ ledger.apply() [source-aware field update, full history]
→ memory.retain() [best-effort; ledger is the system of record]
→ RiskEngine.evaluate()
→ risk transitions retained in Hindsight
The "best-effort" note on retain() matters: if Hindsight is temporarily unavailable, the ledger event is committed and the risk recomputes correctly. The memory write is swallowed. That's a deliberate tradeoff—delivery state must never stall because a memory API call failed.
The ledger's apply() method implements a source authority hierarchy that I'm genuinely proud of:
# app/ledger.py
SOURCE_AUTHORITY = {
"github": 5,
"approval": 4, "dashboard": 4, "feedback": 4,
"task_update": 3,
"standup": 2,
"meeting": 1, ...
}
CONFLICT_WINDOW = timedelta(days=3)
A meeting statement saying "NW-236 SSO is still in progress" arriving two hours after a GitHub PR merge does not overwrite the merge. The meeting has authority 1, GitHub has authority 5, and the merge happened within the conflict window. The disagreement is recorded as an open conflict and surfaces as a conflicting_evidence risk. This is one of the cases covered by the eval suite, and it caught a real bug during development—before I added the window check, late-arriving meeting summaries were silently reverting GitHub-confirmed statuses.
Every rejected field change goes into item_history with applied=0 and a reason field: "older than the current value (set by github on 2025-09-27)" or "sub-step reported done; ticket status unchanged." When a team member asks why a ticket still shows in review when someone mentioned it was done in a standup, the answer is already in the database.
The Query Side: Combining State With Memory
On the read side, the agent pulls evidence from both stores before generating any answer. For an employee question, that means three parallel recalls—personal scope, team-wide, and corrections-only—plus the ledger's current verified state:
# app/agent.py
async def _evidence(self, question, person_id, caller, *, budget, state_limit=10):
recalls = [
self.memory.recall(q, budget=budget, label="team-wide"),
self.memory.recall(question, tags=["kind:correction"],
tags_match="any_strict", budget="low", label="corrections"),
]
if person_id:
recalls.insert(0, self.memory.recall(
q, tags=[f"person:{person_id}"], tags_match="any_strict",
budget=budget, label=f"personal:{person_id}"))
results = await asyncio.gather(*recalls)
corrections = [f for f in results[-1] if _is_correction(f)][:4]
memories = [f for f in _dedupe(corrections, *results[:-1], limit=22)
if visible_to(f, caller)]
state = self.delivery.state_facts(question, person_id, limit=state_limit)
return state, memories
Current state entries (GitHub events, approvals, dashboard edits) are numbered alongside Hindsight memories and presented to the LLM with the instruction that CURRENT STATE entries are authoritative for what's true now and historical memories are what was said and when. That framing is doing a lot of work—without it, the model treats a standup from three days ago as equally credible as a PR merge from this morning.
For manager-level summaries, reflect() does the synthesis directly. The bank is configured with two mental models—"sprint-status-board" and "customer-commitments"—that Hindsight consolidates periodically. When the manager loads the delivery radar, reflect() is called with a JSON schema and those mental models as context. This is the part of Vectorize agent memory that maps most cleanly to the problem: I'm not managing a RAG loop, I'm configuring the memory system to maintain the right abstractions and then asking it to reason over them.
Citation Checking as a Lightweight Hallucination Guard
One pattern I'm keeping: post-processing every LLM answer to strip citation numbers that point beyond the evidence array.
# app/agent.py
_CITE_RE = re.compile(r"\[\s*(\d+(?:\s*[,–-]\s*\d+)*)\s*\]")
def check_citations(answer: str, n: int) -> tuple[str, dict]:
cited, invalid = set(), set()
def fix(m):
nums = _cite_numbers(m.group(1))
keep = [x for x in nums if 1 <= x <= n]
cited.update(keep)
invalid.update(x for x in nums if not 1 <= x <= n)
return "".join(f"[{x}]" for x in keep)
out = _CITE_RE.sub(fix, answer or "")
return out, {"cited": sorted(cited), "invalid_removed": sorted(invalid)}
This caught a real failure mode: gpt-oss writes citations using fullwidth brackets (【1】 instead of [1]), which bypassed the check entirely. The fix was a typographic normalization table applied before parsing. That kind of thing is hard to find without a working eval suite—I only discovered it because the eval cases compare cited evidence against expected content and an entire category of citation checks was silently passing as uncited.
Lessons Learned
Separate the oracle from the author. The risk engine is pure SQLite reads and Python logic. The LLM explains what the engine found; it doesn't find it. This makes risk output reproducible, debuggable, and testable without paying for API calls. Every rule has an eval case. That's only possible because "CI is failing" is a database read, not a language model judgment.
Tag your memories like you mean it. Hindsight's tag-scoped recall is only as useful as your tag discipline. I settled on a taxonomy—source:, person:, sprint:, ticket:, customer:, kind:—early and applied it consistently to every retained item. The corrections flow works (corrections are recalled first because they're tagged kind:correction and queried with any_strict matching) only because that tag was there from the start.
Source authority beats "newest wins." It's tempting to implement "the most recent update wins" as a single timestamp comparison. That rule breaks the moment a meeting summary arrives two hours after a PR merge. Modeling authority explicitly—GitHub events outrank standup lines, dashboard edits outrank task updates—and recording conflicts instead of silently overwriting produces a system that's transparent about disagreements rather than hiding them.
Memory writes should be best-effort; the ledger should not. The two stores have different durability requirements. SQLite with WAL mode gives you synchronous, transactional writes for operational state. Hindsight gives you semantic memory that improves over time. Coupling their availability would mean a Hindsight outage stalls ticket updates. Keep them independent, make the memory write fault-tolerant, and let the ledger be the ground truth.
Eval cases are documentation that runs. The 25 eval scenarios in evals/cases.py are the clearest description of what the system guarantees. Each one specifies an exact sequence of events (meeting says X, GitHub says Y, standup says Z two days later), the expected ledger state, and the expected answer behavior including what the system must refuse to say when information is missing. Writing those cases before the code was stable forced precision about edge cases that would otherwise have been hand-waved: what happens when a standup says "approval still pending" on a ticket that was already approved? (Nothing. That statement fails the APPROVED_RE regex, no dependency resolves, and the ticket stays unblocked.)
The thing I most want engineers to take away from this architecture is that Hindsight changed what was worth building. Without a memory layer that handles semantic retrieval, temporal reasoning, and structured synthesis, I would have spent the project building and debugging retrieval infrastructure. Instead I spent it on the domain logic: source authority rules, risk detection, citation checking, and the ingestion pipeline that makes standup lines and GitHub webhooks land in the same system with the same traceability guarantees. That's the work that makes the product useful, and it's the work that only happens when the memory layer is already solved.





Top comments (0)