This is not a shortage of information. Every answer a department needs has usually been worked out already, by someone, in a meeting that happened. The problem is where it went.
It went into chat threads, policy articles, document folders, meeting notes, email, calibration discussions, and individual memory. Seven places, none of which is searchable together, and the last of which walks out of the building at some point.
What that costs is specific and recognizable. The same questions get answered repeatedly, slightly differently each time. Prior decisions are hard to find, so they get re-litigated instead of applied. Conflicting guidance exists and stays buried until it causes a problem. Calibration outcomes are lost once the meeting ends. New team members depend on finding whoever happens to remember.
The last one is the tell. When onboarding runs on a person instead of a record, the organization lacks knowledge. It has employees who hold knowledge, which is a different and much more fragile thing.
Start with what it never does
Most descriptions of a knowledge agent start with capabilities. This one starts with the boundary, because the boundary is the design and everything else is downstream of it.
The agent never makes the determination itself. It never creates policy. It never resolves a conflict between two sources. It never overrides existing guidance. It never treats a discussion thread as official guidance.
What it does instead is retrieve, cite, compare and escalate. When two approved sources disagree, it does not pick a winner. It reports that they disagree and routes the question to the people whose job that is.
Humans decide. The agent remembers. That sentence is the whole governance model, and it is what makes the rest safe to build. An agent that answers authoritatively becomes a policy source nobody approved. An agent that only ever hands you the record, with its provenance attached, cannot quietly become one.
The agent lives where the work already happens
It sits in the channels the department already uses, so nothing has to be adopted. No new destination, no separate portal, no habit to build. Someone asks a policy or calibration question the way they already ask it, and the agent answers in the thread.
It searches approved policy articles, prior calibration decisions, escalation outcomes, channel history, meeting summaries, and a lightweight knowledge store.
It returns the source articles linked, prior decisions with dates, related discussions, known conflicts named as conflicts, and an escalation recommendation when the record leaves the question open.
Returning a known conflict is the part that is easy to undervalue. A system that always produces a confident answer will produce one when the underlying guidance contradicts itself, and the person asking will never learn that. Surfacing the contradiction is more useful than resolving it, and it is the honest output.
The build is small on purpose. It runs on what the department already has: the chat platform, its agent builder, the existing policy hubs, the calibration memory files, the escalation log, and a small local database holding policies, decisions, escalations, questions and conflicts. Nothing here is exotic, and that is the point. A knowledge system that requires a procurement cycle will not be tested this quarter.
The loop that turns an answer into an institution
Eight steps, and most of them are human.
- A question is asked in the channel.
- The agent searches articles and history.
- An answer is found, a conflict is found, or no guidance exists.
- An unresolved item is escalated.
- The calibration team reviews.
- Leadership approves where required.
- The approved decision is added to the knowledge base.
- Future askers receive the approved answer.
question ---asks---> THE AGENT ---reads only---> +--------------------+
^ | | APPROVED |
| | | KNOWLEDGE STORE |
+--returns sources----+ +--------------------+
and named conflicts | ^ |
| | |
escalates when unresolved | |
v | |
CALIBRATION TEAM | |
| THE ONLY WRITE |
v | |
LEADERSHIP ---approves----------+ |
|
future askers receive the approved answer <-------------------+
The agent reads and never writes. Every path that changes what the organization believes runs through a person, and exactly one arrow writes into the store. That is what stops the agent quietly becoming a policy source nobody approved, and it is why the loop compounds: step seven happens once, step eight is free forever.
Steps four through seven are human. Correctness enters at step seven and only at step seven, from the calibration team, from leadership, and from policy owners where applicable. Only approved decisions are written back. A discussion in a channel, however senior the person in it, is not a decision and does not enter the store.
This is institutional learning, and the distinction matters. The model learns nothing about the domain. The organization accumulates decisions it already made, in a place where the next person will find them. Step eight is where the compounding happens, and step eight costs nothing, forever, once step seven has happened once.
Where the memory lives
Three layers, as plain files in the document stores the teams already use, surfaced through chat. The structure is deliberately boring. The permission model is where the care is required, because whether people write honestly into a personal layer depends entirely on knowing who can read it.
A personal layer, one per person, readable only by its owner. Root context describing role, tools and working style. Live tasks, blockers and next actions. Decisions made, each with its reasoning. Recurring issues with their causes and resolutions. The standards that person works to, versioned. Plus one folder that is deliberately shared upward, holding flags the team needs to know and open questions being surfaced.
A team layer, team readable and lead writable. Team identity and active work. Team decisions with history. Patterns aggregated across the team. Active standards, version stamped. Anonymous signals arriving from the personal layers.
A department layer above both teams, readable by both leads. Cross team context. Patterns surfacing across both teams. Decisions affecting both. An anonymous aggregate from both team layers.
+-----------------------------------+
| DEPARTMENT LAYER | both leads
+-----------------------------------+
^ ^
a pattern seen in MORE THAN ONE team
| |
+---------------------+ +---------------------+
| TEAM LAYER eng | | TEAM LAYER product | team read
+---------------------+ +---------------------+ lead write
^ ^ ^ ^ ^ ^
a pattern, person stripped out
| | | | | |
+-----+ +-----+ +-----+ +-----+ +-----+ +-----+
| own | | own | | own | | own | | own | | own | owner only
+-----+ +-----+ +-----+ +-----+ +-----+ +-----+
named entries, written freely because only the owner can read them
Signal rises and loses identity at every boundary. A person writes named entries into a layer only they can read. What reaches their team is a pattern with the person removed. What reaches the department is a pattern that appeared in more than one team.
Signal moves upward and loses identity as it goes. An individual writes freely because the layer is theirs. What reaches the team layer is a pattern with the person stripped out. What reaches the department layer is a pattern that appeared in more than one team. Each boundary is a permission boundary, so the privacy property holds structurally and nobody has to keep remembering a rule.
Two phases, and why the order carries the weight
Phase 1 runs to completion before Phase 2 begins. That sequencing is the most consequential decision in the test, and it is easy to mistake for scheduling.
Phase 1, weeks one and two, is engineering only. Personal layers for each engineer. The team layer capturing architectural decisions and patterns. Session context assembled automatically at the start of each session. Decisions logged with their reasoning attached, the why alongside the what. Bug patterns and root causes captured in structured form. Standards made explicit, versioned and retrievable.
Phase 2, weeks three and four, adds product while engineering continues unchanged. Product onboards using the same architecture. The shared department layer activates. Engineering decisions become visible to product context and the reverse. Cross team patterns surface that neither team sees alone.
When product onboards in week three, engineering does not learn alongside them. Engineering has been living in the system for two weeks and helps with the setup, so the second team takes days, against the two weeks the first team needed. Run the phases in parallel and that advantage disappears: two teams learning at once, and nobody in the room who has done it before.
Engineering and product are chosen because they hold the most consequential knowledge gap in most technology organizations. Engineers know how a thing was built and which alternatives were rejected. Product knows what was decided, why it was prioritized, and what the customer signal was. Those two records are almost entirely separate today, and the value of connecting them needs no domain expertise to evaluate.
What would count as working
Stated in advance and deliberately narrow. Every criterion is a thing that either exists or fails to. Nothing here is a satisfaction score, for reasons given in the next section.
Prior decisions are findable in under thirty seconds, timed, by someone other than whoever filed them. A repeated question gets the same answer twice, asked again in week four from a different account. Conflicts are escalated rather than buried, with at least one case running all the way to an approved decision. Responses carry their sources, and a spot check confirms the link supports the claim. An approved decision from week one is still retrievable in week four with its date and reasoning. A cross team pattern surfaces that neither team saw alone. People are still asking questions in week four without being reminded to.
That sixth one is the core claim of Phase 2 and the only signal that cannot be produced by either team in isolation. If nothing appears in two weeks, the department layer has not earned its place.
Three predictions, recorded before the test starts
A protocol published after the fact is a story. The author already knows how it turned out and the criteria drift, quietly, toward whatever happened. Publishing first is the only cheap way to stop that.
I am strict about this because I recently paid for it. An experiment of mine returned a clean, total result: one hundred percent in the treated arm, zero in both controls, across all four cases. It was worthless. Every task in it could be answered by copying a sentence out of the text being injected, so the harness was measuring reading comprehension. The only reason I caught it was a prediction written to disk beforehand saying the result should be messy and mixed. The clean number contradicted the prediction, and that contradiction was the entire signal. Two later versions of the same instrument were also wrong, each in a way that flattered whatever I was hoping for.
So here are three predictions, on the record, before any data.
Everything below rests on three measurements taken on 2026-08-11, on a system I run, against codebases I did not write. They are small, and they are stated with their size so they can be argued with.
| What was measured | Setup | Result |
|---|---|---|
| Where durable lessons come from | 4 conditions, 2 foreign codebases | Ordinary use 1, isolated test 0, agreeing with 20 recorded fixtures 3 |
| Whether a healthy system generates material | 1 full suite, 476 tests passing | 0 durable lessons |
| Whether irrelevant context is harmless | 12 trials per arm, 1 model, 1 scenario | Irrelevant note wrong 12 of 12, control merely vague |
One. The layers will fill unevenly, and the pattern will not be effort.
People working where their output has to agree with someone else's, shared interfaces, contested policy, anything with a boundary in it, will produce full layers. People working on self contained work will produce nearly empty ones while working just as hard. Read that as engagement and half the team gets judged for the shape of their work instead of its quality.
The basis, measured 2026-08-11 across four conditions on two codebases I did not write: ordinary use produced one durable lesson, adding an isolated test produced none, and making one component agree with twenty recorded fixtures produced three in about thirty minutes. The variable was integration surface, not hours worked.
Two. Week one will feel like nothing is happening.
A new layer is empty, and a process running normally generates very little worth recording. The material appears when something is contested, revisited, or wrong. A team expecting usefulness by day three will conclude the test failed before the compounding has begun.
The basis, measured the same day: a full test suite on a healthy codebase, 476 tests passing in about a minute, produced zero durable lessons worth recording. Nothing was broken. There was simply nothing to learn, because nothing resisted.
Three. Asking whether it felt useful will return yes, and will mean nothing.
Delivered context makes people feel better informed whether or not the context was any good. This is why every criterion above either exists or fails to, and why none of them is a rating.
The basis, measured the same day at twelve trials per arm, one model, one scenario: against a control given nothing, injecting a genuinely irrelevant note produced a confidently wrong answer in twelve of twelve cases, while the control was merely vague. The bad context went past unhelpful into actively misleading, and it read as authoritative while doing it. Treat that as a lead, with the sample size and the single scenario counting against it.
That third one is also the strongest argument for the boundary at the top of this piece. An agent that hands over sourced records lets a person check. An agent that hands over confident answers removes that option.
What I do not know
There is one instance of this architecture running in production over a long period, and it is mine. One organization, one operator, and every figure quoted above comes from it. This document generalizes to any two teams in any organization, and that generalization is a hypothesis rather than a finding. This test is the first attempt to check it against somebody else's work.
Three further things it cannot settle, listed because a careful reader will find them anyway.
Whether people write honestly into a layer their employer hosts. The permission model is designed for this and design is not proof. A personal layer that quietly becomes a performance record stops receiving true entries immediately, and that failure is silent.
Whether a cross team pattern is a pattern or a coincidence. Two teams generate enough signal that something will always look like a connection. Requiring that neither team saw it independently is a guard, and a weak one.
Whether four weeks is long enough for anything to compound. It is long enough for the layers to fill and for first patterns to appear. It is almost certainly not long enough to observe the property the whole design aims at, which is a team keeping its knowledge across a departure.
The test does not end at four weeks. What ends at four weeks is the part anyone is willing to make predictions about.
If you have run something like this inside a real organization, I am most interested in the first unknown above. The permission boundary is the part I am least able to prove from my own instance, because in my instance the owner and the operator are the same person.
Top comments (0)