DEV Community

ellytan
ellytan

Posted on AI-assisted

Make Each Mistake Only Once: Team Memory in Practice with TencentDB Agent Memory

GitHub: TencentDB Agent Memory

As execution becomes abundant, how should organizations rebuild the flow of information and the inheritance of experience?

Coding agents are rapidly expanding what one person can accomplish. But the bottlenecks in enterprise development do not disappear automatically. Once a task crosses people, sessions, code branches, and permission boundaries, the scarce resource is often no longer code generation. It is the bandwidth of collaboration.

This article explains how TencentDB Agent Memory turns task trajectories, project knowledge, code structure, and working methods into governed memory assets. We discuss what we learned from 2,600 sessions, real Redis development cases, and developer interviews, and how those findings changed our product design. In a related-case evaluation using SWE-bench tasks, learning from earlier cases increased completion from 60% to 80%. In a separate set of 50 exceptionally long, difficult tasks, memory reduced cost by about 19% while also improving success. These are scoped experimental results, with limitations discussed below.

1. Execution is growing faster than collaboration bandwidth

While building Agent Memory, we kept seeing the same contrast: models were becoming stronger and individuals could move faster with an agent, yet delivery slowed when work entered a setting with multiple people, agents, and projects. The issue was less often whether code could be generated and more often whether information could be inherited accurately across people, sessions, branches, and permission domains.

A person and their agent usually share the same immediate task context: the goal, what has already been tried, and why an approach was abandoned. When another person or agent takes over, the handoff may contain only a conclusion, a file, or a diff. The recipient sees the result without the background, constraints, and reasoning that produced it. The information exists, but the task still has to be understood from scratch.

Organizations usually do not need more messages. They need usable context that the right person or agent can understand and apply within the correct permission boundary. AI lowers execution costs, but it does not automatically increase that bandwidth. Faster execution can make the costs of lost context, repeated exploration, and branch conflicts more visible.

1.1 What the one-person company reveals

Here, a one-person company, or OPC, means an AI-native organizational model: one principal decision-maker works with multiple agents, delegating research, development, content, and operations while retaining responsibility for goals, judgment, and acceptance.

Agents are expanding the range of tasks an individual can execute. METR's measurements of frontier models suggest that, for self-contained tasks primarily in software engineering, machine learning, and cybersecurity, the task duration associated with a 50% success probability has historically doubled roughly every seven months. Anthropic's privacy-preserving analysis of approximately 400,000 Claude Code sessions also reported that coding-agent activity in GitHub projects had more than doubled since late 2025, with the users in that analysis averaging about 20 hours per week. One vendor's data cannot represent an entire industry, but it suggests that agents are becoming a persistent work interface for some users.

OPC efficiency also comes from fewer handoffs, fewer organizational barriers, and concentrated goals. In Carta's startup sample, about 36% of companies founded in 2025 had a solo founder, up from 31% in 2024. This does not establish that mature one-person companies are already widespread. It does point to a useful research question: as individuals gain access to more execution capacity, how much work can be done with fewer interpersonal interfaces?

The OPC does not eliminate collaboration. It turns much of it into collaboration between one person and several agents sharing context. It illustrates one possible upper bound for AI-native execution, but enterprises cannot simply copy the model.

1.2 Enterprise scale adds boundaries

An enterprise differs from an OPC in more than headcount. A real team must manage roles, project isolation, branches and versions, data security, and accountability. A useful approximation is:

Collaboration cost ≈ number of handoffs × cost per handoff + conflict-resolution cost + permission and compliance cost.

This is not a financial model. It separates three questions that are often conflated: Can the information be found? Can it be understood? Is its use authorized? Meetings, messages, and documents are not bandwidth in themselves. Information becomes effective context only when it is complete, trustworthy, applicable to the current version, and usable in the task.

The most dangerous conflicts may be small differences between branches. Two deployment skills can look almost identical, yet one reads config/prod/ while another branch uses conf/prod/. Two test procedures may adjust the same parameter, but one multiplies it by 0.5 and the other by 2. A system that selects the “best” asset solely by semantic similarity can execute a plausible workflow in the wrong directory or environment and even cause an outage.

Our operational definition is therefore stricter:

Collaboration bandwidth is the amount of effective context that the relevant person or agent can correctly understand and directly use per unit of time, within the appropriate permission boundary.

1.3 From an “AI GitHub” to a three-layer organization

A talk by Analemma founder Sun Tianxiang, published by Sequoia China, proposed an organizational metaphor: people work with high cohesion in their own branches and merge at suitable moments. Another person's agent can understand the work trajectory, so the team does not need to repeat trials that have already been validated.

The proposed AI-native organization has three layers: management and resource allocation at the top, people and AI working together in the middle, and a shared foundation of skills, trajectories, and supporting infrastructure at the bottom.

We find the direction useful, but enterprise implementation must include the mechanisms that make GitHub collaboration work: versions, permissions, review, merging, conflict resolution, and rollback. Raw trajectories may contain credentials, customer data, failed attempts, and unconfirmed judgments. Visibility does not make a conclusion trustworthy or appropriate for everyone to see.

Our interpretation is: the upper layer governs information boundaries; the middle layer carries real tasks; the lower layer preserves organizational assets that can be inherited. Together, they form a path from tasks to assets and back to tasks.

1.4 Why specifications matter

Spec-driven development offers a related signal. As agents implement faster, more effort shifts from writing code to expressing intent, constraints, plans, and acceptance criteria accurately—and preserving that task state.

GitHub Spec Kit organizes work as:

Specify → Plan → Tasks → Implement

Requirements and constraints → technical design → executable tasks → implementation and validation

A specification is treated as an evolving shared source of truth, rather than a static document completed before development. Agents amplify the cost of ambiguity: one missing boundary can produce an entire incorrect implementation, while unclear acceptance criteria move rework into design review and code review.

Specifications make intentions machine-readable, move constraints out of chat, preserve task state across sessions and agents, and give reviewers an explicit contract instead of forcing them to infer everything from a diff.

But a specification primarily serves the current task. It does not automatically explain whether a similar problem occurred before, why a historical approach was abandoned, which module experienced a related failure, whether a rule conflicts with another branch, or whether the current task may access that information. Nor does it inherently manage provenance, freshness, versions, and negative feedback.

Specifications make one task's state expressible. Team memory makes validated task state inheritable. The former demonstrates the need for structured context; the latter adds reuse across tasks and enterprise governance.

2. How TencentDB Agent Memory is structured

2.1 The goal is fewer wrong turns in the next task

We define team memory as turning the background, knowledge, code relationships, and methods produced by real tasks into agent assets that can be retrieved, combined, traced, authorized, and continuously updated—and then assembled as needed for future tasks.

This is more than retaining every chat or adding a vector index to a knowledge base. Traditional RAG emphasizes what can be found in a corpus. Team memory must also answer who owns information, which version it applies to, who may use it, which agent should receive it, and how it should be corrected after use. It manages the lifecycle from creation and abstraction to review, use, and retirement.

Memory Hub handles governance and scheduling. A Memory Pack enters the task context shared by a person and an agent. Four types of Memory Asset form the shared foundation. Rather than putting everything into one giant knowledge pool and expecting the agent to find the right information, we keep complete assets in the pool and assemble only what the current identity, task, and permissions require.

The asset pool holds the complete record. The current task receives what is relevant.

2.2 Four asset types preserve four kinds of state

  • Chat Memory preserves background, constraints, decisions, preferences, and historical interactions. It answers why a judgment was made and which approaches have already been tried, restoring task state without repeated explanation.
  • Wiki preserves product knowledge, architecture, standards, runbooks, and established conclusions. It provides stable facts and design constraints.
  • CodeGraph preserves symbols, files, calls, dependencies, and impact paths. It helps determine which modules a change may affect.
  • Skill preserves repeatable methods, tool calls, boundaries, and validation rules. It makes a verified workflow reusable.

These are not four unrelated document directories. A debugging task may recover historical context from Chat Memory, read architectural constraints from Wiki, identify the impact surface through CodeGraph, and use a Skill to investigate and validate the fix.

They also enter the task at different levels of detail. Chat Memory has layers from raw conversation to stable understanding; Wiki, CodeGraph, and Skill use summaries, bindings, and tool entry points for progressive access.

2.3 Narrow the boundary before searching for relevance

A flat memory pool forces raw records, abstract conclusions, fixed rules, and temporary context to compete for attention. The most similar result may not be the most reliable, nor applicable to the current identity, branch, or task. We therefore separate how memory is formed from how assets are assembled.

Content layers: from evidence to stable understanding

Chat Memory uses an L0 → L1 → L2 → L3 pipeline. Each layer distills the previous one without replacing it:

  • L0 Conversation: original conversations, timestamps, and full context. High fidelity, low density; useful for checking exact statements and recovering evidence.
  • L1 Atom: facts, constraints, decisions, events, and instructions. Atomic and structured; useful for precise retrieval of actionable information.
  • L2 Scenario: memory blocks organized around a project or scenario. Denser context for quickly resuming a category of work.
  • L3 Core / Persona: stable patterns, long-term profiles, and higher-level understanding. Highly abstracted and updated less frequently; useful for entering the user and team context.

L2 and L3 can establish context quickly. When a task needs specific facts, keyword search, vector retrieval, and source anchors lead back to L1 and L0. The upper layers reduce reading; the lower layers protect against distortion through repeated summarization.

Routing layers: from identity to a context bundle

  1. Identity and scope: establish Team, User, Agent, Task, project, visibility, and ACL. Unauthorized assets never enter the candidate pool; they are not retrieved first and removed later.
  2. Fixed bindings: explicitly attached role rules, task constraints, Wiki resources, and required Skills enter the assembly scope directly. An incidental similarity ranking must not displace organizational requirements.
  3. Dynamic retrieval: within the authorized scope, retrieve additional Memory, Skill, Wiki, and CodeGraph candidates relevant to the task and request intent.
  4. Relevance fusion: use exact retrieval such as BM25 for error codes, filenames, and symbols, and vector retrieval for intent, failure patterns, and similar tasks. Combine them with reciprocal rank fusion, or RRF.
  5. Context assembly: build a Memory Pack according to role, priority, binding, version, and token budget. Asset versions can be pinned for a task to avoid inconsistent rules during execution.

Progressive disclosure: expose the entry point first

Assembly does not mean copying every selected asset into the prompt. Chat Memory can expose a scenario summary before retrieving an original turn. A Skill can expose its name, trigger conditions, and purpose before loading its full procedure and resources. Wiki and CodeGraph initially expose resource names, summaries, and tool entry points. An agent discovers capabilities through /v3/tools/list and retrieves pages, source code, callers, or impact paths through /v3/tools/call.

The prompt explains what is available and why it might help. Tool calls provide details when needed. Result counts, individual lengths, total characters, and timeouts also constrain the process. The goal is a reliable starting point with as little context as necessary.

2.4 Assets belong to the team and are assembled per task

Personal memory preserves continuity between a person and their agent. Team memory preserves organizational experience. Assets therefore need to belong to the Team rather than to a particular agent. Experience trapped in one account, conversation format, or framework cache is closer to a private cache than a team asset.

Framework independence is about ownership and continuity, not just compatibility. Chat Memory, Wiki, CodeGraph, and Skill use a common asset model carrying source, owner, scope, version, ACL, evidence, and state. Different agents read and write through gateways, HTTP APIs, SDKs, or tool interfaces. The agent decides how to plan and execute; Memory lets it start from what the team already knows.

This does not mean giving every agent the same content. Bug fixes may need historical failures, CodeGraph, and troubleshooting Skills. Requirements analysis may need Wiki, Chat Memory, and business constraints. The team owns the complete pool; each role and task receives the relevant subset.

The infrastructure connects all three organizational layers without making business decisions for managers or imposing one agent workflow.

2.5 Turn task outcomes into organizational assets

Raw chat contains evidence, temporary guesses, repetition, and invalidated approaches. Treating an entire trajectory as long-term memory can increase noise and conflict as retrieval grows.

We divide asset formation into four steps:

  1. Segment evidence: split sessions, documents, and repositories into identifiable tasks, turns, pages, symbols, and commits, rather than preserving only a summary.
  2. Extract candidates: identify context, decisions, rules, code relationships, procedures, failed paths, and validation results as candidate Atoms or Skills.
  3. Bind scope: attach Owner, Team, Repo, Branch, Path, Version, Time, ACL, and source evidence. These fields matter as much as the text.
  4. Promote after validation: keep unverified trajectories as low-weight background. Promote content toward stable scenario memory or shared Skills when tests, commits, or human review support it.

At task completion, new context, judgments, relationships, validation methods, and negative feedback become candidates that return to Memory Hub after review. The aim is to retain what was learned, avoid restarting after a handoff, and begin the next task from a saved state.

An asset is a combination of evidence, abstraction, scope, and validation status. Its test is whether it reduces uncertainty in the next task.

2.6 Cold start without waiting years

Teams should not have to wait years for memory to accumulate. Historical sessions can seed Chat Memory and Skills, existing documents can seed Wiki, and repositories can seed CodeGraph. New tasks continuously add evidence.

The infrastructure first captures information the team already has, then lets new tasks validate, correct, and extend it. The purpose is to help the team start from zero less often.

2.7 A worked example: resuming a coding task with memory

Consider a real code path in the TencentDB Agent Memory repository. Knowledge Service exposes Wiki and CodeGraph tools through /v3/tools/list and executes read-only queries through /v3/tools/call. We can use the addition of a CodeGraph impact query to illustrate how an internal coding task could resume with memory.

The modules, interfaces, and constraints below come from the repository. The task narrative is an engineering reconstruction to explain the workflow, not a verbatim account of one person's work or a particular commit. It makes no additional claim about measured efficiency.

The external short name impact must enter the read-only allowlist and map through toCodeGraphToolName() to codegraph_impact. MemoryKnowledge/src/routes/code-graph.ts must register the direct query route from the same tool list and validate fields such as symbol and depth. Execution then passes through MemoryKnowledge/src/engines/code/bridge.ts to the CodeGraph engine.

If the mapping is omitted, the agent may discover the tool but receive a 403 unknown tool error when calling it. If a separate interface bypasses the shared route, it may omit x-tdai-service-id isolation, parameter allowlisting, or ready-state handling. A locally correct change can still leave the system incomplete.

A fresh session must reconstruct these relationships. Finding CODE_GRAPH_TOOLS by filename does not immediately reveal all the team's existing constraints:

  • Agents may call query tools; management operations such as create, delete, and sync must not enter the tool allowlist.
  • External short names such as impact map consistently to internal names such as codegraph_impact.
  • service_id comes from the x-tdai-service-id request header. Resource-ID queries still require tenant constraints, and cross-tenant access returns 404.
  • Queries return safe empty results until CodeGraph reaches ready status, rather than using an incomplete index.
  • Tool definitions, direct routes, unified tool routes, and the MCP surface must remain consistent.

These facts are distributed across interface designs, READMEs, historical sessions about multitenancy, code comments, type constraints, and earlier tool extensions such as search, explore, and callers. A task-specific Memory Pack brings them together:

  • Chat Memory: the decision to separate external short names from the internal codegraph_ prefix and keep one source of truth for registration. This helps prevent updating the description but forgetting the execution mapping.
  • Wiki: the progressive-disclosure protocol and the boundary between agent-accessible queries and management operations. This helps prevent accidental exposure of sync or delete.
  • CodeGraph: the createToolsRoutes → executeCodeGraphTool → toCodeGraphToolName → executeTool path, plus the direct route's dependency on shared constants. This helps prevent missing a route or bridge outside the current file.
  • Skill: a checklist covering schema, allowlist, mapping, tenant isolation, state branches, type checks, and regression cases. This helps prevent happy-path-only validation.

Before editing code, the agent should then produce a repository-level plan:

  1. Add impact(symbol, depth) to the definitions in routes/tools.ts and confirm its read-only allowlist entry.
  2. Reuse CODEGRAPH_QUERY_TOOL_NAMES so the unified entry point and routes/code-graph.ts register the same set.
  3. Check that toCodeGraphToolName("impact") produces codegraph_impact and reaches the engine through bridge.ts.
  4. Preserve x-tdai-service-id and resource-ownership checks; code_graph_id alone must not permit cross-tenant reads.
  5. Test required symbol, depth bounds, unknown-tool 403, cross-tenant 404, safe non-ready behavior, and normal results.
  6. If Panel, SDK, or MCP exposes the capability directly, update types and documentation according to the same contract.

The intended change in the development experience is concrete. At task start, the agent recovers the shared protocol and historical decisions instead of searching thousands of lines for an entry point. During impact analysis, it sees registration, routing, mapping, isolation, and testing instead of “add one item to an array.” Implementation follows shared constants and constraints. Review covers invalid input, index state, and tenant boundaries as well as the happy path. Finally, the implementation path and validation checklist become a reusable Skill for future extensions.

Memory does not replace debugging when a new engine bug has no relevant history. If a historical rule is obsolete, versions, provenance, and test results must support downranking or withdrawing it.

3. What we learned from 2,600 sessions

3.1 Separate related information from reusable experience

Our Task-only analysis asked a specific question:

Before a bottleneck occurred, did earlier tasks already contain actionable experience that could have prevented or reduced it?

We split 2,600 original sessions into 5,081 tasks. A long session often contains several goals, so session-level similarity can confuse “same repository” or “same user” with a useful problem relationship.

Candidate relationships combined semantic vectors; overlapping files, directories, and symbols; errors and traps; action sequences and artifacts; time; and strong anchors. The first pass produced 48,114 candidate pairs. Filtering retained 23,134 for a stricter second review, independent reclassification, and checks against original-turn evidence.

A relationship broadens retrieval; it does not prove that experience is useful. A high-confidence conclusion must point back to the original text and use only information available before the bottleneck.

3.2 What the 38% result actually means

In an early stratified sample, only 38% of same_problem labels were correct as that exact relationship type, although 95% of those pairs had some concrete relationship. For reusable_sop, exact-type correctness was 62%, while 96% had a concrete relationship.

This does not mean that 38% of knowledge is reusable. It means models can find related tasks more easily than they can distinguish the same work item, the same problem, and a transferable procedure. A relationship is not the same as reuse, and reuse does not imply that something can be executed directly.

After stricter review, independent classification, and machine-verifiable anchors, we retained 22,361 canonical relations. Only 42 same_work_item, 135 same_problem, and 54 reusable_sop relations remained strong. The other 22,130 became related_context “Shadow” relations used only for retrieval.

Sampled precision for the three strong types was 95.24%, 69.00%, and 83.33%, respectively. The same_problem sample fell one example short of the predefined 70% threshold. We retained that as a known risk instead of adjusting the definition until it passed.

The resulting funnel was:

  • 2,600 original sessions: a session is not necessarily one task.
  • 5,081 tasks: organize memory around goals and evidence.
  • 48,114 initial candidate relations: broad retrieval finds clues but includes substantial noise.
  • 22,361 canonical relations: verify time and source evidence.
  • 231 strong relations: high-confidence assets should be selective and reliable.
  • 22,130 Shadow relations: background can aid discovery without directly driving execution.

This changed our asset strategy. “Keep” and “delete” are insufficient states. The system needs strong assets, weak hints, and background references, with promotion, downranking, and withdrawal as evidence changes.

3.3 Which losses should memory address first?

We identified 2,203 bottlenecks. At least one appeared in 1,644 of 5,081 tasks, or 32.36%, and in 1,264 of 2,600 sessions, or 48.62%.

  • Logical rework: 1,350. Reversing or redoing a design, implementation, or interpretation.
  • Missing context: 269. Needing repository, business, or historical background.
  • Intent mismatch: 230. Moving away from the goal or acceptance criteria.
  • Repeated failure: 145. Retrying an already unsuccessful path.
  • Lost state: 74. Failing to resume after a session, person, or agent change.
  • Quality iteration: 65. Requiring several revisions to reach the expected quality.
  • Task decomposition: 50. Failing to turn a goal into executable steps.
  • Missing external knowledge: 20. Needing facts or rules outside the current context.

An independent-prompt audit of positive samples achieved 84% precision, 84% accuracy in locating the target turn, and 82% accuracy in locating the causal turn. The zero-window miss rate was approximately 6.5%. Both the production analysis and the audit were model-based, not human gold-standard annotation. These are quality indicators for an internal automated analysis, not general industry findings.

The useful signal is that 1,350 logical-rework cases far outnumbered the 269 missing-context cases. Team memory cannot stop at document retrieval. It must preserve reasoning, failed approaches, applicability conditions, and validation methods; otherwise an agent can find relevant material and still repeat the same reasoning error.

4. Developer interviews and real Redis cases

4.1 Design and review can cost more than code generation

We interviewed seven developers using AI coding tools across complex system design, maintenance, automated testing, large features, and incident handling. The account here omits identifying details about people, projects, and incidents and focuses on shared patterns.

First, design and code review often take more time than writing code. Review requires recovering requirements, understanding tradeoffs, checking impact, and identifying violations of historical constraints. A final diff explains what changed but not all the reasons behind it. This is why a task's Memory Pack combines specifications, Chat Memory, and CodeGraph.

Second, fewer people may collaborate continuously on the same small task. One person often owns a component or coordinates multiple agents, with synchronization concentrated at design review, interface agreements, and merging. Fewer live handoffs shift context-recovery work into review and maintenance. Team memory should allow cohesive work to merge safely when needed, rather than demand constant synchronization.

Third, restoring state across sessions remains a consistent pain point. In established codebases, the missing information is often “why”: why a compatibility branch exists, which incident produced a condition, or whether an apparently redundant check is still necessary. Code graphs and current documentation do not necessarily contain these answers.

Fourth, freshness, provenance, version, and visibility into retrieval determine trust. Developers want to know why a memory was used, which task it came from, whether it applies to the current branch, and how to correct it. Silent injection of plausible text makes it difficult to distinguish model judgment from established team knowledge.

Fifth, memory and deterministic tools have different roles. Memory supplies risks, context, and historical methods before a task. Scripts, tests, and checklists validate results afterward. Neither replaces the other.

The interviews also changed our product sequence: demonstrate immediate personal and project value before relying on future team benefits. Contributions are harder to sustain if they help only an unknown colleague encountering a similar problem later. Restoring the contributor's own state, reducing review explanation, or improving a specification creates an immediate reason to contribute.

4.2 Use real Redis tasks to refine asset selection

Public benchmarks help control variables, but do not capture every project boundary, issue-tracker convention, path difference, and time constraint. Alongside the 5,081 historical tasks and 22,361 relations, we selected 35 real Redis TAPD test tasks to build a historical-task retrieval and asset-selection environment.

This was not another completion-rate benchmark. It asked whether we could find earlier tasks that were valid in time, relevant to the project, and backed by traceable evidence, then distinguish executable experience from weaker hints and background.

The process included broad judgment of candidates, strict second review, verbatim source verification, audits of stratified negatives and out-of-scope candidates, recovery of suspected misses, and deterministic anchoring with unique TAPD IDs.

The results were:

  • 35 target test tasks.
  • 12,574 primary candidate pairs.
  • 57 final related pairs.
  • 18 target tasks with at least one final relation.
  • 54 historical tasks across 45 historical sessions.
  • 32 strong relations, 10 weak relations, and 15 background references.
  • 0 non-Redis candidates and 0 time-invalid candidates in the final set.

The environment helped refine engineering boundaries rather than simply maximize retrieval volume:

  • Scope must be precise. The same repository does not imply the same branch, path, configuration, or environment. Issue identity, version, and time belong in scope.
  • Strong and background relations have different uses. Strong assets can inform plans and Skills; weak relations should carry caution; Shadow relations aid discovery without automatically driving execution.
  • Future information must not leak backward. A later fix cannot become “historical experience” for an earlier target task.
  • Every conclusion must return to its source. Cluster labels and model summaries are indexes; original turns, issues, code, and validation results remain the evidence.
  • Negative examples matter. Similar-looking but inapplicable cases help define when an asset should not trigger.

Enterprise memory is difficult because applicability must be established carefully. Semantic similarity is an entry point. Project boundaries, chronology, evidence, and validation determine whether an experience can be acted on.

5. Evaluation: learn from earlier cases, test on later ones

The number of memories stored or Skills generated does not establish value. The relevant question is whether an agent can learn from completed tasks and perform better on subsequent related work.

5.1 Experience inheritance on SWE-bench tasks

SWE-bench Original contains 2,294 real GitHub issue/PR tasks from 12 popular Python repositories. SWE-bench Verified contains 500 tasks that software engineers confirmed to be solvable.

Our evaluation modeled the gradual accumulation of team experience rather than supplying a collection of manually written answers. Three constraints mattered:

  1. Transfer experience only between tasks in the same repository or with a genuine relationship.
  2. Respect chronological order: a later case cannot receive its own answer or information from future cases.
  3. Use task tests to determine success rather than asking the model whether memory helped.

In this related-case evaluation, team memory increased task completion from 60% to 80%: an absolute gain of 20 percentage points, or about 33.3% relative improvement. This is the combined result of the four asset types, not an isolated score for each asset type.

5.2 Exceptionally long and difficult tasks

We selected 50 of the most difficult SWE-bench cases and combined them into a continuous sequence of long problems within one session. The purpose was to examine experience inheritance, information retrieval, and goal consistency during exceptionally long work.

The reported success rate without team memory was 17%, at a cost of $887.64. With team memory, the reported success rate rose to 20%, while cost fell to $717.78. Total tool-call turns also declined by nearly 19%.

5.3 What these results do—and do not—establish

The results support a specific finding: assets formed from earlier tasks can help agents complete later, related software-engineering tasks. This is consistent with the direction of research such as Agent Workflow Memory: historical trajectories can contain reusable workflows, not merely logs.

They do not mean that deploying memory improves every development task by 20 percentage points. Results depend on case selection, the baseline agent, model version, asset extraction, retrieval settings, repetition count, and statistical interval. Before broadening the claim, we need a fixed evaluation subset, reported sample sizes and confidence intervals, separate measurements of helpful, neutral, and harmful transfer, and more testing across repositories, versions, and longer time spans.

SWE-bench and the Redis environment play different roles. Executable tests measure final task completion. Real project cases refine applicability, evidence chains, and confidence levels. One asks whether memory helps; the other asks how to avoid using it incorrectly in a real team.

6. Team memory is also a governance system

6.1 Six design principles

1. Preserve state that others can inherit. Personal memory supports continuity; team memory supports inheritance. After a change of person, session, or agent, goals, constraints, decisions, and unfinished state should remain available.

2. Judge assets by the uncertainty they remove from the next task. An asset earns promotion by reducing wrong turns, restoring necessary background, clarifying code impact, or enabling deterministic validation. Memory count is not a success metric.

3. Preserve both abstraction and evidence. Upper layers offer usable conclusions, rules, and procedures. Lower layers retain tasks, original turns, issues, commits, tests, and versions. Abstraction without evidence is difficult to trust; evidence without abstraction requires rereading all the history.

4. Make assets composable and assemble them per task. A task receives Chat Memory, Wiki, CodeGraph, and Skills relevant to its goal, role, project, and permissions. Atomicity means a clear trigger, one responsibility, and independent validation—not endlessly splitting knowledge into smaller fragments.

5. Give memory a lifecycle. Assets need creation, review, publication, forks, merges, downranking, pinning, expiration, and deletion. Human negative feedback must affect future routing instead of relying on the model to make a better judgment next time.

6. Decouple memory from models and agent frameworks. Models and workflows change quickly. Organizational experience should survive account, session, and framework migrations.

6.2 Six engineering challenges in shared memory

Conflict: similar assets may prescribe different actions for different branches, paths, or environments. Scope filtering, version priority, conflict detection, and, when needed, human review must resolve the difference.

Freshness: an interface upgrade, configuration migration, or business-rule change can invalidate a once-correct conclusion. Assets need valid_from, valid_to, last-validation time, and code-version information.

Permissions: sharing does not mean visibility to everyone. Identity and ACL filtering must precede relevance retrieval, following least-privilege principles and the basic requirements of zero-trust architecture.

Provenance: users should be able to inspect which asset an agent used, where it came from, and which planning step it influenced. Evidence chains support review, audit, and correction.

Negative feedback: incorrect, outdated, or mistakenly retrieved assets need flags, downranking, and withdrawal, with feedback informing the trigger conditions of related assets.

Cost: more context is not automatically better. Retrieval and assembly must balance usefulness, tokens, latency, and exposure of private information.

Team memory therefore needs more than retrieval. Like a code repository, it needs versions, branches, permissions, review, merging, and rollback. Without governance, increasing information bandwidth can also accelerate the spread of incorrect information.

Conclusion: make experience available before the next mistake

The OPC shows what can happen when goals, responsibility, and context are concentrated: one person and multiple agents can execute efficiently. Enterprises cannot eliminate information boundaries, and complete transparency of every trajectory is not the goal. They need effective context to move across people, sessions, agents, and projects under permission, version, and evidence constraints.

When that bandwidth is low, a mistake becomes a temporary fix in one task. Incident background stays in chat, an incorrect assumption remains in someone's head, and the root cause appears only indirectly in a diff. The next person and agent cannot access the experience before execution, so the organization repeats investigation, trial, and rework.

Team memory makes the background, failed path, root cause, fix, applicable version, permissions, and validation method inheritable. It delivers the appropriate parts to the appropriate person or agent when a similar task begins. The gain is in the reach, clarity, and usability of experience—not message volume.

Our work follows this chain: internal task analysis shows that related context is much more common than reliably reusable experience; interviews place design, review, and state recovery at the center; real Redis cases demand finer scope, chronology, and confidence levels; and the scoped SWE-bench result of 60% → 80% provides initial evidence that earlier experience can affect later outcomes.

“Make each mistake only once” is an aspiration, not an absolute promise about organizational behavior. When effective context can flow safely and accurately, each task can contribute something the next task inherits.

This work is still exploratory. Asset boundaries, conflict handling, updates, and assembly across agents need real tasks to expose their limitations. Internal data and benchmarks provide a starting point. The practical test is whether the system reduces repeated explanation, lost state, repeated trial and error, and gaps in inherited experience.

If your team struggles to resume work across sessions, trace historical decisions, avoid repeating investigations, recover design-review context, or govern knowledge across people and agents, we welcome real use cases, issues, and contributions.

Try the project, open an issue, or contribute a pull request: TencentDB Agent Memory on GitHub.

Top comments (3)

Collapse
 
cubl9snp71hm profile image
cubl9snp71hm •

Team memory is one of those things that sounds obvious until you try to scale it past 5 people. We've all been in the "who knew about this?" meeting after a production incident that someone fixed six months ago but never wrote down.

What's worked for us: treating postmortems as living documents rather than one-off writeups. We tag them with symptoms, root causes, and — crucially — the exact commit or config change that fixed it. Then we surface those tags in on-call runbooks and PR templates. It's not perfect, but it turns "tribal knowledge" into "searchable knowledge" without requiring everyone to read every doc.

The hard part isn't capturing the lesson — it's keeping the retrieval friction low enough that people actually check before they repeat the mistake. Curious how the agent handles stale/conflicting memories over time. That's usually where these systems break down PS: the tool I meant is on labagent .tech

Collapse
 
koev3kcjausd profile image
koev3kcjausd •

Đọc xong bài viết là nhớ ngay những lần team mình lặp lại bug y hệt cái đã fix 6 tháng trước — chỉ vì knowledge nằm rải rác trong Slack thread, PR comment, hoặc pire, đầu của anh dev đã nghỉ việc.

Thực tế "make each mistake only once" nghe đẹp nhưng implementation mới là phần khổ: ai curate knowledge base? Ai decide cái nào đáng ghi nhớ vs noise? Ai remind team trước khi deploy feature tương tự? Tooling giúp capture context (decision rationale, trade-offs, failed approaches) ngay lúc đang code/review thì mới hiệu quả — đợi sprint retrospective thì context đã bay màu.

Thấy approach dùng agent memory để index và retrieve lesson learned theo context (ticket, code area, error pattern) khá hợp lý. Còn missing piece ở chỗ: feedback loop — how do we know lesson đã được apply thành công? Metrics nào track "mistake prevented" vs "mistake repeated"?

Curious team bạn handle conflict giữa lesson learned cũ vs tech stack mới như thế nào — deprecate hay versioning? — found it via LabAgent, site: labagent .tech

Collapse
 
ywnigcsmku2m profile image
ywnigcsmku2m •

Đúng trọng tâm. Vấn đề không phải là thiếu công cụ ghi chú — mà là khi nào và như nào kiến thức được nắm bắt lại vào quy trình. Hầu như team nào cũng có Notion/Confluence đầy ắp doc, nhưng khi incident xảy ra, ai cũng quay lại hỏi nhau trên Slack vì doc đã lạc hậu hoặc không gắn liền với context thực tế (commit, PR, ticket, runbook).

Thấy bài viết nhấn mạnh "execution becomes abundant" — đây chính là điểm mấu chốt. Khi code gen và deploy nhanh hơn, tốc độ sinh ra sai lầm mới cũng tăng theo. Nếu không có cơ chế tự động gắn lesson-learned ngay vào luồng làm việc (ví dụ: fail test → suggest fix pattern từ incident trước; deploy canary → nhắc rolling-back strategy đã từng dùng), thì knowledge base chỉ thành "kho lưu trữ chết".

Một điểm thường bị bỏ qua: ownership của memory. Ai chịu trách nhiệm cập nhật entry khi pattern thay đổi? Nếu không có owner rõ ràng (hoặc rotation), agent memory cũng sẽ bị drift giống như doc thủ công. Curious xem TencentDB Agent Memory handle phần lifecycle management của memory entry như thế nào — có TTL, review cycle, hay auto-deprecate khi pattern không còn match codebase hiện tại không? PS: the tool I meant is on labagent .tech