Your bot fixes failing tests on its own. Before each attempt it looks up how similar failures were fixed before, and the lookup is ranked by similarity. So the fix it tried last Tuesday, the one that turned the suite red, keeps coming back first, because it is still the closest text to today's error. The fix that actually worked sits at position four.
You want the ones that passed to rise and the ones that failed to sink, without a person rating anything.
The short answer: a code bot already has the verdict a chat bot lacks. The test run exits zero or it does not. Send that result back to the memory store as a rating on the memories the attempt actually followed, and a store that reads ratings at ranking time will weigh what worked into the order of the next lookup. The rest of this article is the three decisions that make that loop honest, and a check with a control.
Disclosure: I work on Mnemoverse, the memory layer the code below uses.
The outcome is already there
A bot that answers people has to infer an outcome from the conversation, and most conversations say nothing. A bot that changes code has a build, a test run, a linter, a command that exits zero or does not. That is a verdict produced by something other than the bot's own opinion of its work, and it arrives at a known moment.
pytest, for example, documents its exit codes: 0 is "All tests were collected and passed successfully", and 1 is "Tests were collected and run but some of the tests failed". Codes 2 to 5 mean the run was interrupted, broke internally, was called wrongly, or collected nothing, and current pytest adds 6 for too many warnings. None of those is a verdict on your fix, with one trap for a bot that edits code: by default a test file that fails to import stops the run with code 2, so a patch with a syntax error would look like no verdict. --continue-on-collection-errors turns that case into code 1. An import error inside conftest.py still exits 4, so if the bot can touch code that conftest imports, compile the changed files first and count a failure there as red.
Three decisions, in the order they come up
When the call goes. After the test run, not inside the lookup. At lookup time you only know what was similar. The outcome exists once the suite has finished, so the rating goes out then. Rating in the same breath as the lookup rates the retrieval, not the fix.
Which ids. Keep the ids of the items the patch actually followed with the attempt that used them: not with the session, and not every item the read returned. Items that are always rated together always move together, so a fix that failed and a fix that worked, returned for the same failure, would never come apart. If the bot makes three attempts in one run, rating whatever was read last hands the second attempt's red run to the third attempt's memories.
What to send when there is no verdict. Nothing. A run killed by a job time limit, an interrupted run, a suite that collected no tests: none of these says whether the fix was right. Our library page puts the rule in one sentence: "outcome is a required argument with no default, and nothing calls feedback() for you". An attempt you do not rate leaves those memories' valence where it was. A small positive on every run with no verdict is not the same as silence, because it moves every memory it touches and tells the ranking nothing.
The loop in Python
This uses our Python SDK, mnemoverse 0.3.1. The bot's own calls are placeholders.
import subprocess
from mnemoverse import MnemoClient
client = MnemoClient(api_key="mk_live_YOUR_KEY")
failure = "payments integration test times out after the retry change"
recall = client.read(failure, top_k=5)
# the bot returns its patch and the recalled items it actually followed
patch, followed = bot.propose_fix(failure, recall.items)
used = [item.atom_id for item in followed] # the ids this attempt is built on
bot.apply(patch)
run = subprocess.run(["pytest", "-q", "--continue-on-collection-errors", "tests/integration"])
if run.returncode == 0:
outcome = 1.0 # collected and passed
elif run.returncode == 1:
outcome = -1.0 # ran and failed, including a test file the patch broke on import
else:
outcome = None # no verdict: interrupted, internal error, usage error, nothing collected, killed
if outcome is not None and used:
client.feedback(atom_ids=used, outcome=outcome, query_concepts=recall.query_concepts)
The loop above only reads and rates. What it reads are notes the bot wrote earlier with client.write(), one short note per fix it applied, and the ratings attach to those notes, not to the failure text. A write can be refused by the write gate, so check stored on the response.
What the rating changes
Two things, both documented on our library page. The memory's valence, an outcome polarity from -1.0 to +1.0 that every returned item carries, moves with the rating. And the links between the query's concepts and the memory's concepts are updated, which is why the call takes query_concepts as well as ids.
At ranking time the valence modulates relevance, so a fix with a history of green runs gains ground and one with a history of red runs loses it, and a fix that worked can out-rank a closer match that kept failing. Nothing is deleted. The record of what failed stays in the store and stays readable, which is useful the day someone asks why the bot stopped trying the obvious fix. The coefficient is not published, so this article makes no promise about how many positions one rating moves.
If your bot talks MCP
A bot built on an MCP client, such as a coding agent running headless, gets the rating half of the same loop through the memory_feedback tool, which takes memory_ids and an outcome from -1.0 to 1.0. It does not send query_concepts, so it moves the memories' valence but not the links between the question and the memories; those are taught from the Python path. Its description tells the agent when to call it: "right after you act on (or reject) recalled memories". Put the test-result rule in the agent's standing instruction: rate the memories you followed after the suite finishes, send +1 on a green run and -1 on a red one, and send nothing when the run gave no verdict.
Check it with a control
-
Write three notes for the same failure: fix A, which you know passes; fix B, which you know fails; and fix C, similar wording, which you will never rate. Check
storedon each write, then read once and note each item'svalence. - Rate each fix by its own run. Apply A, run the suite, and send the verdict for A's id only. Do the same for B.
- Read the failure again. A's valence should be above where it started and B's below. The update can land after the call returns, so if nothing has moved yet, wait and read again before concluding anything.
- Control. C's valence should be where it started. If it moved, a rating reached an id the attempt did not use. C's position may still shift, because the ratings also updated the links between the query's concepts and the memories' concepts. That is not a leak, which is why the control reads valence, not position.
There is a 12-minute walkthrough of the same loop for a bot that answers people, where the outcome has to come from the conversation instead of a test run:
The library page covers where the update rule comes from and the chat-bot case in full. What does your bot currently do with a red test run, apart from retrying?
The Python SDK is on PyPI as mnemoverse, and the MCP server package on npm is open source (MIT): github.com/mnemoverse/mcp-memory-server.
Top comments (4)
The exit-code mapping is the careful part. It's easy to count code 2 as a failure and punish memories for a broken import they didn't cause.
Three things I'd watch once the ratings start moving the ranking:
Small counts. A fix that passed once and a fix that passed 9 of 10 times can end up with similar valence, and the one-off is often just the luckier of the two. Ranking on a lower confidence bound for the pass rate (Wilson, or a Beta prior on passes and attempts) keeps a single green run from outranking a record.
Flaky tests. A flaky suite hands out random verdicts, and a memory that happens to be followed on a lucky run gets credit it didn't earn. Rating only when a re-run agrees costs one extra test run and removes most of that noise.
Which failures a fix was matched to. Some fixes get tried mostly on easy failures, and easy failures pass more, so a fix's pass rate partly measures what it was matched to rather than how good it is. Comparing fixes within similar failures (same test, same error class) keeps that from inflating the popular ones.
The control check you describe is the right way to find out whether any of this matters in practice: if ranking by outcome doesn't beat ranking by similarity on a held-out set of failures, the ratings are mostly noise.
The control in the article checks plumbing: that a rating reaches the ids the attempt actually followed and nothing else. Your held-out comparison is the check that matters for the bot, similarity alone against similarity with the ratings on top, on failures the ratings never touched, and it belongs next to the control as the step that says whether any of this is worth running.
Flaky suites fit the article's rule for runs with no verdict: two runs that disagree say nothing about the fix, so nothing is sent, which is your re-run rule. Small counts are harder to promise on. The ranking weighs the memory's valence, each rating nudges it by a prediction-error update, and the size of the nudge is not published, so I would not say where one green run lands next to nine of ten. What the bot does control is the value:
outcometakes any value from -1.0 to 1.0, so a fix with a single pass behind it can go in below 1.0 until its record grows.On the matching bias, the Python call also takes
query_concepts, so part of the credit lands on the link between this failure's concepts and the fix, not on the fix alone. That is closer to comparing within a failure class than a global pass rate would be, but it is not the stratified comparison you describe, and the held-out test is where the difference would show.Scaling the outcome by record length is a sensible workaround for the unpublished step size. Since the step isn't known, the "one pass vs nine of ten" question can be answered by measurement instead: seed a test store with synthetic fixes whose true pass rates you set (say 0.3, 0.6, 0.9), feed them rated attempts drawn from those rates, and check after how many ratings the ranking orders them correctly. That gives the update rule an empirical sample-size curve, and it tells a bot author how many runs a fix needs before its rank means anything.
That measurement answers the question better than a published step size would, and three details make it read cleanly. Write the three notes against the same failure in close but not identical wording, check
storedon each write as the article's control does, and keep them in a domain of their own, reading with that domain so nothing else in the store comes back between them. They will still not start tied, so read the order once before any rating and give the highest true rate to the note that starts last and the lowest to the one that starts first: the curve then counts the ratings it takes to overturn a head start, not ratings that happen to agree with it. And rate the synthetic attempts the way the bot will rate its real ones, including any scaling by record length, so the curve belongs to the bot's rule and not to a flat plus or minus one.What comes out is a curve for that setup and that query, not a constant of the update rule, since recency and associations also weigh in on the order, and on real data they differ from fix to fix. For a bot author that is the useful form anyway: the question is how many runs this bot's fixes need, and the seeded test can be run again whenever the bot's notes or its rating rule change.