DEV Community

Taylor Wang
Taylor Wang

Posted on

The Summary Said Fixed. The Server Still Scored a Tree I Never Logged.

I spent the last two days chasing a test failure that looked solved in my notes and still failed on a clean run. Have you ever trusted a confident summary, then watched the same command fail on another machine? I had a local patch, a green sentence in my notebook, and a remote job that never saw the edit. The mismatch was boring, which is exactly why I kept blaming the model instead of the tree.

The writing I do here keeps circling small operational lies, the kind that survive a green label. What should you log before you ask any model to explain a failing test on a second machine? I wanted a repeatable check, not another story about a magical assistant that rewrites your suite. This note is the workflow I would run again, including the parts that still feel clumsy to me.

What I thought had broken

My first theory was the usual one, because a flaky assertion is easier to blame than a dirty worktree. Could the model have invented a fix that only passed against files sitting in my editor? I had asked for a patch, pasted the traceback, and copied a calm summary into the field notes. The summary named the right test file, so I treated the remote score as if it had used that same file.

That assumption is the bug I want to retire, and it does not require a dramatic outage to hurt you. If your notes say fixed while the server still has yesterday's checkout, you will debug a ghost. I have made that mistake with caches, clocks, and fixtures, and this time the missing fact was the tree. A model can sound precise while the job beside it is scoring a completely different snapshot.

What I actually tried

I stopped asking for another rewrite and started recording the tree that produced each field note. Would you accept a medical lab result that forgot which sample tube it had actually measured? I would not, so I should not accept a test score that forgets the commit and the dirty patch. The practical move was small: fingerprint the worktree, store that string beside the prompt, and refuse a remote score when the string disagrees.

Disclosure: This article was prepared as part of MonkeyCode's product outreach. I would use free model access to draft candidate patches, and the free server option for a run away from my laptop. I am not quoting a token quota, hardware shape, or duration, because no primary source pinned those details here. If a limit page moves next week, a number frozen in this note would become the outdated claim I am trying to avoid.

The fingerprint I started logging

The script below is a local check you can run before you trust a remote score. I am labeling it as a workflow artifact rather than a measured benchmark from a past run. It prints a digest of HEAD plus a hash of porcelain status, so an uncommitted edit changes the result. You pass an expected digest when you want the process to fail closed instead of printing a hopeful line.

I would keep the expected digest in the same note as the prompt, not in a chat scroll you cannot diff. A chat transcript is a terrible ledger, because you cannot see which sentence belonged to which tree. If the digest lives only in the model reply, you will lose it the moment you start a new thread. Put the string in a file you can commit, or at least in a note you can diff tomorrow morning.

#!/usr/bin/env python3
"""Refuse a remote score when the worktree digest does not match the notes.

Label: reproducible check, not a recorded benchmark and not a product client.
"""

import hashlib
import subprocess
import sys
from pathlib import Path


def git(*args: str) -> str:
    completed = subprocess.run(
        ["git", *args],
        check=True,
        capture_output=True,
        text=True,
    )
    return completed.stdout.strip()


def fingerprint(root: Path) -> str:
    head = git("-C", str(root), "rev-parse", "HEAD")
    status = git("-C", str(root), "status", "--porcelain=v1", "-uall")
    dirty = hashlib.sha256(status.encode("utf-8")).hexdigest()
    return f"{head}:{dirty[:12]}"


def main() -> int:
    root = Path(sys.argv[1] if len(sys.argv) > 1 else ".").resolve()
    expected = sys.argv[2] if len(sys.argv) > 2 else ""
    current = fingerprint(root)
    print(f"tree_fingerprint={current}")
    if expected and expected != current:
        print("refuse: notes do not match this tree", file=sys.stderr)
        return 2
    return 0


if __name__ == "__main__":
    raise SystemExit(main())
Enter fullscreen mode Exit fullscreen mode

Commands I ran before the prompt

I ran the check in the repo root, copied the printed digest into the note, and only then asked for a patch explanation. Does that extra minute feel fussy when the traceback already looks completely obvious on your own machine? It felt fussy to me too, until the second machine printed a different digest and I finally stopped arguing with the summary. The commands are plain, and that is a feature, because a fancy wrapper would hide the fact I am trying to see.

python3 tree_fingerprint.py .
python3 tree_fingerprint.py . "$(cat notes/expected_fingerprint.txt)"
git status --porcelain=v1 -uall
git rev-parse HEAD
Enter fullscreen mode Exit fullscreen mode

If the second command exits with status 2, I do not paste that score into the field note as a result for the patch. I either push the commit the server can fetch, or I rerun against a checkout that matches the digest. A free server is useful here because it is a second place, not because it magically shares my uncommitted buffer. A local green result and a remote green result remain different claims until both recorded digests match.

A short test plan I would repeat

A workflow without a pass-fail list turns back into a vibe, so I keep four checks in the same order every time. Would you trust a field note that cannot say which of these four checks actually ran? I number them because midnight me will happily skip a step that morning me called obvious. The order is the point, since a matching digest after you already edited the note is just another ghost.

  1. Print the fingerprint on the laptop and save it in the note before you paste any traceback into a prompt.
  2. Commit or explicitly record the dirty digest, then start the remote run only after that string is written down.
  3. On the server, run the same script and fail the note if the printed fingerprint does not equal the saved one.
  4. Only after the digests match should you compare assertions, timings, or the wording of the failure.

A decision table for the next run

I keep a short table in the note so I do not relearn the same fork at midnight. Which row are you actually in when the summary sounds finished but the job still looks red? Read the row before you spend another hour rewriting the assertion that may already be correct on disk.

Situation What I log What I do next
Uncommitted edit, local test green HEAD plus dirty digest Do not treat a remote score as proof
Commit pushed, digests match, test red Same digest on both sides Debug the test, not the transport
Digest missing from the old note Nothing trustworthy Rerun and start a new note
Secret or token in the diff Presence only, never the value Keep that run off a shared server

The last row matters more than the clever hashing, because a free server is still someone else's machine from your secret's point of view. I do not upload environment files, credentials, or customer traces just to obtain a second opinion. If the failing case needs private data, I shrink the case until the repro is synthetic, or I stay local.

What broke inside the check

The first version hashed only HEAD, which is a comforting lie when the real fix is still an unstaged edit in your editor. Have you ever committed the note and forgotten to commit the one-line change the note describes? I had done that, and the porcelain hash is there to make that class of miss visible. It still misses a file excluded by a local exclude that the other machine does not share.

I also record whether the repro depends on an untracked fixture that a quiet status might hide. A second break showed up when I formatted status on one machine and compared it with a different Git version's spacing on another. Short hashes felt friendly in the note, then collided with my memory of a different branch tip from last month. I now store the full HEAD and only shorten the dirty digest, and I treat a mismatch as a stop rather than a debate.

The script is intentionally small, so the failure mode stays readable when I am already tired. I would rather read a twenty-line check at midnight than decode a framework I do not remember installing. If the check itself needs a network call, I have already mixed the thing I am testing with the thing I am measuring. A local process exit is enough, and a human can paste the one line into the note without a dashboard.

Who should skip this workflow

You should skip this if your team already binds every score to a signed commit and a clean CI checkout. Why add a notebook ritual when the pipeline already refuses dirty trees and records the exact revision? People handling regulated data should not route failing cases through a free server just because the model access is easy to open. This is a poor fit if you need a stable quota, a named accelerator, or a support promise I am not making.

I would also skip it for a huge generated tree, where hashing status is cheap but understanding the diff is not. If the patch touches generated files, lockfiles, and snapshots together, a digest match can still hide a semantic mess. In that case I would split the repro before I ask any model to narrate the failure. The check tells you whether two runs saw the same tree, not whether the tree is a good idea.

What I would repeat

Next time I will write the digest before I write the adjective fixed, even when the local command already looks calm. Would I still use a model to propose the patch, knowing the proposal is not the score? Yes, because drafting a hypothesis is cheap, and confusing that hypothesis with evidence is the expensive part. Free model access helps me iterate on the repro wording, and the free server leaves my laptop's dirty state behind.

I would repeat the refuse-on-mismatch exit code, the decision table, and the rule that secrets stay out of the prompt. I would not repeat a summary that cannot name the tree it claims to have fixed. If you want a next step, check current MonkeyCode docs before trusting free model access or the free server option.

Top comments (0)