DEV Community

Morgan Xu
Morgan Xu

Posted on

Postmortem: An Unbound Cache Key Certified a Stale Completion

An unbound completion cache can certify the wrong model revision. The key left out the server image digest and the prompt hash. Green tests then approved a patch built for an older contract.

This note reconstructs that lab failure class for review practice. It does not claim a dated production outage or customer loss. The durable fix binds cache identity and rejects missing stop metadata.

Incident summary

An evaluation job reused a stored model completion from an earlier run. The cache key held only the repository name and the test path. A later prompt edit never reached the model at all.

The evaluation host still returned the older completion text. The applicator treated that text as a fresh unified diff. Unit tests passed because those fixtures came from the same stale run.

Review accepted the green suite as certification of the change. The suite never called the model on the edited prompt. Certification and execution had quietly split apart in that job.

Why this review still matters

AI coding tools now sit inside ordinary continuous integration jobs. A passing suite no longer proves the model saw the current prompt. Hackathon demos and review bots can hide the same stale cache.

Earlier postmortems in this series covered different certification faults. Those notes involved parsers, lockfiles, workspaces, streams, schemas, and imports. This note covers cache identity for model completions instead.

The overlap with current AI tooling talk is the review gap. Generated patches look finished long before identity checks exist. A portfolio demo can ship that gap into a real gate.

Timeline

The reconstructed timeline uses relative marks rather than claimed clock times.

  1. A reviewer stored a completion under a short cache key.
  2. The prompt later gained a failure clause for partial diffs.
  3. The evaluation image was rebuilt with a newer diff applicator.
  4. A later job hit the cache and skipped the model call.
  5. The applicator applied the old diff and the suite passed.
  6. Review marked the change certified because the CI result was green.
  7. A manual check against the prompt hash exposed the mismatch.

Contributing factors

Several misses combined before the stale completion could pass. Each miss looked harmless when reviewed alone in the job log.

  • The cache key ignored the raw prompt bytes.
  • The cache key ignored the model revision string.
  • The cache key ignored the server image digest.
  • Missing finish metadata was treated as a normal stop.
  • The test fixture was written by that same cached run.
  • The merge gate checked process exit codes only.

No single miss would have certified the patch by itself. The combination made yesterday's text look like fresh evidence.

What the gate proved

The gate proved that one stored text applied without a conflict. It did not prove that the current model produced that text. It did not prove that the current image parsed the diff.

It also did not prove that the new prompt clause was seen. That distinction holds on any host, whether paid or free. An inexpensive host does not relax completion identity rules.

Disclosure: This article was prepared as part of MonkeyCode's product outreach. MonkeyCode is an open-source project for coding assistance workflows. The operator states that free model access and a free server option exist.

This note does not restate model names, token ceilings, or hardware. Those values change and must be read from current project docs. A free server option still needs a pinned image digest.

Failure analysis of the short key

The short key collided across prompt edits in the reconstructed lab. Two different prompts shared one repository name and one test path. The second job never observed that collision in its exit code.

A hash of the prompt bytes would have split those entries. A hash of the image digest would have split them again. Either field alone is still too weak for a merge gate.

Fixture generation must be part of the key as well. Otherwise an old completion can satisfy fixtures it wrote itself. That self-agreement is not independent evidence of correctness.

Durable fix

Bind every cache entry to four explicit identity inputs. Use the prompt hash, model revision, image digest, and fixture generation. Refuse a hit when any of those fields is absent.

Refuse a completion when stop metadata is missing or unknown. Treat a length stop as truncation, not as a successful patch. A truncated hunk can still compile and mislead the suite.

Proposed check

The Python below is a proposed local check, not an executed run. Adapt paths and field names before placing it in CI.

import argparse
import hashlib
import json
import pathlib
import sys

def sha256_bytes(data: bytes) -> str:
    return hashlib.sha256(data).hexdigest()

def completion_cache_key(prompt: bytes, model_rev: str, image_digest: str, fixture_gen: str) -> str:
    if not model_rev or not image_digest or not fixture_gen:
        raise ValueError("cache identity fields must be non-empty")
    payload = {
        "prompt_sha256": sha256_bytes(prompt),
        "model_rev": model_rev,
        "image_digest": image_digest,
        "fixture_gen": fixture_gen,
    }
    raw = json.dumps(payload, sort_keys=True, separators=(",", ":")).encode()
    return sha256_bytes(raw)

def accept_completion(record: dict, expected_key: str) -> None:
    reason = record.get("finish_reason")
    if reason != "stop":
        raise ValueError("only an explicit stop can certify a patch")
    if record.get("cache_key") != expected_key:
        raise ValueError("cache key does not match live identity")
    if not record.get("diff"):
        raise ValueError("completion has no diff to review")

def main() -> int:
    parser = argparse.ArgumentParser()
    parser.add_argument("--prompt", required=True)
    parser.add_argument("--model-rev", required=True)
    parser.add_argument("--image-digest", required=True)
    parser.add_argument("--fixture-gen", required=True)
    parser.add_argument("--record", required=True)
    args = parser.parse_args()
    prompt = pathlib.Path(args.prompt).read_bytes()
    record = json.loads(pathlib.Path(args.record).read_text())
    key = completion_cache_key(
        prompt,
        args.model_rev,
        args.image_digest,
        args.fixture_gen,
    )
    accept_completion(record, key)
    print(key)
    return 0

if __name__ == "__main__":
    sys.exit(main())
Enter fullscreen mode Exit fullscreen mode

Compilation of a partial hunk is not certification. The function rejects every finish reason except an explicit stop. A cache hit still fails when the live key differs.

Commands that make the key observable

Compute the image digest outside the model call itself. Store that digest beside the job log for later review. Do not trust a floating image label such as latest.

docker image inspect "$EVAL_IMAGE" --format '{{.Id}}' | tee image_digest.txt
sha256sum prompt.txt | awk '{print $1}' | tee prompt.sha256
sha256sum fixtures/generation.txt | awk '{print $1}' | tee fixture_gen.sha256
python3 check_cache_key.py \
  --prompt prompt.txt \
  --model-rev "$MODEL_REV" \
  --image-digest "$(cat image_digest.txt)" \
  --fixture-gen "$(cat fixture_gen.sha256)" \
  --record completion.json
Enter fullscreen mode Exit fullscreen mode

Compare the live key with the cached record before any apply. Fail the job when the printed key differs from the record. Do not apply the stored diff after a key mismatch.

Pass the fixture generation hash as the fourth identity field. A missing flag should abort before the diff applicator starts. That abort is part of the durable fix.

Test plan

Run the following checks on a scratch branch first. Treat them as the minimum bar before a merge decision.

  1. Change one prompt byte and confirm the cache key changes.
  2. Change only the image digest and confirm the key changes.
  3. Replay a record with an empty finish reason and confirm rejection.
  4. Replay a record whose finish reason is length and confirm rejection.
  5. Replay a record from the prior fixture generation and confirm rejection.
  6. Apply a fresh completion only after all five checks pass.

Negative test

Keep the old short key in a dedicated negative test. That negative test must fail closed on every run. A green negative test means the identity gate has regressed.

Decision table

Use this table when choosing a free host for the check. The table is a control guide, not a quota promise.

Situation Free model access Free server Extra control
Prompt and fixtures still moving Exploration only Clean image only Bound key, no merge
Patch is a candidate for main Fresh uncached call Pinned image digest Full test plan
Secrets in the prompt Do not use Do not use Keep prompt off shared hosts
Named model SLA required Do not assume Do not assume Read current docs first
Only a truncated result exists Do not certify Do not certify Split the task and rerun

Who should not use this approach

Skip this flow when the prompt contains secrets or private source. A shared free server is the wrong place for that material. Keep those prompts on a host the team already controls.

Skip this flow when the team needs a named model SLA. Skip it when a fixed token ceiling must be guaranteed. This article does not establish those commercial terms.

Skip it when the only evidence is a cached green suite. That pattern is the incident, not a mitigation. Also skip it when every hunk needs a named human owner.

The cache check does not replace that human owner. It only stops a stale completion from looking certified.

Limitations

This reconstruction does not measure latency, cost, or quality. It does not name a current model or a token amount. It does not claim that free access will remain available.

Operator-supplied availability can change without notice to readers. The sample code does not call a network API. It does not prove that a vendor cache is safe.

The sample only shows a local identity check a gate can enforce. Image digests must come from the runtime that applies the patch. A laptop digest does not certify a remote free server.

Record both digests when both runtimes exist in the path. Certify only the digest from the runtime that applied the diff. Anything else repeats the original identity mistake in a new log.

What to keep after the incident

Keep the bound key in the job log after every run. Keep the finish reason beside that key in the same record. Keep the negative test that rejects the old short key.

Drop any dashboard that counts green runs without those fields. A later reader should see why the old completion was refused. The log should show the four identity fields and the stop reason.

Absence of those fields is a failed job, not a blank success. Teams that already review model diffs can rehearse this check first. A free server is enough for that rehearsal when the prompt is public.

Read the current MonkeyCode docs before relying on free access. Then keep this same gate even if the host later changes.

Top comments (0)