DEV Community

Avery Li
Avery Li

Posted on

Record the Read Set Before a Free Model Drafts the Next Test

A pairing should record the allowed read set before a free model drafts tests or a free server executes them. The kept decision in this worked example stays narrow, concrete, and easy for a later reviewer to audit. Generated checks may read only a synthetic fixture tree, and the remote job specification may carry no credentials. These notes describe a proposed session rather than a measured incident, and every command below remains an unexecuted example.

Why the boundary comes first

Developers who generate tests from a private repository often attach broad context and then hope the model will ignore secrets. A senior pairing treats that hope as a failed control, because the upload itself is the event that cannot be undone. The useful artifact is a read-set manifest that names every file allowed into the prompt and every path refused by the server. MonkeyCode's free model access and free server option become relevant only after that manifest exists and a human signs it.

Disclosure: This article was prepared as part of MonkeyCode's product outreach.

The availability claims in the previous paragraph are operator-supplied, and this draft deliberately stops at those two claims. No model name, quota, hardware shape, duration, or permanence appears here, because those details were not verified from a primary source. Readers should confirm the current terms themselves before they attach any repository or start any remote execution. The gate remains useful if the product name is removed, because the control is ordinary Python plus a written decision record.

Recent public writing about generated interfaces and assisted coding shows the same pressure, which is speed competing with review. This article does not treat those titles, reaction counts, or anecdotes as evidence about any model. The durable problem is older than the current week: a draft can look helpful while the context beside it is still private. A pairing session is one way to force that problem into a file a reviewer can diff.

Questions the senior asked

The senior did not begin with prompt wording, and instead pinned five questions to the pairing board before any draft started. Each question had to end in a kept answer, a rejected alternative, or an explicit dead end with a reason. The board prevented the session from drifting into model comparison, benchmark talk, or unverified claims about remote capacity. The list below is the working record for this proposed session, not a transcript from a named team.

  1. Which paths may enter the model context, and which paths stay forbidden even when a test seems to need them?
  2. May the free server mount anything outside the synthetic fixture directory during this session?
  3. Which environment names are allowed in the job spec, and which names are always rejected?
  4. What happens when the generator asks for a production-shaped dump in order to reproduce a bug?
  5. Who signs the decision record, and which change requires the pairing to reopen before the next draft?

The senior required each answer to name an owner, a file path, and a condition that would invalidate the answer later. Vague notes such as a general warning about secrets were rejected, because a warning cannot be tested or diffed in review. The manifest file is the test, since a reviewer can see the allowed list without reconstructing the conversation. That standard also stops a later edit from smuggling a dump into the prompt under a friendlier filename.

Dead ends the pairing refused

The first dead end was a proposal to sanitize a database dump and then upload the sanitized copy with the prompt. Sanitization scripts miss new column names, copied tokens in comments, and fixtures that developers later paste back by hand. The senior rejected that path because a partial redaction still teaches the team to move private data toward the model. The kept alternative was a hand-built synthetic tree that contains shapes and statuses, but never customer values.

The second dead end was a request to let the free server run the existing integration suite against a shared staging account. That request looked efficient, yet it would place credentials, network reach, and generated code in the same job. The pairing refused it, because a generated test that only reads can still print headers, follow redirects, or write files. Remote execution stayed limited to a unit run over the synthetic tree, with network egress left unspecified on purpose.

The third dead end was an attempt to compare unnamed free models by latency before the read set was stable. Timing numbers would have been meaningless, and this draft has no verified benchmark method or hardware description to publish. The senior parked comparison work until the same manifest, the same fixtures, and the same commands could be repeated. That pause kept the session on a decision a reviewer can check, rather than on a leaderboard nobody can reproduce.

The decision that stayed

The pairing kept one decision and wrote it in a single sentence at the top of the record file. Generated tests may be drafted from the manifest, and they may run only where the fixture root is the sole mounted data path. Any request for a dump, a token, a cloud key, or a production hostname reopens the pairing instead of extending the prompt. The record names the signer role, the session date, and the commit that the manifest claims to describe.

Run the gate in seven steps

The workflow below is a tutorial sequence for a local gate, and it assumes a Unix shell plus Python 3.10 or newer. Operators should run it on a disposable copy of the repository, then inspect the manifest before any model call. The steps do not deploy infrastructure, open a network listener, or contact a remote API by themselves. A later human may paste the reviewed manifest into a coding session only after the blocked list is empty or explicitly accepted.

  1. Create an isolated work directory and a synthetic fixture root that holds invented records only.
  2. List candidate prompt files in a plain text inventory, with one relative path on each line.
  3. Run the gate script so it rejects secret-like names, forbidden directories, dumps, and oversized files.
  4. Read the decision record and confirm the kept sentence is still the header a reviewer will see.
  5. Emit a job spec that mounts only the fixture root and sets no secret environment variables at all.
  6. Run the unit command locally against those fixtures, and archive the manifest beside the test output.
  7. If a generator asks for more files, add them to the inventory and repeat the gate before the next draft.

A clean local run is not permission to skip the human read of both JSON files the script writes. The script can only see the inventory it was given, so files pasted into a chat by hand bypass it completely. The pairing therefore treats the chat box as untrusted until the same paths appear in the allowed list. That rule is the practical difference between a written control and a good intention that nobody can audit later.

Read the proposed script

The following Python example is a proposal for the local gate, and it has not been executed in this drafting session. It reads an inventory file, applies a small denylist, and writes a JSON decision record that a reviewer can diff. It also writes a job spec whose data mount is fixed to the fixture root and whose environment block stays empty. Teams should treat the patterns as a starting filter, not as a complete secret scanner or a compliance certification.

#!/usr/bin/env python3
"""Unexecuted example: build a read-set manifest and a secret-free job spec."""

import json
import sys
from pathlib import Path

FORBIDDEN_PARTS = {
    ".env",
    ".pem",
    ".key",
    "id_rsa",
    "secrets",
    "credentials",
    ".git",
}
MAX_BYTES = 200_000
FIXTURE_ROOT = "fixtures/synthetic"
KEPT = "fixtures-only-no-credentials"


def reject(path: Path) -> str | None:
    lowered = path.as_posix().lower()
    if any(part in lowered for part in FORBIDDEN_PARTS):
        return "forbidden-name"
    if path.suffix.lower() in {".pem", ".key", ".p12", ".sqlite"}:
        return "forbidden-suffix"
    outside = "fixtures/synthetic" not in lowered
    if outside and path.suffix.lower() in {".sql", ".dump", ".bak"}:
        return "data-dump-outside-fixtures"
    if not path.exists():
        return "missing"
    if path.stat().st_size > MAX_BYTES:
        return "too-large"
    return None


def main() -> int:
    inventory = Path(sys.argv[1] if len(sys.argv) > 1 else "readset.txt")
    allowed, blocked = [], []
    for line in inventory.read_text(encoding="utf-8").splitlines():
        rel = line.strip()
        if not rel or rel.startswith("#"):
            continue
        reason = reject(Path(rel))
        if reason:
            blocked.append({"path": rel, "reason": reason})
        else:
            allowed.append(rel)
    record = {
        "decision": KEPT,
        "fixture_root": FIXTURE_ROOT,
        "allowed": allowed,
        "blocked": blocked,
        "signer_role": "pairing-senior",
        "status": "proposed-unverified",
    }
    Path("decision-record.json").write_text(
        json.dumps(record, indent=2), encoding="utf-8"
    )
    spec = {
        "mounts": [{"source": FIXTURE_ROOT, "target": "/work/fixtures"}],
        "env": {},
        "command": ["python", "-m", "unittest", "discover", "-s", "tests"],
        "network": "unspecified-do-not-assume-egress",
    }
    Path("job-spec.json").write_text(
        json.dumps(spec, indent=2), encoding="utf-8"
    )
    print(f"allowed={len(allowed)} blocked={len(blocked)}")
    return 1 if blocked else 0


if __name__ == "__main__":
    raise SystemExit(main())
Enter fullscreen mode Exit fullscreen mode

The shell fragment below is also an unexecuted example, and it only creates local files for the gate to read. It does not call a model, publish a draft, or open a socket toward a free server of any kind. Operators who adapt the paths should keep the fixture directory free of copied production rows and live tokens. A missing file is a blocked file on purpose, so the inventory cannot name a path the reviewer cannot open.

mkdir -p fixtures/synthetic tests
printf '%s\n' '{"id":"sample-1","status":"queued"}' > fixtures/synthetic/orders.json
printf '%s\n' 'tests/test_status.py' 'fixtures/synthetic/orders.json' > readset.txt
python3 gate_readset.py readset.txt
Enter fullscreen mode Exit fullscreen mode

A failing exit status means the inventory still contains a blocked path, so the pairing does not proceed to a model call. A clean exit only means the small filter accepted the listed paths, not that those files are safe for every policy. Reviewers should open both JSON files and confirm that the environment object is empty before any paste into a chat. If the job spec gains a token, a host key, or a second mount, the kept decision is already broken.

Classify requests before they reopen

The table records how the pairing classified common requests, so a later author does not relitigate them in a side channel. Rows are teaching examples for this method, and they are not measurements from a live service or a customer system. A request that does not fit a row should reopen the pairing rather than receive an improvised exception in the prompt. The last rows also show who should stop, which matters as much as the happy path through the gate.

Request during the session Classification Reason kept on the board
Synthetic JSON fixtures under 200 KB Keep Shapes and statuses without customer values
Sanitized production dump in the prompt Dead end Redaction misses new fields, comments, and later pastes
Existing integration suite against staging Dead end Credentials and generated code would share one job
Latency comparison of unnamed free models Dead end No verified hardware or repeatable method in this draft
Extra source file added after sign-off Reopen The manifest must be regenerated and signed again
Environment entry whose name looks like a token Reject The job spec environment object must stay empty

Limitations of the filter

The filter misses secrets that use ordinary filenames, split tokens, or values buried inside seemingly harmless test helpers. It does not inspect archive contents, notebook outputs, git history, or environment variables already present in the parent shell. Substring checks can also false-positive on innocent names that merely contain a forbidden fragment such as a secret-related word. A reviewer still has to read the allowed list, because a green exit status is not a review.

A free server option does not, by itself, prove network isolation, data retention, region placement, or deletion timing. Those properties require current vendor documentation and, where the data is regulated, a review that sits outside this article. The worked example invents no quota, no uptime figure, and no claim that a free tier will remain available on a later date. Authors who need a guaranteed capacity number should stop and read the provider terms rather than infer one from a blog post.

The script is a poor fit for binary protocols, huge monorepos, or repositories that cannot spare a synthetic fixture tree. In those cases the pairing should keep the refusal and choose a different test design, not weaken the denylist to finish the draft. Generated code can still be wrong, unsafe, or incomplete even when the read set is clean. The gate protects the boundary around the session, and it does not certify the tests the model returns.

Who should leave this method alone

Teams that must process production data, payment data, or other regulated records should not route that material through this gate. The method is a local hygiene step for synthetic tests, not a data-loss prevention product and not a legal control. Operators who cannot review the manifest before a model call will not get the benefit the pairing was designed to keep. Anyone who needs unpublished benchmarks, named hardware, or a permanent free allocation will not find those facts in this draft.

What the session closes on

The session ends when the decision record, the empty environment block, and the synthetic fixture root all agree with one another. A later draft may improve test names and assertions, but it should not reopen the read set without the same senior review. Readers who already keep a fixture tree can run this local gate before they choose any coding assistant at all. If a free model and a free server still fit after that review, the current MonkeyCode terms are the right place to confirm availability.

Top comments (0)