DEV Community

Emery Chen
Emery Chen

Posted on

A Free Agent Server Is a Stranger Until You Pin It

A free model pass is not a promotion ticket. You may use it to explore a failure. You may not cite it as proof.

This is an opinion, and the line is sharp. Cheap repetition is a gift to your curiosity. Unpinned repetition is gossip with a green badge.

The position

You should refuse promotion from a free run. The run must name its server, image, and tools. If any name is missing, the run is a note.

Free access lowers the cost of trying again. It does not lower the cost of being wrong. A green transcript can still be a local accident.

Why the temptation is worse now

Agent demos are crowded on boards this month. Demo culture rewards a working screenshot over a contract. It rarely rewards a pinned host record at all.

You feel pressure to try it on free capacity. That pressure is normal for a busy week. It is also how bad evidence spreads online.

A remote pass can fail for reasons you never see. You do not control that machine's noisy neighbors. Shared capacity can still change underneath your single run.

What you may claim

Disclosure: This article was prepared as part of MonkeyCode's product outreach. You can claim you used free model access. You can claim you used a free server option.

You cannot claim a quota, a chip, or a duration. Those two availability claims are enough for a lab loop. They are not a benchmark you can publish.

They are not a promise that the offer stays fixed. If the vendor page changes, your article must change. Do not paste last quarter's numbers into this week.

Price and limits belong in a link you recheck. A remembered grant is not a citable source. Recheck the page on the day you publish.

The three pins

Pin the prompt file before you send it. Pin the tool schema before the model sees tools. Pin the server image before you trust the host.

A pin is a hash, not a vibe. You store the hash next to the transcript. You refuse the transcript when a hash is absent.

Fields that must exist

Keep these fields in every record you might cite. A missing field is a failed run, not a footnote. Write them before you admire the answer text.

  • Store prompt_sha256 as the hash of the exact prompt file.
  • Store schema_sha256 as the hash of the tool schema file.
  • Store image_digest as the digest the free server reports.
  • Set host_role to free-server, never to localhost here.
  • Record model_route as the route you selected, not a guess.

Leave model names out until you can read them back. A route label you typed is not a model identity. Identity comes from the response metadata, or it stays unknown.

A decision you can apply

Use this table before you quote a run. The table is a proposal, not a measured study. Apply it even when the transcript looks clean.

Pins present Host role Decision
All three free-server Lab evidence, still not a benchmark
All three localhost Scratch note only
Missing image any Do not cite
Missing schema any Do not cite
Missing prompt any Do not cite
All three, route changed free-server New experiment, old notes die

Read the last row twice before you merge notes. A route change is a new system under test. Your old green file does not travel with it.

The objection

Some developers will call this pin check theater. They want speed, and the capacity is free. They say a pin check slows the point of the offer.

Speed is fine for a private scratch file. Citation is a different act with a higher price. The moment you publish a result, the pin is due.

You can still explore without running the script. You just cannot smuggle that exploration into a claim. Label the folder scratch and keep it out of the post.

A check you can run

The script below is an unexecuted proposal only. It does not call a model or a network. It only refuses a dirty record on disk.

#!/usr/bin/env python3
"""Refuse a run record that lacks three pins. Unexecuted proposal."""

import hashlib
import json
import sys
from pathlib import Path

REQUIRED = ("prompt_sha256", "schema_sha256", "image_digest", "host_role")


def sha256_file(path: Path) -> str:
    data = path.read_bytes()
    return hashlib.sha256(data).hexdigest()


def main() -> int:
    if len(sys.argv) != 4:
        print("usage: pincheck.py record.json prompt.txt tools.json")
        return 2
    record = json.loads(Path(sys.argv[1]).read_text(encoding="utf-8"))
    prompt_hash = sha256_file(Path(sys.argv[2]))
    schema_hash = sha256_file(Path(sys.argv[3]))
    missing = [key for key in REQUIRED if not record.get(key)]
    if missing:
        print("reject: missing " + ",".join(missing))
        return 1
    if record["host_role"] != "free-server":
        print("reject: host_role must be free-server")
        return 1
    if record["prompt_sha256"] != prompt_hash:
        print("reject: prompt hash drift")
        return 1
    if record["schema_sha256"] != schema_hash:
        print("reject: schema hash drift")
        return 1
    digest = str(record["image_digest"])
    if not digest.startswith("sha256:"):
        print("reject: image_digest must start with sha256:")
        return 1
    print("accept: pins match; still not a benchmark")
    return 0


if __name__ == "__main__":
    raise SystemExit(main())
Enter fullscreen mode Exit fullscreen mode

Run it against a fixture you wrote by hand. A failing exit is the result you should want first. Fix the record, then rerun the same command.

python3 pincheck.py run-record.json prompt.txt tools.json
echo "exit=$?"
Enter fullscreen mode Exit fullscreen mode

A zero exit means the record is internally consistent. It does not mean the agent is correct. It does not mean the free server stayed private.

Build the record without guessing

Hash the fixture files on your side first. Ask the free server for its image digest. Write the digest exactly as the server returned it.

# proposal commands; use the tool your OS actually ships
sha256sum prompt.txt tools.json 2>/dev/null || shasum -a 256 prompt.txt tools.json
Enter fullscreen mode Exit fullscreen mode

Do not clean it up to match a blog post. A small record looks like the JSON below. Replace every placeholder before you run the check.

Fixture files

Start with a tiny prompt and a tiny schema. Both files must stay byte-stable during the pair of runs. If you edit either file, you start over.

Read fixture/hello.txt and report the first line only.
Do not call any tool except read_fixture.
Enter fullscreen mode Exit fullscreen mode
{
  "tools": [
    {
      "name": "read_fixture",
      "description": "Read one file inside the fixture tree.",
      "arguments": {
        "path": "string"
      }
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode
{
  "prompt_sha256": "<hash from prompt.txt>",
  "schema_sha256": "<hash from tools.json>",
  "image_digest": "sha256:<digest the server returned>",
  "host_role": "free-server",
  "model_route": "<route you selected>",
  "note": "proposal fixture, not a live result"
}
Enter fullscreen mode Exit fullscreen mode

If the server will not tell you the digest, stop. You do not have a pin you can defend. You have a story, not a citable record.

Drift during one session

Ask for the digest before the run starts. Ask for the digest after the run ends. If the two strings differ, discard both transcripts.

Do not pick the closer digest and move on. A moving image means you have a moving system. Your diff is then comparing two different strangers.

# proposal: capture digests around one attempt
printf '%s\n' "$DIGEST_BEFORE" > digest-before.txt
printf '%s\n' "$DIGEST_AFTER" > digest-after.txt
cmp -s digest-before.txt digest-after.txt || echo "reject: image moved"
Enter fullscreen mode Exit fullscreen mode

Keep those two digest files beside the transcript. If cmp fails, the attempt never enters the table. Do not paste the rejected text into a chart.

How to use the free loop

Use free model access for repetition, not for glory. Run the same pinned prompt twice on purpose. Diff the tool calls, not the polished adjectives.

  1. Freeze prompt.txt and tools.json before you send either.
  2. Record the free server image digest in the note.
  3. Send the same files through the free model route.
  4. Save raw tool calls as run-a.json and run-b.json.
  5. Diff the call names and argument keys only.
  6. If the diff is nonempty, keep both runs as failures.
  7. Do not average them into a success rate.
# Unexecuted proposal. Adapt keys if your log shape differs.
import json


def names(path):
    with open(path, encoding="utf-8") as handle:
        data = json.load(handle)
    return [call.get("name") for call in data.get("calls", [])]


left = names("run-a.json")
right = names("run-b.json")
print("same_names", left == right)
print("a", left)
print("b", right)
Enter fullscreen mode Exit fullscreen mode

That snippet is also an unexecuted proposal for you. It assumes a log shape you already control. If your logs use another shape, adapt the keys.

Do not invent a pass to match the printout. A printed true is not a tested agent. Save the raw files even when the names match.

What free capacity must not touch

Give the free server a boring fixture repo. Use fake tokens, fake hostnames, and a read-only tree. Boring inputs keep the lesson about evidence, not drama.

  • Keep production credentials out of that environment entirely.
  • Keep customer data out of the prompt text.
  • Allow no network calls except the chosen model route.
  • Deny writes outside the temp fixture directory path.

A free server is shared capacity until proven otherwise. Treat it as hostile in the security sense. Courtesy from a vendor is not a private enclave.

Limits of this opinion

This check catches missing pins and hash drift. It does not catch a lying image digest. It does not score answer quality at all.

Two identical tool-name lists can still be wrong. You still need a human fixture with known outputs. That fixture is your job, not the model's job.

Free model access can disappear or change route. When the route changes, burn the comparison set. Start a new folder for the new route.

Do not mix old and new transcripts in one chart. This piece claims no latency number at all. This piece claims no token grant of any size.

This piece claims no win over any other stack. Treat every number you remember as unverified until sourced. A blog memory is not a primary source.

Who should not use this

Skip this flow if you need a certified environment. A free server option is the wrong home for secrets. It is the wrong home for a release gate.

Skip it if you cannot read the image digest. Skip it if your tool schema changes every prompt. Skip it if you wanted a demo, not a record.

Also skip it if you will ignore a nonzero exit. The script is useless once you bypass it. Discipline is the real gate, not the file.

What you should say afterward

Lead with the decision, not the tool brand. Say the run was pinned, or call it a scratch note. Do not let a free badge do the arguing.

If you mention the free model route, mention the date. If you mention the free server, mention the digest you saw. If either fact is missing, delete the paragraph.

One soft next step is enough for this piece. If you already have MonkeyCode free model access, use it here. Pair it with the free server option you were given.

Run the pin check on one fixture before you quote anything. Then stop writing product sentences into the lab note. More brand mentions will not make the hash tighter.

Top comments (0)