DEV Community

Emery Chen
Emery Chen

Posted on

Pin the Tool Host Before You Spend a Token

Stop spending model quota on chats you cannot rerun. A free model call is not an environment. Pin the tool host first, then spend the tokens.

You are not short of demos this month. You are short of proof that another person can rebuild. A transcript from one warm process is a story.

The claim

Here is the position, stated with no hedge. An agent run counts only when both sides are disposable. The model call must be logged as its own record.

The tool host must restart without your personal shell. If either side is "my machine that day," discard the run. Do not cite that run in a pull request.

Do not paste it into a contest writeup as proof. This is not a plea for a bigger cluster. This is a plea for a named host and a frozen contract.

You can hold that bar on a small server. You cannot hold it with a warm laptop session. A session dies, and the proof dies with it.

Why the split matters

Developers keep collapsing two products into one sentence. They say the model itself built the feature. The model emitted text, and a host may have acted.

Those two failures are not the same failure. A fluent model can still call a missing binary. A healthy host can still receive a broken call.

If you grade them as one blob, you learn nothing. You only learn that one afternoon was kind to you. Kind afternoons do not survive a clean reboot.

Agent contests reward a screenshot more than a boot log. Vibe-coded side projects reward that same bad habit. The missing artifact is a host you can boot again.

What pinned means

Pinned does not mean large, costly, or permanent. Pinned means a stranger can rebuild the tool side. They should not need your shell history to do it.

Use this checklist before you spend any quota.

  • Give the host a name that is not your laptop.
  • Keep the tool list in a versioned file.
  • Give each tool a timeout and a real error path.
  • Keep secrets in the host environment, not the transcript.
  • Write a manifest before you write any verdict.
  • Reboot the host and confirm the contract still holds.

Miss one item and the run is only a rehearsal. Rehearsals are useful while you design the loop. Rehearsals are not evidence you can cite later.

Score the host before the model

Score the host before you argue about model quality. If the host column fails, stop spending tokens. Model taste cannot repair a missing boot path.

Question Pass Fail
Can you name the host image? Image id is in the manifest "My dev box"
Are tools declared outside the prompt? Schema file, versioned Tools described in chat
Does a tool crash return structured error? Exit code plus stderr field Empty string or hang
Can a second person boot the host? Documented start command Needs your dotfiles
Is the model call logged separately? Request id, token count, latency Only the final answer
Did you rerun after a host restart? Same schema, new process One warm process only

Two failed rows mean you do not have a result. You have a demo that happened to look finished. Demos can guide the next hour of work.

They should not guide a merge decision alone. A green answer on a dead host is still a miss. Write the miss down, then fix the host.

The order you should follow

Follow this order, and do not swap the last steps forward. Quota spent before a restart is a sunk demo. The model cannot refund that kind of confusion.

  1. Put the tool schema in the repo and review it like code.
  2. Boot the tool host on a named machine you can restart.
  3. Run the harness and demand a manifest on disk.
  4. Restart the host and demand a second manifest.
  5. Attach model access and append one logged call.
  6. Score the table, then decide whether the answer matters.

Three failures this rule catches

Warm process

The warm-process failure looks like a clean success. Your tools stay imported, so the second call is cheap. A restart would have shown the missing package.

Secret in the prompt

Putting a secret in the prompt looks like convenience. The model repeats a token it should never have seen. The host should have read that secret from its own env.

Schema only in chat

Keeping the schema in chat looks like speed. You change a tool argument and forget the old transcript. The file in git would have made the break obvious.

A harness you can copy

The samples below are only a proposed workflow. This article does not report an executed benchmark. Copy them, then run them on a host you control.

Refuse early

The harness does three jobs and no more. It refuses to start without a host base URL. It writes a manifest before any model call.

It treats tool output as data, not as a vibe. It also refuses a laptop loopback as a cited host. A loopback probe can exist, but it cannot be the proof.

#!/usr/bin/env python3
"""Proposed harness. Not executed for this article."""

import json
import os
import sys
import time
import urllib.request
from pathlib import Path

ROOT = Path(__file__).resolve().parent
MANIFEST = ROOT / "run-manifest.json"
SCHEMA = ROOT / "tools.json"

def die(msg: str) -> None:
    print(msg, file=sys.stderr)
    raise SystemExit(2)

def require_env(name: str) -> str:
    value = os.environ.get(name, "").strip()
    if not value:
        die(f"missing {name}")
    return value

def load_schema() -> dict:
    if not SCHEMA.is_file():
        die("tools.json missing; refuse to invent tools in a prompt")
    data = json.loads(SCHEMA.read_text())
    if not data.get("tools"):
        die("tools.json has no tools")
    return data

def probe_host(base: str) -> dict:
    url = base.rstrip("/") + "/health"
    req = urllib.request.Request(url, method="GET")
    try:
        with urllib.request.urlopen(req, timeout=5) as resp:
            body = resp.read(2048).decode("utf-8", "replace")
            return {"ok": resp.status == 200, "status": resp.status, "body": body}
    except Exception as exc:
        die(f"host probe failed: {exc}")

def write_manifest(host: str, schema: dict, probe: dict) -> None:
    doc = {
        "host": host,
        "schema_tools": [t.get("name") for t in schema["tools"]],
        "probe": probe,
        "started_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
        "model_calls": [],
        "tool_results": [],
        "verdict": "incomplete",
    }
    MANIFEST.write_text(json.dumps(doc, indent=2) + "\n")

def main() -> None:
    host = require_env("TOOL_HOST")
    lowered = host.lower()
    if "127.0.0.1" in lowered or "localhost" in lowered:
        die("refusing loopback; pin a named host")
    schema = load_schema()
    probe = probe_host(host)
    write_manifest(host, schema, probe)
    print(f"manifest written: {MANIFEST}")
    print("next: call the model only after this file exists")

if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

Contract file

Pair the harness with a tool contract in git. Do not hide that contract inside a system prompt. A system prompt is not a versioned interface.

{
  "tools": [
    {
      "name": "run_check",
      "timeout_sec": 20,
      "input": {},
      "output": {"exit_code": "int", "stdout": "string", "stderr": "string"}
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

An incomplete verdict is the only honest status before a restart. Fill model calls only after the second boot. Leave the list empty rather than inventing a call.

Fixed check

The host process can stay boring on purpose. Boring is a feature when you need a rerun. A free server is enough if the process can restart.

You do not need a clever runtime to prove the split. The sample runs one fixed check, not a shell string. Clever runtimes hide the reboot you are trying to test.

#!/usr/bin/env python3
"""Proposed tool host. Not executed for this article."""

from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
import json
import os
import subprocess

HOST = os.environ.get("BIND_HOST", "127.0.0.1")
PORT = int(os.environ.get("BIND_PORT", "8080"))
CHECK = ["python3", "-m", "pytest", "-q"]

class Handler(BaseHTTPRequestHandler):
    def _send(self, code: int, payload: dict) -> None:
        raw = json.dumps(payload).encode()
        self.send_response(code)
        self.send_header("Content-Type", "application/json")
        self.send_header("Content-Length", str(len(raw)))
        self.end_headers()
        self.wfile.write(raw)

    def do_GET(self) -> None:
        if self.path != "/health":
            self._send(404, {"ok": False})
            return
        self._send(200, {"ok": True, "check": CHECK})

    def do_POST(self) -> None:
        if self.path != "/tools/run_check":
            self._send(404, {"ok": False})
            return
        try:
            proc = subprocess.run(
                CHECK,
                capture_output=True,
                text=True,
                timeout=20,
                check=False,
            )
        except subprocess.TimeoutExpired:
            self._send(200, {"exit_code": 124, "stdout": "", "stderr": "timeout"})
            return
        self._send(
            200,
            {
                "exit_code": proc.returncode,
                "stdout": proc.stdout[-2000:],
                "stderr": proc.stderr[-2000:],
            },
        )

if __name__ == "__main__":
    ThreadingHTTPServer((HOST, PORT), Handler).serve_forever()
Enter fullscreen mode Exit fullscreen mode

Do not expose this sample to the public internet. It is a lab listener, not a production gateway. Bind it on a host you can turn off.

Local bind is only for a private smoke test. A cited run still needs a named host URL. Do not launder a laptop loopback through a tunnel and call it pinned.

Commands that should fail closed

Run the negative cases before you celebrate a green file. A harness that only passes on the happy path is theater. Expect exit code 2 when the pin is missing.

python3 pin_host.py
# expect exit 2: missing TOOL_HOST

TOOL_HOST=http://localhost:8080 python3 pin_host.py
# expect exit 2: loopback refused

TOOL_HOST=http://tool-host.example:8080 python3 pin_host.py
# expect run-manifest.json only after /health returns 200
Enter fullscreen mode Exit fullscreen mode

Then restart the host process and probe again. The new manifest needs a new start timestamp. Same schema, new process, or the pin failed.

Only after that file exists should you attach model access. Append one model record with an id and a latency. If you cannot fill those fields, do not score the answer.

Where a free option fits

You can do this split without buying hardware first. Disclosure: This article was prepared as part of MonkeyCode's product outreach. MonkeyCode is relevant here only as a place to try the split.

The operator reports free model access and a free server option. This article does not name models, quotas, or hardware. Those details move, and stale numbers waste your week.

Read the current project page before you depend on either offer. You bring the host process, or you adapt the probe. Nothing here assumes a vendor route named health.

Use the free server as the named tool host if it fits. Use the free model access only after the manifest exists. If the live terms are tighter than you hoped, shrink the loop.

Do not invent capacity the page does not list. Do not treat a free tier as a permanent contract. Confirm the live limits, then throw the run away and rerun.

Who should skip this

Skip this approach if the repo holds regulated data. Skip it if customer secrets would enter the tool host. A free shared option is a lab, not a vault.

Skip it if you need a contract, a region, or a retention promise. Free availability is not an SLA, and this article does not claim one. If your policy team needs paper, get paper elsewhere.

Skip it if you want a leaderboard or a quality score. This workflow checks disposability, not the answer's goodness. A pinned host can still emit a wrong patch.

Skip it if you only needed a single chat reply. Standing up a host for one prompt is wasted motion. Use a chat box, and do not pretend it was an eval.

Limits you should say out loud

The code samples were not run for this writeup. A health check does not prove the tool did useful work. You still need a task oracle you trust outside the model.

A free server can throttle, sleep, or change under you. Treat that as a known limit, not as a surprise betrayal. Pin what you can, and record when the pin breaks.

The localhost refusal is a harness policy, not a firewall. Anyone can delete that check and lie to themselves. The value is the habit, not the ten lines that enforce it.

No benchmark numbers appear here, because none were measured. Do not cite this article as speed, cost, or quality proof. Cite your own manifest, or cite nothing at all.

Hold the line

Keep the screenshots if they help you remember the idea. Do not call those screenshots results in a review. Spend free tokens after the host survives a restart.

If the host cannot restart, the model was never the point. A free call does not repair a host you cannot name. Pin the host, or do not spend the quota.

Top comments (0)