DEV Community

Taylor Lin
Taylor Lin

Posted on

Split the Model Call From the Host Before an Agent Job Starts

A review note once collapsed two resources into one phrase: the agent had finished on "the free setup." The diff renamed a column in a migration. The log did not name the machine, the toolchain, or whether a credential had been in the environment. That note is a composite illustration, not a report from a named team. It is still a placement bug worth fixing before the next job starts.

A model call can be inexpensive and still be the wrong place to execute. A free server can be convenient and still be the wrong place to keep secrets. This walkthrough separates those choices with a glossary, a numbered procedure, a small classifier, and one concrete job at every outcome. The classifier is a proposal. It has not been timed, load-tested, or run against a production fleet.

1. Name the resources before you spend either

Treat "free" as an availability flag, not as permission to run every job.

A model call consumes inference. It can draft a patch, explain a traceback, or propose a command. It does not, by itself, prove that a command ran anywhere.

An execution host runs processes. It has a filesystem and, usually, a network. Output from that host is evidence only about that host. If the review note never names the host, the evidence is incomplete even when the exit code is zero.

Disclosure: This article was prepared as part of MonkeyCode's product outreach. MonkeyCode is relevant only as an operator-supplied option for free model access and a free server. This article does not state model names, token quotas, hardware, duration, or permanence. If those terms are not in the product docs you can open today, leave them out of the job record.

2. Glossary

Use one meaning in the job record, the review note, and the log header.

Term Meaning Not the same as
Model call An inference request that returns text or a patch A process exit on a machine you can name
Execution host The machine or container that runs commands The service that served the model
Scratch workspace A disposable directory that contains fixtures only A checkout, a home directory, or a CI secret store
Private host A machine already allowed to hold credentials or internal routes A shared free server
Evidence host A pinned environment you are willing to cite in a release or merge note Any host that printed ok
Placement The leaf chosen before the job starts A label added after the log looks green

If a field is missing, placement is incomplete. Do not infer it from a friendly summary line.

3. Fill the record with inspections

The booleans below are claims about a workspace. Set them from inspection, not from the task title.

  1. List the files the job will read or write. A draft that only quotes a traceback needs no execution. A command that applies a migration does.
  2. Search the candidate tree for credential material before any copy step. A hit means touches_secrets=true until you prove the hit is a fixture with no live value.
  3. Mark writes_outside_scratch=true if the command targets the repo, a shared volume, a package registry, or a database.
  4. Mark reaches_private_network=true if the command resolves an internal hostname, a VPC address, or a webhook you would not put on a public laptop.
  5. Mark claimed_as_release_evidence=true if a person will cite the result in a merge decision, a changelog, or a security note. The same command can change leaves when the claim changes.
# Inspection sketch for a workspace you already trust. Not a vendor command.
git status --short
rg -n -i "api_key|secret|token|password|database_url" -g '!*.lock' || true
# If rg is unavailable, use the search your repo already documents.
Enter fullscreen mode Exit fullscreen mode

A clean search is not a proof that no secret exists. It is a minimum check before you even consider a disposable host. If the search fails closed because the tool is missing, do not treat the workspace as clean.

4. Decide in order

Later answers do not soften an earlier answer that demands a stricter host.

  1. Will someone cite this result as release evidence, merge evidence, or a security claim? If yes, the leaf is evidence_host. A free server is not that host unless a separate, written parity check already exists. This article does not assume one.
  2. If the result will not be cited that way, does the job need execution at all? If it needs only an explanation or a draft, the leaf is explain_only. Use a model call. Do not start a server.
  3. If it needs execution, does it touch secrets, write outside a scratch directory, or reach a private network? If any flag is true, the leaf is private_host. Keep the free server out, even if it has spare capacity.
  4. Otherwise the leaf is scratch_server. A free server option is eligible only here: fixtures in, disposable workspace, no credentials, and no claim that the run certifies production.

Job size is not a branch. A ten-line script that reads DATABASE_URL is still private_host.

5. Encode the procedure

The Python below is a decision aid, not a client and not a benchmark. It encodes the order in section 4. Run it on a record you wrote.

from dataclasses import dataclass

@dataclass(frozen=True)
class Job:
    needs_execution: bool
    touches_secrets: bool
    writes_outside_scratch: bool
    reaches_private_network: bool
    claimed_as_release_evidence: bool

def place(job: Job) -> str:
    if job.claimed_as_release_evidence:
        return "evidence_host"
    if not job.needs_execution:
        return "explain_only"
    risky = (
        job.touches_secrets
        or job.writes_outside_scratch
        or job.reaches_private_network
    )
    if risky:
        return "private_host"
    return "scratch_server"
Enter fullscreen mode Exit fullscreen mode

Hand-written expectations, not measured results:

explain a traceback, no commands                 -> explain_only
pytest on fixtures under /tmp/job-1842          -> scratch_server
alembic upgrade head with a live database URL   -> private_host
same pytest, cited as merge evidence            -> evidence_host
Enter fullscreen mode Exit fullscreen mode

If place and your review note disagree, fix the note or the booleans. Do not "correct" the function so a convenient host wins.

6. One job at each leaf

These four cases are illustrations. They are not logs from a live provider.

explain_only. The ticket asks why a traceback mentions a missing id column. No files change. No sockets open. Placement stops at the model call. Paste the explanation into the ticket. There is no server bill to attach, because no server was required. If the explanation proposes a migration, that proposal is a new job and must be classified again.

scratch_server. The record says needs_execution=true and every sensitive flag is false. The command is pytest tests/test_parse.py against fixtures. A disposable workspace is eligible. A free server may be that workspace only if the docs you checked still offer it. Copy fixtures. Do not copy .env. When the process exits, delete the workspace. Keep the junit file as a note about that disposable host, not as a release citation.

# Illustration only. Confirm place(job) == "scratch_server" first.
test ! -f .env
mkdir -p /tmp/job-1842
cp -R tests/fixtures /tmp/job-1842/
# Start the disposable workspace with the tool your team already uses.
# Delete /tmp/job-1842 after you save the junit note.
Enter fullscreen mode Exit fullscreen mode

private_host. The command is alembic upgrade head and the environment contains a real database URL. The classifier returns private_host even if a free server is idle. Capacity is not trust. Run the command on a machine that is already allowed to hold that URL. If no such machine is available, the honest result is "not run." Redacting the log afterward does not move the leaf.

evidence_host. The command is the same fixture test, but the reviewer will cite it in the merge note. The leaf flips because the claim changed, not because the test changed. Pin the image, record the commit, and run on the host your team already uses for citations. A green log from a disposable server can remain a sketch. It is not the citation.

7. Limits, and who should skip this

The classifier does not choose a model, a region, or a price. It does not prove that a free tier is currently up. It does not sanitize prompts, and it will lie if the booleans are wishes. touches_secrets=false is valid only after an inspection, not because the title said "demo."

Skip this approach for production customer data, live credentials, or any artifact you cannot delete. Skip it when policy forbids third-party execution. Skip it when you need a number this article does not have: quota, retention, hardware, or how long a free server remains free. Those facts belong to current primary docs, not to a placement sketch.

Write the five booleans, run place, and put the leaf name in the review note before an agent starts. Use a free model call or a free server only on the leaves that allow them. Leave the free server unused when the leaf is private_host or evidence_host.

Top comments (0)