DEV Community

Taylor Wang
Taylor Wang

Posted on

48-Hour Field Notes: I Parked the Prompt Until the Runtime Card Existed

I had a tiny Python job that passed in my editor and then died on a clean remote interpreter. My hands went straight back to the chat box, because the model had drafted the snippet and I wanted a new patch. Have I ever handed that model the interpreter it was about to meet, or only the traceback from my own laptop? That missing card is the hole I would spend the next forty-eight hours closing, before I trust another rewrite.

The blame I would park for a day

I would not open with a model swap, and I would not open with a guessed package install. I would park the prompt until two JSON cards existed, one from the laptop and one from the job that failed. Would a longer prompt fix a missing tomllib import on an older interpreter, or would it just narrate the same crash? That loop wastes a full afternoon when the failing interpreter never appears in the prompt at all.

I would also write down the question I am actually asking, because "make it work" hides two different failures. One failure is a snippet that is simply wrong on every interpreter I can still find nearby. The other failure is a snippet that is legal only on the laptop that happened to draft it. Which of those two failures am I actually holding, if I have not printed both sides yet?

The card I would run unchanged

I would keep one file, run it with the same command shape on each host, and refuse to edit it between those runs. The point of the file is not a pretty dashboard that I would paste into a status meeting. The point is a stable key set so a textual diff stays honest across the two machines. If a key appears on only one side, I would treat that as a probe bug, not as a runtime fact.

Keys I would refuse to freestyle

I would record the version triple, the implementation name, and the machine string before I record anything clever. I would record isolated mode, the user-site flag, and the safe-path flag, because those flags change imports without changing my source file. I would probe tomllib and zoneinfo by name, then ask the default SSL context how many CA certificates it loaded. I would allowlist a few environment names and leave every secret key out of the dictionary on purpose.

"""Runtime capability card.

Unexecuted proposal: run this file unchanged on each interpreter.
Do not treat any printed JSON as a result already measured.
"""
from __future__ import annotations

import importlib.util
import json
import os
import platform
import ssl
import sys
import tempfile

ALLOW = (
    "TZ",
    "LANG",
    "LC_ALL",
    "SSL_CERT_FILE",
    "SSL_CERT_DIR",
    "PYTHONNOUSERSITE",
    "PYTHONSAFEPATH",
    "PYTHONPATH",
)

def probe(name: str) -> str:
    spec = importlib.util.find_spec(name)
    if spec is None:
        return "missing"
    return "present"

def ssl_default() -> str:
    try:
        ctx = ssl.create_default_context()
    except Exception as exc:
        return type(exc).__name__
    return f"certs:{len(ctx.get_ca_certs())}"

def card() -> dict:
    names = getattr(sys, "stdlib_module_names", ())
    return {
        "version": list(sys.version_info[:3]),
        "implementation": platform.python_implementation(),
        "machine": platform.machine(),
        "isolated": int(getattr(sys.flags, "isolated", 0)),
        "no_user_site": int(getattr(sys.flags, "no_user_site", 0)),
        "safe_path": int(getattr(sys.flags, "safe_path", 0)),
        "prefix": sys.prefix,
        "temp": tempfile.gettempdir(),
        "tomllib": ("tomllib" in names) if names else probe("tomllib"),
        "zoneinfo": ("zoneinfo" in names) if names else probe("zoneinfo"),
        "ssl_default": ssl_default(),
        "env_allow": {key: os.environ.get(key) for key in ALLOW},
    }

if __name__ == "__main__":
    print(json.dumps(card(), indent=2, sort_keys=True))
Enter fullscreen mode Exit fullscreen mode

Commands I would keep boring

I would save the file as runtime_card.py and redirect stdout into laptop.json from a directory I chose. I would copy the same file to the remote job and run it there without a quiet local edit in between. I would take a third snapshot under isolated mode when I suspect a user site is leaking imports. I would then run a unified diff on the JSON files and read the hunks before I reopen any chat.

python runtime_card.py > laptop.json
python -I runtime_card.py > isolated.json
diff -u laptop.json remote.json
Enter fullscreen mode Exit fullscreen mode

How I would read a mismatch

I would not ask a model to interpret a raw diff until I had applied a short table myself. A version gap and a missing CA bundle ask for different fixes, and mixing them produces a confident wrong patch. These rows are reading rules for the next run, not timings from a host I am pretending to have measured. I would write the matching row beside each hunk so the later prompt cannot wander.

Diff you see What I would assume What I would not do
version crosses 3.9, 3.10, or 3.11 A stdlib name may simply be missing Install a random backport first
isolated or no_user_site flips A user site or sitecustomize is involved Rewrite the function body first
ssl_default is an exception or certs:0 CA loading differs across the two hosts Rewrite the HTTP client first
temp or prefix differs Those paths are host-local Treat the path as a logic bug
Allowlisted env values differ The job may have missed an export Let a model invent the value
  1. I would run the card locally, then on the remote job, then under isolated mode, before I open a chat.
  2. I would redact allowlisted values and only then paste the two JSON objects into a prompt.
  3. I would apply the table myself and reject any patch that skips a printed version boundary.
  4. I would re-run the same unchanged card after the patch, instead of trusting a green sentence from the model.

Where a second interpreter fits

Disclosure: This article was prepared as part of MonkeyCode's product outreach.

I would try MonkeyCode's free server option as that second interpreter only when it can execute this same plain Python file. I would use free model access only after both cards exist, and only to draft a mismatch list I can check against the table. I am not naming a model, a quota, a machine size, or a date when that free option might change. If those details matter to your schedule, read the current product notes before you plan a batch around them.

Would I paste both cards into the model and ask it to make the tests pass on the first turn? No, because that request invites a silent rewrite that hides the version boundary I still need to see. I would ask for keys whose values differ, then for one change that still runs on the older version triple. I would reject any answer that invents an install command when the table already explains the gap.

You can run this same loop with a local virtualenv and any chat window you already trust today. The notes stay useful if you never open the product, which is the bar I want for a field note. If you already have access, try the card on the free server and redact paths before any model sees them. Only then would I spend one free-model turn on the diff, and I would keep that turn singular on purpose.

What I expect to break

I would budget the early hours for the card itself, because a sloppy probe can lie with a straight face. get_ca_certs() reports what the default context loaded, not every certificate file sitting on disk, so an empty count is only a clue. Printing PYTHONPATH can expose local directories I should not paste into a shared chat, so I would redact values first. A future allowlist key might hold a token rather than a directory, and I would rather drop the key than leak it.

I would also expect the textual diff to look louder than the real bug I am chasing. Prefix strings change across operating systems even when the language contract is identical, and I would note that noise without chasing it. A model that sees the noise may propose a path rewrite and miss the version boundary sitting a few lines above. Have you ever shipped that kind of fix and watched the same import fail again on the remote job?

I would not write a pass rate, a timing figure, or a saved-hour claim into this note. Those numbers do not exist until the two cards are real files, and a made-up figure would make the comparison useless. If the remote side cannot execute plain Python, I would stop and pick a different probe instead of forcing this file. Forcing the file onto a non-Python job would only create a new failure that looks like evidence.

What I would repeat

I would repeat the unchanged file, the three saved JSON snapshots, and the rule that version gaps outrank style complaints. I would repeat the redaction step every time the allowlist grows, even when the new name looks harmless on my laptop. I would drop any urge to let the model choose the Python version for me, since the card already printed the triple I must support. I would also repeat a local self-check that does not depend on any hosted service at all.

On a machine I control, I would run a tomllib import under Python 3.11 or newer and expect that import to succeed. I would run the same check under an older interpreter and expect ModuleNotFoundError, which is a language fact I can verify myself. I would do the same for zoneinfo against a Python 3.8 interpreter if I still have one nearby. If that self-check fails, the card is wrong and I would fix the probe before I blame either host.

Who should skip this approach

This card is a debugging aid for a script that already runs, or fails, under a normal CPython job you can invoke. You should skip it if you need a signed software bill of materials, a locked production image, or a formal security review. You should also skip it if the remote side is not a Python process, because a capability card cannot explain a missing runtime. A container image digest, a lockfile, and a real staging deploy still matter more than this JSON when you are shipping.

Do not use this workflow if environment values might contain secrets, and do not paste unredacted JSON into any model. Skip it if you need a guaranteed model identity, reserved capacity, or a support commitment, because a free-access note does not establish those things. If your failure is a product incident rather than an interpreter mismatch, file that incident instead of diffing sys.prefix. Why would a runtime card answer an outage that never reaches the Python process you thought you were debugging?

The note I would keep

I would end the forty-eight hours with two JSON files and a written reason for the next edit, not with a thicker prompt. If the cards match and the job still fails, then I would finally blame the snippet, with the remote traceback kept beside the diff. Until those files exist, the model is guessing about a machine it has not seen, and I would not spend another hour on that guess. The repeatable part is the card, the redaction, and the refusal to rewrite before the mismatch has a name.

Top comments (0)