DEV Community

Taylor Wang
Taylor Wang

Posted on

48-Hour Field Notes: I Would Inventory Probe Tools Before I Blame the Endpoint

I spent the first evening blaming a health endpoint that still returned 200 when I opened it in a browser. The generated probe kept dying on a missing command, so I nearly filed the whole thing as an upstream outage. Have you ever watched a green page and a red job argue about the exact same URL? I had, and the real mismatch lived in the tool list rather than in the service itself.

The recent developer chatter keeps returning to generated fixes and to the blame that follows a red job. I do not need a trending title to recognize that habit, because I just practiced it on a probe. A model can suggest a clean command and still assume a laptop that the remote box does not match. The useful response is a diff of tools, not a louder theory about the health endpoint itself.

What I thought had failed

I assumed the route had drifted, because that is the story I reach for whenever a check goes red. The page loaded, the status looked fine, and the job log only showed a short command failure. Why would I inspect the remote PATH when the browser had already proved the host was alive? I would now, because a browser tab and a remote shell are not the same runtime at all.

The second wrong guess was a client timeout that I had copied from a laptop note. I had reused a probe flag I remembered, and the remote shell rejected that flag before any packet left. A rejected flag is not a slow service, even when the log line looks dramatic and urgent. I wasted an hour raising the timeout number while the binary I needed was simply absent.

What I tried in the first twelve hours

I stopped editing the probe and started writing down the box in a file I could diff later. I wanted a record I could compare after a reboot, a package install, or a fresh login elsewhere. Would you trust a generated fix you cannot replay after the next clean login on that box? I would not, so the notes became a small script instead of another paragraph in a chat window.

The inventory is deliberately boring, and it refuses to call the health URL until a client exists. It writes a marker under the home directory so I can see whether that home survives the next login. It prints the shell, the user, and the resolved binaries without installing a single extra package. The script below is the artifact I would commit beside the job, not a benchmark and not a product claim.

A snapshot script, not an installer

#!/bin/sh
# probe-inventory.sh - snapshot only, no installs
set -eu
out="${TMPDIR:-/tmp}/probe-inventory.txt"
{
  date -u +when=%Y-%m-%dT%H:%M:%SZ 2>/dev/null || date
  id -un 2>/dev/null | sed 's/^/user=/' || echo user=unknown
  printf 'shell=%s\n' "${SHELL:-unknown}"
  printf 'home=%s\n' "${HOME:-unset}"
  printf 'path=%s\n' "$PATH"
  for cmd in curl wget python3 python sh bash; do
    if command -v "$cmd" >/dev/null 2>&1; then
      printf 'have_%s=%s\n' "$cmd" "$(command -v "$cmd")"
    else
      printf 'have_%s=missing\n' "$cmd"
    fi
  done
  marker="${HOME:-/tmp}/.probe-inventory-marker"
  if [ -n "${HOME:-}" ] && touch "$marker" 2>/dev/null; then
    echo home_writable=yes
  else
    echo home_writable=no
  fi
} > "$out"
printf 'wrote %s\n' "$out"
Enter fullscreen mode Exit fullscreen mode

I run that same file on the laptop and on the remote shell, then I diff the two snapshots. The point is not a pretty report I can paste into a status channel. The point is a disagreement I can see before I accept a generated command.

sh probe-inventory.sh
cp /tmp/probe-inventory.txt /tmp/laptop-probe-inventory.txt
# repeat on the remote shell, copy that snapshot back, then:
diff -u /tmp/laptop-probe-inventory.txt /tmp/remote-probe-inventory.txt
Enter fullscreen mode Exit fullscreen mode

What broke when I trusted the generated probe

The first generated probe called curl with long flags for failure, silence, errors, and a five-second limit. That line is reasonable on a laptop where curl is installed and those long options are accepted. On the remote shell, the inventory printed have_curl=missing, so the job died before it could connect. Have you ever raised a timeout for a program that was not even present on the remote disk?

A stdlib fallback I would keep labeled

The fallback I kept uses Python's standard library, because that inventory showed python3 and not curl. This fallback is a proposal I would store beside the notes, not a claim that every image ships Python. If python3 is missing too, the honest result is no client, not another flag invented from memory.

#!/usr/bin/env python3
# Minimal probe. Stdlib only. Unexecuted until inventory shows python3.
import sys
import urllib.request

def main(url):
    req = urllib.request.Request(url, method='GET')
    try:
        with urllib.request.urlopen(req, timeout=5) as resp:
            print('status=%s' % resp.status)
            return 0 if 200 <= resp.status < 300 else 1
    except Exception as exc:
        print('probe_error=%s: %s' % (type(exc).__name__, exc), file=sys.stderr)
        return 2

if __name__ == '__main__':
    if len(sys.argv) != 2:
        print('usage: probe_url.py URL', file=sys.stderr)
        sys.exit(2)
    sys.exit(main(sys.argv[1]))
Enter fullscreen mode Exit fullscreen mode

The second break was persistence, which I noticed only after I celebrated a green user-site install. A later login did not have that tree, and the marker file I expected was gone as well. A disposable shell is allowed to forget your home, your packages, and your shell history. Why did I treat one successful install as a contract the next session had to honor?

The third break was the prompt, not the network, and it was entirely my own copying habit. I pasted a log that still contained a query string with a token, then asked for a shorter probe. A model can repeat whatever you paste, and a shared shell can keep history you never meant to leave. I would redact the URL before either step, even when the failure looks purely mechanical and safe.

# redact query secrets before any note leaves the box
sed -E 's/([?&](token|key|sig)=)[^& ]+/\1REDACTED/g' job.log > job.redacted.log
Enter fullscreen mode Exit fullscreen mode

What I would repeat on hour forty-eight

I would not start from the generated command, even when the wording looks cleaner than mine. I would start from the inventory, then allow only the clients that both sides actually have. A small decision table keeps me from negotiating with a red log at the end of a long day. Would I skip that table just because I still remember which tools my laptop had yesterday?

The only probes I would allow

Inventory result Probe I would allow What I would not do
have_curl is a real path curl fail, silent, max-time 5 Assume wget flags match
curl missing, python3 present stdlib urlopen with timeout 5 pip install to match the laptop
both clients missing stop and record no client Invent a one-liner from memory
home_writable=no write notes under /tmp only Hide state in a user site
marker gone after relogin treat installs as session-scoped Call one green run permanent

The repeatable loop is short enough to run while I am tired, and I want it to stay boring. I would rather repeat four dull checks than invent a fifth client flag from a half-read man page. None of these steps install software, change a service, or require a named model to be useful. If a step fails, the failure is the note, and I stop instead of papering over it.

  1. Run the inventory locally and on the remote shell, then save both snapshots with a timestamp.
  2. Diff the two files and circle every missing line before you open a chat window.
  3. Pick a probe from the table, and refuse any suggestion that calls a binary you do not have.
  4. Redact tokens from the log, then ask for a review of that chosen probe only.
  5. Reopen the session and check the marker before you describe any install as permanent.

Where a free model and a free server fit

Disclosure: This article was prepared as part of MonkeyCode's product outreach. I kept the inventory as the real method, and the product entered only after that snapshot already existed. The only availability I am willing to use here is free model access plus a free server option. I do not have a primary source for quotas, hardware, duration, or permanence, so those figures stay out.

The useful split is simple, and it does not require me to treat a laptop result as universal. I would run the inventory on the free server, then paste only the redacted snapshot into the model. I would ask it to reject any probe that calls a binary the snapshot marks as missing. I would run the accepted probe on that same server, then reopen the session and check the marker.

If the free server forgets the home directory, the fix belongs in the job script or the image. A one-shot install is a rehearsal note, not a permanent remedy I would hand to the next on-call. Would I let a model choose the server image just because a free option happened to be available? No, a free server is a place to rehearse the diff, not proof that production matches it.

If your team already pins a base image with a required client, keep that contract and use the inventory as a regression check. A free option does not replace that pin, and a green rehearsal does not promote itself into production. I would still want a human to read the diff before anyone restarts a shared health check. If you already have that free model access and free server, run this inventory there before you accept a generated probe.

Who should skip this, and what it will not catch

This approach is a bad fit when your incident process forbids ad hoc shells on the failing host. It is also a bad fit when the only legal client is a signed binary that PATH will not reveal. People who would paste live credentials into a prompt should not use the model step at all. Teams that already diff a locked image bill of materials do not need this as their source of truth.

The inventory will not catch an application bug that returns 200 with an empty or stale body. It will not catch a TLS intercept, a bad clock, or a signed URL that expires between two runs. It will not tell you whether a free server stays up next week, because I never measured that. A missing measurement is not a zero, and I will not decorate the gap with a borrowed number.

A snapshot can go stale the moment someone installs a package or rotates the underlying image. I would date the file and refuse to reuse yesterday's have_curl line without running the script again. If the remote shell is not a POSIX sh, the script can fail before it reports a single tool. That failure is still a useful fact, and I would record it instead of blaming the health endpoint.

I would repeat the inventory, the diff, and the marker check whenever the next probe goes red. I would not repeat the hour I spent raising a timeout for a command that was never installed. The endpoint can still be wrong, but I want that conclusion to survive a tool list. Have I earned the right to blame the service before those three checks are on disk?

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •
You need to verify your account.
Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to