DEV Community

niuniu
niuniu

Posted on

The Weekly Digest Failed Open

You are sitting in Monday standup when a teammate asks why the weekly digest looks strangely blank. The bot posted on time, the heading is intact, and the body claims there were no wins worth keeping. You know the page is wrong, because Thursday's pull requests are still sitting in the repository history. The channel goes quiet in the particular way that means people are rewriting the week to match the page.

You open the job log before you open the calendar, because a green check is not the same thing as a true report. The workflow finished cleanly, which is why nobody paged you and why the blank week survived until standup. A closer read shows the model call failed on billing, the catch block swallowed it, and the renderer shipped your polite fallback. The process that was meant to stay reachable had been a terminal on a laptop, and that laptop slept at 18:40.

The timeline is short enough to retell without slides, and that shortness is part of why it was missed. At 17:02 the scheduled job started through the laptop tunnel and called the paid model endpoint you had been using all week. At 17:02 that endpoint rejected the key, your client treated the refusal as an empty result, and the renderer wrote the quiet-week sentence. At 18:40 the lid closed, the tunnel dropped, and Monday inherited a false record instead of a missed post.

Three contributing factors turned an outage into something the team could mistake for a product truth. First, language generation and host reachability were both required, yet the code treated each of them as optional. Second, the fallback copy was written to be polite, so an outage and a genuinely quiet week produced the same sentence. Third, the host was a machine with a sleep policy, which is a fine sketchpad and a bad place to keep a promise.

People did what you would expect with a clean heading and an empty body on a Monday morning. The standup treated the blank week as a social fact, and the conversation shifted toward explaining a slow week that had not happened. That is the expensive part of a failed-open report, because the outage edits human memory before anyone reads the log. A red job would have been cheaper, since a red job starts an argument about the bot instead of an argument about the team.

Raising the quota would only have moved the same failure from this Monday to a later Monday. A longer sleep timer fails the same way, because a commute, a dead battery, or a system update can still drop the tunnel. You need the channel to say which dependency failed, and you need the scheduler to live somewhere that does not share your sleep policy. Until both of those are true, a larger allowance only buys a later surprise with the same polite sentence.

You can picture the setup as a fire alarm wired through a lamp that only works while you are still in the room. When the lamp is dark, you cannot tell whether there was no fire or whether the watchman went home. The first repair is therefore to make failure loud before you try to make either dependency cheap. You want separate host and model probes, and you want those two results to stay distinct in the channel.

The sketch below is an unexecuted example, and you should adapt it only to endpoints you actually control. It does not call a vendor SDK, and it does not assume a quota, a model name, a region, or a hardware shape. Replace the two URLs, keep the exit codes stable, and treat any timeout as a down dependency rather than as an empty week. If you cannot explain an exit code in one sentence, do not let that code paint the job green.

# Unexecuted example: split host health from model health.
# Replace both URLs with endpoints you control. Never echo the token.
set -u
HOST_URL="${HOST_URL:-http://127.0.0.1:8787/health}"
MODEL_URL="${MODEL_URL:-https://example.invalid/v1/models}"

host_code=$(curl -s -o /dev/null -w "%{http_code}" --max-time 5 "$HOST_URL" || echo "000")
model_code=$(curl -s -o /dev/null -w "%{http_code}" --max-time 8 \
  -H "Authorization: Bearer ${MODEL_TOKEN:-}" "$MODEL_URL" || echo "000")

if [ "$host_code" != "200" ]; then
  echo "HOST_DOWN code=$host_code"
  exit 2
fi
if [ "$model_code" != "200" ]; then
  echo "MODEL_DOWN code=$model_code"
  exit 3
fi
echo "BOTH_UP"
Enter fullscreen mode Exit fullscreen mode

You then change the renderer so a non-success never becomes the sentence you reserved for an empty week. A down host should say the digest was skipped because the host was down, and it should say nothing else about the week. A refused model should say the digest was skipped because the model refused, using a different sentence so on-call can route the page. Only a successful call that returns no items should use the quiet-week sentence, and that sentence should name the filter you applied.

# Unexecuted example: renderer plus a local self-check. Not run against a live vendor.
def render_digest(status: str, items: list[str]) -> str:
    if status == "host_down":
        return "Digest skipped: host down. Do not treat this week as empty."
    if status == "model_refused":
        return "Digest skipped: model refused. Do not treat this week as empty."
    if status == "ok" and not items:
        return "Digest completed: no items passed the filter."
    if status == "ok":
        return "Digest completed:\n" + "\n".join(f"- {item}" for item in items)
    return "Digest skipped: unknown status. Do not treat this week as empty."

def self_check() -> None:
    for status in ("host_down", "model_refused", "unknown"):
        text = render_digest(status, ["a merged fix"])
        if "Do not treat this week as empty." not in text:
            raise SystemExit(f"WORDING_COLLAPSED status={status}")
        if "no items passed the filter" in text:
            raise SystemExit(f"QUIET_WEEK_LEAK status={status}")
    quiet = render_digest("ok", [])
    if "no items passed the filter" not in quiet:
        raise SystemExit("QUIET_WEEK_MISSING")
    print("WORDING_OK")

if __name__ == "__main__":
    self_check()
Enter fullscreen mode Exit fullscreen mode

A small test plan keeps that wording split from rotting the next time someone refactors the renderer. Stop the host process, run the probe, and expect exit 2 plus the host-down sentence in a dry-run channel. Point the model probe at a bad token and expect exit 3, with the job failing closed instead of exiting zero. Save the renderer as render_digest.py, run python3 render_digest.py, and accept a wording change only when it prints WORDING_OK.

After the local checks pass, close the laptop lid and wait for the sleep policy to kill the tunnel. The next scheduled run should fail on the host probe, and it should not post the quiet-week sentence. If the scheduler still lives on that laptop, you have reproduced the incident rather than closed it. Move the scheduler only after those four observations are written into the runbook where the next on-call will see them.

The durable fix is to stop parking a scheduled promise on a machine that is allowed to sleep. Two operator-supplied MonkeyCode options fit this gap: free model access, and a free server for the host. Disclosure: This article was prepared as part of MonkeyCode's product outreach. This draft does not treat a token quota, duration, model name, or hardware shape as a verified fact.

You can still send the rare, high-stakes summary to a paid frontier model after the probes pass. The lesson from the blank Monday is narrower than a shopping comparison, because reachability and generation failed separately. A polite fallback becomes a second incident when it looks like a result and lands in a channel people trust. If a green check can still hide either failure, the runbook is not finished, no matter which host you pick.

This approach is a poor fit when the digest can see customer data, secrets, or text you would not paste into a third-party log. It is also a poor fit when you need a contractual uptime promise, a fixed quota, or a named model that must not change under you. Free access can move, rate-limit, or disappear, and a weekend bot should degrade into an explicit skip rather than a fabricated quiet week. If you cannot explain the failure in one sentence in the channel, you are not ready to automate the summary.

Before you automate another recap, run the four checks against a stub and write down which sentence each failure must produce. That note is the durable artifact, and it stays useful even if you never adopt an outside host. Keep the paid path for the few summaries that actually need it, and keep the cheap path behind the same probes. If you later want a non-laptop host, start from the current project docs and wire the probes before you trust another green check.

Top comments (0)