DEV Community

Chen Yuan
Chen Yuan

Posted on Originally published at dispatch-blog.hashnode.dev

A Retry Is Not a Recovery Plan: Designing Publish Jobs for Ambiguous Success

A publishing automation can fail after doing exactly what it was asked to do. The browser clicks Publish, the site accepts the article, and the connection drops before the response arrives. The worker sees a timeout. The publication sees a new article. A scheduler that treats every timeout as permission to retry creates a duplicate. A scheduler that treats the timeout as a permanent failure may leave a perfectly good article unrecorded.

This is not peculiar to writing platforms. Payment requests, webhook consumers, email deliveries, and infrastructure changes all encounter the same ambiguity: the caller knows that it sent an operation, but cannot prove whether the remote system committed it. For browser publishing, the ambiguity is especially visible because there may be several screens between the editor and the public URL.

The useful engineering goal is not to pretend that network delivery can be made infallible. It is to design a small recovery protocol that makes uncertainty explicit. The protocol should survive a crashed process, a restarted scheduler, and a different operator taking over the next morning. A single retry counter cannot represent these facts.

The dangerous interval is after submission, not before it

Consider a job that prepares a long Markdown article, opens the editor, applies tags and a cover image, and presses Publish. Before the button press, most failures are reversible. If a title did not fill correctly, the worker can inspect the editor, correct the field, and try the local interaction again. Once the press may have reached the server, the risk changes.

A common implementation wraps the entire job in a broad exception handler. If any stage raises an error, the outer queue retries the whole task. This is convenient for an image-download failure and unsafe for an uncertain publish response. The same retry policy should not govern both operations.

The exact point of uncertainty is not merely a successful HTTP status. For browser tasks it begins when a client dispatches an action with possible permanent effects. The page may navigate without returning a clean automation response. The platform might save a draft, schedule publication, or publish immediately. These outcomes need different follow-up checks.

Treat the action boundary as a state transition. Before calling the browser, write durable evidence that the attempt is about to happen. If the worker later disappears, another worker can distinguish an untouched task from one whose commit result is unknown. The transition should happen before the action, even though that means a crash may mark an attempt that never actually reached the platform. Conservatively verifying a nonexistent publication is much safer than blindly publishing twice.

Give each job a stable identity that survives a restart

A queue job ID is often an implementation detail. It may change when the message is moved, replayed, or re-created by a new scheduler. A publishing identity should describe the intended content and target: publication account, approved article revision, and the planned slot. A deterministic key can help you identify the same logical request even when the worker process changes.

The revision part matters. A correction to an article should not silently be treated as an identical request. Conversely, starting the same content in a new worker should not manufacture a fresh publishing identity. The system must choose that policy explicitly rather than relying on timestamps generated inside a retry loop.

Here is a minimal example. Its hash is for local matching; it is not a platform-provided idempotency guarantee.

import hashlib
import json

def publish_identity(site, account, slot, title, body):
    revision = hashlib.sha256(body.encode("utf-8")).hexdigest()
    payload = {
        "site": site,
        "account": account,
        "slot": slot,
        "title": title.strip(),
        "revision": revision,
    }
    encoded = json.dumps(payload, sort_keys=True).encode("utf-8")
    return hashlib.sha256(encoded).hexdigest()
Enter fullscreen mode Exit fullscreen mode

Keep the identity and a content hash alongside the approved source. Later, if a published item has the same title but a different body, that is not proof that the intended job succeeded. It could be an old article, another revision, or a name collision. Identity should make verification stricter, not provide an excuse to accept a weak match.

Store the attempt before touching the irreversible button

A durable run record can be remarkably small. It needs a job identity, source revision, target account, current state, attempt timestamp, candidate draft URL, any known public URL, and a brief history of observed transitions. The record should not contain passwords, live session cookies, or a copy of browser authentication material.

One practical state machine separates local preparation from platform reconciliation:

PREPARED
  -> EDITOR_VERIFIED
  -> COMMIT_POSSIBLE
  -> PUBLISHED_VERIFIED

COMMIT_POSSIBLE
  -> RECONCILE_ONLY
  -> PUBLISHED_VERIFIED

PREPARED or EDITOR_VERIFIED
  -> REPAIR_REQUIRED
Enter fullscreen mode Exit fullscreen mode

The important state is COMMIT_POSSIBLE. It means the system cannot promise the action was never delivered. The next invocation is not allowed to submit again merely because the worker lost its response. Instead, it must inspect the platform through read-only paths.

Write state changes atomically where practical. On a local filesystem, a process can write a complete JSON record to a temporary file and replace the prior record. On a database, use a transaction and a uniqueness constraint around the logical job identity. In a distributed queue, acquire a lease or use compare-and-swap so that two workers cannot both cross the action boundary. Each method has different crash and recovery behavior, but all are stronger than keeping the attempt only in process memory.

Do not label an attempt as successfully published when the button was clicked. A click confirms an action was issued, not that the public page now exists. That distinction is the core of the protocol.

Reconciliation is a read operation with a precise success contract

After an ambiguous response, visit the platform's own published list or the expected public article URL in the same authorized account context. First look for a durable identifier: the site's post ID, canonical URL, or publication record. Then verify enough content to ensure that this is the intended revision.

A useful success contract checks the author and publication, exact normalized title, a distinctive passage near the end, and the expected publication status. For technical posts it can also inspect the number of code blocks and the cover image. A slug alone is not sufficient when platforms permit edits, collisions, or redirects. The public page must actually be visible; a draft editor URL is not proof of publication.

The contract can be expressed as data rather than a vague boolean. An illustrative verifier might return:

{
  "status": "published_verified",
  "account": "dispatch-example",
  "public_url": "https://example.dev/reliable-publish-recovery",
  "checks": {
    "author": true,
    "title": true,
    "tail_passage": true,
    "code_blocks": true,
    "canonical": true
  },
  "verified_at": "2026-10-09T15:00:00Z"
}
Enter fullscreen mode Exit fullscreen mode

These are example values, not evidence of a real publication. In production, every true value should come from a concrete read-back of the page or platform API. Record the raw URL, source hash, and observation time so a later maintainer can assess what was actually checked.

If the site is slow to hydrate, the verifier may need a bounded polling window. A single empty render immediately after navigation is weak failure evidence. However, bounded polling must not become an infinite wait. If the window expires without reliable proof, preserve RECONCILE_ONLY and report the uncertainty accurately.

An idempotency key is useful only when the receiver enforces it

An API can offer a documented idempotency key and store the result of the first accepted request. That is valuable when the server guarantees that a repeated key resolves to the same operation. A browser editor typically does not expose such a contract. Putting a key in the local job file does not make a second click harmless.

Even real API idempotency contracts have details: key retention periods, request-body consistency, account scope, and behavior during concurrent requests. Read the specific API documentation rather than assuming that a string header solves everything. For sites without a supported key, the recovery plan depends on remote observation and a strict at-most-one automatic commit policy.

At-most-one is deliberately different from exactly-once. There are unavoidable cases where a process records COMMIT_POSSIBLE and crashes before it sends the click. The cautious system can then fail to publish, pending verification or manual adjudication. This is a trade-off. For a public article, a missing item can be diagnosed and published later; duplicated posts may produce confusing subscriptions, bad search indexing, or separate comment histories.

Choose the cost of each failure consciously. If the platform provides an official create endpoint with verified idempotency semantics, prefer that contract. If it does not, never describe browser clicks as exactly-once delivery.

Separate editor repair from remote-outcome repair

A resilient runner should decide its next action from the saved state and the latest platform evidence, not solely from the exception type. For example, an editor element becoming hidden before the commit boundary can be fixed by refreshing a snapshot and using a known selector. An exception after the commit boundary calls for verification, not another click.

The decision rule can be written as a simple function. It intentionally avoids calling a network or browser tool itself:

def next_step(state, remote_evidence):
    if remote_evidence.get("verified_publication"):
        return "finalize_receipt"
    if state in {"COMMIT_POSSIBLE", "RECONCILE_ONLY"}:
        return "read_only_reconcile"
    if state == "EDITOR_VERIFIED":
        return "commit_once"
    if state == "PREPARED":
        return "finish_editor"
    return "inspect_and_record"
Enter fullscreen mode Exit fullscreen mode

This is a planning function, not an entire publishing system. It assumes that some trusted component already gathered remote evidence and that the caller enforces the single-commit boundary. For a real implementation, make the transition itself durable and reject invalid transitions rather than trusting a caller-provided state string.

When an external challenge interrupts the public-page check, preserve the exact URL and current attempt state. A normal login or simple challenge may be completed within the existing authorized browser session. It should not trigger creation of a new browser profile, a second worker, or a fresh article. If evidence remains unavailable, stop at a truthful uncertainty state.

Make the receipt useful to the next operator

The final receipt should answer operational questions without requiring a replay of the whole session: which account was targeted, which source revision was approved, which action was attempted, whether publication was independently verified, what public URL was observed, and what work remains. Include a stable run ID and elapsed duration. Separate a technical script result from a business outcome.

A publish script that exits with status zero after filling the editor has not necessarily completed the business task. Likewise, a script that times out immediately after publication may have produced a valid public article. A reliable report can say the script had an error while the business outcome was independently verified. Collapsing both facts into one success flag loses information.

Here is a compact receipt shape:

{
  "run_id": "article-20261009-01",
  "content_revision": "sha256-of-approved-source",
  "platform_state": "PUBLISHED_VERIFIED",
  "attempt_count": 1,
  "public_url": "https://example.dev/reliable-publish-recovery",
  "verification": "public_page_and_content_checked",
  "next_action": "record_metrics_on_next_review"
}
Enter fullscreen mode Exit fullscreen mode

For a blocked run, keep the same keys and use a specific failure stage. Do not invent zero views or zero interactions when those metrics have not been collected. The honest value is unmeasured. Archive the evidence needed to resume; leave unrelated browser tabs and user files alone.

The deployment checklist is shorter than the incident report

The first improvement is to identify every action that may create a permanent remote object. Surround each one with a persisted attempt marker and a platform-specific read-back. The second improvement is to write the verification contract before production: what exact URL, title, identity, and tail passage constitute proof? The third is to constrain retry logic so that uncertainty cannot create a second post.

Next, test three deliberately unpleasant cases in an authorized test environment: crash before the commit call, crash after dispatch but before receiving its result, and successful publication followed by a failed verifier. Each case should produce a distinct state and a deterministic next step. Re-run the worker from a new process to prove that local memory is not the source of truth.

Finally, measure outcomes across actual runs. Useful metrics include ambiguous commit events, publications verified after a tool error, duplicate incidents, and time spent in reconciliation. An automation system that reports only green executions can hide its most expensive defect.

A reliable publishing worker is not one that always finds a way to click Publish. It is one that knows when to click, when to look, and when the evidence is too weak to do either again. The retry queue still has a place, but recovery begins with a durable record and a fresh observation of what the platform actually did.


Originally published on Dispatch.

Top comments (0)