A watermarking job must have a deadline, an explicit terminal-state contract, and a recoverable lease before it reaches production. For healthtech documents shared outside the organization, the deciding constraint is fidelity versus render cost: rerendering every uncertain job protects delivery only if duplicate work is controlled, while waiting forever protects nothing.
Short answer: treat queued, rendering, succeeded, failed, and canceled as a closed state machine; store a monotonic attempt number and lease expiry; stop client polling on every terminal state; and let a reconciler reclaim expired work. A rendering label is an observation, not proof that a worker is alive.
How do you debug a PDF job stuck in progress forever?
The usual failure is missing ownership semantics. A worker claims a document, begins rasterization or PDF rewriting, and disappears before committing a result. The database still says rendering, so the polling client faithfully repeats a request that can never reveal progress. A lost response creates the mirror image: the result exists, but the client never saw the terminal transition.
No magic timeout fixes both cases. The service needs two clocks with different jobs: a short renewable lease says who may currently render, while an overall deadline says when the workflow must stop trying. Make both server-side policy. Browser tabs close, mobile radios sleep, and callers retry; none of them should own liveness.
Start with ownership.
Start by writing the transition table. Keep it small enough to review during an incident.
| Current state | Accepted event | Next state | Operational meaning |
|---|---|---|---|
queued |
lease acquired | rendering |
one attempt owns work until expiry |
rendering |
artifact committed | succeeded |
checksum and location are durable |
rendering |
permanent render error | failed |
retry requires a new decision |
queued or rendering
|
deadline reached | canceled |
polling must stop |
rendering |
lease expires | queued |
reconciler may schedule another attempt |
Reject every other transition. In particular, a late worker must not overwrite a newer attempt. That comparison prevents an old, low-fidelity render from replacing a later artifact after the job was reclaimed.
This design has a real limitation: leases are a poor fit for work that cannot be retried or fenced at its output boundary. If the rendering library writes directly to a shared final file, use a single durable queue consumer or add attempt-specific staging before enabling reclamation. Otherwise, lease expiry creates concurrency without giving the system a reliable winner. The trade-off is deliberate: a longer lease reduces duplicate render cost but delays recovery, while a shorter lease recovers sooner and demands tighter heartbeat and capacity margins. Neither value can be copied from another workload.
Implement bounded polling and fenced commits
The following Go example keeps the transport generic. The illustrative policy allows a caller to wait for 45 seconds, uses capped backoff, honors cancellation, and recognizes every terminal state. Change those values from load tests and the document-sharing SLO; they are configuration, not universal constants.
package watermark
import (
"context"
"errors"
"fmt"
"math/rand"
"time"
)
type State string
const (
Queued State = "queued"
Rendering State = "rendering"
Succeeded State = "succeeded"
Failed State = "failed"
Canceled State = "canceled"
)
type Job struct {
ID string
State State
Attempt int64
Artifact string
Error string
}
type Reader interface {
Get(context.Context, string) (Job, error)
}
func Wait(ctx context.Context, r Reader, id string) (Job, error) {
ctx, cancel := context.WithTimeout(ctx, 45*time.Second)
defer cancel()
delay := 250 * time.Millisecond
for {
job, err := r.Get(ctx, id)
if err != nil {
return Job{}, fmt.Errorf("read job: %w", err)
}
switch job.State {
case Succeeded:
if job.Artifact == "" {
return Job{}, errors.New("succeeded job has no artifact")
}
return job, nil
case Failed, Canceled:
return job, fmt.Errorf("job ended as %s: %s", job.State, job.Error)
case Queued, Rendering:
// Continue within the caller's finite wait budget.
default:
return Job{}, fmt.Errorf("unknown job state %q", job.State)
}
jitter := time.Duration(rand.Int63n(int64(delay/4) + 1))
timer := time.NewTimer(delay + jitter)
select {
case <-ctx.Done():
timer.Stop()
return Job{}, fmt.Errorf("poll deadline: %w", ctx.Err())
case <-timer.C:
}
if delay < 4*time.Second {
delay *= 2
}
}
}
Polling timeout is not job failure. It means this caller exhausted its observation budget. Return the job identifier so another request can inspect it, and keep the durable workflow independent of that HTTP connection.
Keep those meanings separate.
The worker needs fencing too. On claim, atomically increment attempt, set lease_expires_at, and move the row to rendering. Heartbeats may extend that lease only when both the job ID and attempt still match. The final commit uses the same predicate. Zero updated rows means the worker is stale; it must discard its output rather than publish it.
func CommitArtifact(ctx context.Context, db DB, jobID string, attempt int64, uri, sum string) error {
result, err := db.ExecContext(ctx, `
UPDATE watermark_jobs
SET state = 'succeeded', artifact_uri = ?, artifact_sha256 = ?,
lease_expires_at = NULL, finished_at = CURRENT_TIMESTAMP
WHERE id = ? AND state = 'rendering' AND attempt = ?`,
uri, sum, jobID, attempt)
if err != nil {
return fmt.Errorf("commit artifact: %w", err)
}
n, err := result.RowsAffected()
if err != nil {
return fmt.Errorf("read commit count: %w", err)
}
if n != 1 {
return errors.New("stale render attempt")
}
return nil
}
Here DB is the application's narrow database interface. Write artifacts to attempt-specific temporary keys, validate them, and make only the winning key visible after its fenced database commit succeeds.
Budget fidelity instead of hiding it in retries
Watermarking can preserve vector content, fonts, links, and page geometry, or it can rasterize pages and rebuild a document. Those paths have different resource profiles and different chances of changing what a recipient sees. Because PDF behavior is standardized by ISO 32000-2, validation should target the output contract, not the mere fact that a renderer returned success.
A platform team should choose the mode per document class before capacity planning. A clinical summary that must remain searchable may justify a structure-preserving path; a preview intended only for visual review may permit rasterization. Record the selected mode in the job so an operator can explain latency and so a retry cannot silently change fidelity.
| Approach | Fidelity check | Render-cost pressure | Failure handling |
|---|---|---|---|
| Structure-preserving watermark | text extraction, page boxes, fonts, links | parser complexity and validation | fail closed on unsupported input |
| Page rasterization | pixel comparison and page dimensions | CPU, memory, and output size | retry at a declared resolution |
| Hybrid by document class | contract chosen at submission | two paths to operate | route by stored policy, never by retry count |
This is a buy-versus-build decision as much as a rendering decision. A managed renderer can reduce parser maintenance but adds a remote dependency and constrains how leases, retention, and diagnostics integrate with the platform. A self-hosted renderer exposes resource controls and failure evidence but puts patching, sandboxing, and capacity on the on-call team. The useful comparison is ownership of failure modes, not a feature tally.
Keep the SLO split. Submission availability, time to terminal state, and successful fidelity validation are separate signals; combining them lets fast corrupt output mask rendering trouble. Capacity models should use pages, input bytes, rendering mode, and concurrent attempts rather than job count alone. One 200-page raster job is not operationally equivalent to one two-page metadata watermark.
Verify the fix under failure, not only success
A healthy-path test proves very little. Run a deterministic fault matrix in staging: terminate a worker after claim, after temporary upload, and immediately before commit; delay status reads; duplicate delivery; and let the lease expire while the old worker continues. Each case must end in one durable artifact or an explicit terminal error, with no stale attempt able to win.
Then test the document contract. Use synthetic fixtures that cover multiple page sizes, rotated pages, embedded fonts, searchable text, and link annotations. Compare page count and geometry, extract text where preservation is required, verify the watermark on every intended page, and open the output with more than one independent PDF reader. Do not use real patient data in these fixtures.
The minimum operational signals are job age by state, lease-expiration count, attempts per job, render duration by mode and page-count bucket, terminal outcome, and fidelity-validation failures. Alert on the SLO symptom, such as an excessive tail of nonterminal job age, rather than on a single worker restart. Logs should carry job ID and attempt, but document content does not belong in them.
Short tests catch the expensive bug: after a caller times out, confirm that the server either completes within its own deadline or reaches failed or canceled. Never accept permanent rendering as an outcome.
Forever is not a state.
Roll back without reviving stale work
Rollback is a state transition plan, not just a binary deployment. Pause new claims, allow leases for the current version to drain, and keep the status reader compatible with all defined states. If the new rendering mode is the problem, route newly queued jobs back to the previous mode while preserving each existing job's stored policy and attempt number.
Do not bulk-change rendering rows to queued while workers still hold valid leases. That creates duplicate work and removes the fencing evidence needed to decide which output may publish. Reclaim only expired leases, in bounded batches, and watch terminal-state latency plus fidelity failures during recovery.
The durable rule is plain: every accepted job reaches a terminal state within a declared server-side budget, and every published artifact belongs to the current fenced attempt. Polling then becomes a bounded way to observe the workflow, not the mechanism keeping it alive.
References
- ISO 32000-2, Portable Document Format: https://www.iso.org/standard/75839.html
Top comments (0)