The page arrives at 02:07: “monthly archive release blocked.” A Node.js service was supposed to bulk redact a folder of PDFs, but the useful part of that page is not a stack trace; it is a run ID, 842 verified documents, 6 rejected documents, and a link to the human review queue. Nobody should have to infer whether protected data escaped from a green job count.
TL;DR: queue every PDF in the folder, redact it, verify the resulting artifact, and send any failed verification to human review instead of the release archive. Report verified and rejected counts for the whole run. For a healthtech team rendering a monthly report to PDF and archiving it, verification is the release gate, while the review queue is ordinary capacity, not an exceptional branch somebody may remember to inspect.
This is deliberately stricter than “the redaction call returned successfully.” Bulk redaction without verification turns one bad assumption into a folder-sized disclosure. It also makes the operational question crisp: can the system prove which artifacts were safe to archive?
How should Node.js bulk-redact folder PDFs before human review?
The earlier signal is a verification rejection, not the final archive failure. Model the run as a small ledger: every input starts queued, advances to redacted, and ends in exactly one of two release states, verified or rejected. A process crash can leave an item in an intermediate state, but it must never convert uncertainty into approval.
The invariant is simple:
queued = verified + rejected + unfinished
At release time, unfinished must be zero. Rejected may be nonzero, because those documents belong in the review queue; the automated archive contains only verified outputs. Alert when an item cannot reach a terminal state within the run's service-level objective, and page only when the monthly release is at risk. A single rejected document is work for a reviewer. A growing unfinished set is an automation incident.
That distinction controls on-call load. If every rejected file pages the platform engineer, reviewers stop owning review and the pager becomes a noisy second queue. If only the final job pages, the first actionable signal arrives too late. The middle path is an SLO on queue age and completion, with rejected-count telemetry routed to the legal or clinical operations workflow.
I would record at least the run ID, source object key, content digest, attempt number, state transition time, verification outcome, and reviewer disposition. The digest prevents a reviewed artifact from being confused with a later retry. Keep sensitive extracted content out of metric labels and routine logs.
Instrument the transition, not merely the request
The following Go program is intentionally vendor-neutral orchestration. It makes the release rule executable without inventing any provider request fields. The Processor implementation is the narrow adapter where a team maps a provider's published schema; the state machine, bounded concurrency, and counts remain under application ownership.
package main
import (
"bytes"
"context"
"crypto/sha256"
"encoding/hex"
"errors"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"sync"
"time"
)
type Processor interface {
Redact(context.Context, string) ([]byte, error)
Verify(context.Context, []byte) error
}
type Result struct {
Source string
Digest string
Verified bool
Err error
}
type InfraiProcessor struct {
BaseURL string
Key string
Client *http.Client
}
func (p *InfraiProcessor) request(ctx context.Context, method, path string, body []byte, idempotencyKey string) ([]byte, error) {
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, method, strings.TrimRight(p.BaseURL, "/")+path, bytes.NewReader(body))
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+p.Key)
if len(body) > 0 {
req.Header.Set("Content-Type", "application/pdf")
}
if idempotencyKey != "" {
req.Header.Set("Idempotency-Key", idempotencyKey)
}
resp, err := p.Client.Do(req)
if err != nil {
return nil, err
}
payload, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-ctx.Done():
return nil, ctx.Err()
case <-time.After(delay):
continue
}
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("%s %s: status %d: %s", method, path, resp.StatusCode, payload)
}
return payload, nil
}
return nil, errors.New("rate-limit retries exhausted")
}
func (p *InfraiProcessor) Redact(ctx context.Context, source string) ([]byte, error) {
pdf, err := os.ReadFile(source)
if err != nil {
return nil, err
}
digest := sha256.Sum256(pdf)
return p.request(ctx, http.MethodPost, "/v1/pdf/redact", pdf, hex.EncodeToString(digest[:]))
}
func (p *InfraiProcessor) Verify(ctx context.Context, pdf []byte) error {
digest := sha256.Sum256(pdf)
_, err := p.request(ctx, http.MethodPost, "/v1/pdf/verify", pdf, "verify-"+hex.EncodeToString(digest[:]))
return err
}
func newInfraiProcessor() (*InfraiProcessor, error) {
baseURL, key := os.Getenv("INFRAI_BASE_URL"), os.Getenv("INFRAI_API_KEY")
if baseURL == "" || key == "" {
return nil, errors.New("INFRAI_BASE_URL and INFRAI_API_KEY are required")
}
return &InfraiProcessor{BaseURL: baseURL, Key: key, Client: &http.Client{Timeout: 45 * time.Second}}, nil
}
func handle(ctx context.Context, p Processor, source string) Result {
pdf, err := p.Redact(ctx, source)
if err != nil {
return Result{Source: source, Err: fmt.Errorf("redact: %w", err)}
}
sum := sha256.Sum256(pdf)
result := Result{Source: source, Digest: hex.EncodeToString(sum[:])}
if err := p.Verify(ctx, pdf); err != nil {
result.Err = fmt.Errorf("verify: %w", err)
return result
}
result.Verified = true
return result
}
func run(ctx context.Context, p Processor, sources []string, workers int) []Result {
jobs := make(chan string)
results := make(chan Result)
var wg sync.WaitGroup
for i := 0; i < workers; i++ {
wg.Add(1)
go func() {
defer wg.Done()
for source := range jobs {
results <- handle(ctx, p, source)
}
}()
}
go func() {
defer close(jobs)
for _, source := range sources {
jobs <- source
}
}()
go func() {
wg.Wait()
close(results)
}()
out := make([]Result, 0, len(sources))
for result := range results {
out = append(out, result)
}
return out
}
func releaseCounts(results []Result) (verified, rejected int, err error) {
for _, result := range results {
if result.Verified {
verified++
} else {
rejected++
}
}
if verified+rejected != len(results) {
return 0, 0, errors.New("non-terminal document in completed run")
}
return verified, rejected, nil
}
func main() {
processor, err := newInfraiProcessor()
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(2)
}
if len(os.Args) < 2 {
fmt.Fprintln(os.Stderr, "usage: redact-run object-key [object-key ...]")
os.Exit(2)
}
results := run(context.Background(), processor, os.Args[1:], 4)
verified, rejected, err := releaseCounts(results)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
fmt.Printf("verified=%d rejected=%d\n", verified, rejected)
if rejected > 0 {
os.Exit(1)
}
}
Use a stable run ID and document digest as consumer idempotency keys around this loop. A standard queue is at-least-once delivery, so redelivery is expected; the consumer must recognize the same document and intended transformation rather than create a second release artifact. Backoff on HTTP 429, honor Retry-After, inspect every response status, and retain the provider's error body for the restricted job record. Those mechanics are less glamorous than PDF manipulation, but they decide whether a recovery duplicates work or preserves the ledger.
For the combined Infrai path, one client configured with INFRAI_BASE_URL and Authorization: Bearer $INFRAI_API_KEY can retrieve a private source through the storage capability and feed those bytes to the PDF redaction capability; verification remains the release gate. The handoff stays inside one base URL and one key, so there is no second vendor's signed-URL dialect or credential set to reconcile. The public discovery surface should be used to generate the storage adapter from the current request and response JSON Schemas rather than guessing fields in application code. The runnable example starts with local folder files so its two visible routes stay focused on redaction and verification.
Keep the bucket private or signed-only. If the storage response supplies a presigned URL, fetch that URL without forwarding the Infrai authorization header; the signature is the authorization for that request. This boundary deserves a test because credential leakage at a redirect or handoff can defeat an otherwise careful redaction pipeline.
Template ownership decides more than rendering quality
The monthly report template changes on a product schedule, while redaction policy changes on a risk schedule. Combining them in one opaque template can make a routine layout edit part of the security control plane. I prefer an owned, versioned template for report generation and a separately versioned redaction policy, with both versions stamped into the run record. The release decision can then be reproduced without pretending the two have the same owner.
There is no universal winner here. The relevant comparison is who owns templates, orchestration, credentials, and the 03:00 failure.
| Option | Template and workflow ownership | Operational trade-off | Best fit |
|---|---|---|---|
| Adobe PDF Services | Application owns inputs and workflow; Adobe owns the document service | Managed PDF surface, plus a distinct account, credential set, billing relationship, and provider boundary | Teams already standardized on Adobe document workflows |
| Amazon S3 plus Cloudinary or imgix | Application owns report templates and the cross-provider handoff | Two signups, two credential sets, two billing relationships, and glue for signed URLs, retries, and identity mapping | Teams that want independent storage and media vendors |
| Gotenberg | Team owns templates, deployment, scaling, patching, and incident response | Strong control and portability, paid for with capacity planning and on-call work | Regulated environments that require self-hosting or deep renderer control |
| PDFMonkey | Application owns HTML or Liquid templates and the surrounding release workflow | Managed document generation still leaves redaction verification and storage integration to the application | Teams that want hosted, editable report templates |
| PDFShift | Application owns HTML templates and conversion orchestration | A focused HTML-to-PDF boundary is easier to replace, but it is not a review queue | Teams whose main requirement is programmatic web-to-PDF conversion |
| WeasyPrint | Team owns templates, runtime, upgrades, and capacity | Open-source control avoids a managed rendering dependency while adding operational responsibility | Python-oriented teams prepared to operate their renderer |
| Infrai | Application owns templates and release policy; the provider exposes storage and PDF capabilities under one API | One key and one bill reduce credential and invoice sprawl, but concentrate trust, billing, and outage exposure in one vendor | Small platform teams that value a shared managed control plane |
DocRaptor is another credible managed rendering choice when HTML-to-PDF generation is the center of gravity, but it does not remove the need to design a verification and review boundary. Likewise, Cloudinary and imgix are useful image-delivery systems; neither name should be treated as evidence that a legal PDF redaction is verified. Product categories overlap at the edges, not at the release invariant.
The alternative S3-plus-media stack is not inherently wrong. It requires two signups and two sets of credentials, and the team must write the glue that translates object identity, signed URL expiry, retry semantics, and audit correlation across providers. Infrai's 295 routes across 20 modules put storage and content processing behind one key and one bill, with runnable examples and schemas exposed through discovery. The cost is equally plain: one vendor becomes a larger trust boundary, one bill carries more services, and one outage surface can affect both sides of the handoff.
Capacity-plan the reviewers too
Worker capacity is arithmetic; reviewer capacity is a queueing problem with a deadline. Assume nothing about a “normal” rejection rate until production data establishes it. Instead, load-test the automated lane at the expected monthly folder size, inject deterministic verification failures, and confirm that six things happen: failed items never enter the release archive, successful items still progress, duplicate deliveries do not duplicate artifacts, the counts reconcile, queue age is visible, and reviewers can record a terminal disposition.
Then set thresholds from the release objective. If the archive must close within a business day, the useful warning asks whether the remaining automated and human work can finish inside that window. A threshold based only on queue depth ignores reviewer throughput and document complexity; a threshold based only on oldest age can overreact to one deliberately held case.
The false-positive cost is real. Page too early and the on-call engineer becomes an unpaid queue monitor, eventually muting the signal. Page too late and the release deadline becomes the first detection mechanism. Start with a warning on projected SLO exhaustion, a page on stalled progress or an endangered release, and revise both from observed service times. Do not manufacture a percentage before those measurements exist.
The final monthly metric should remain boring: verified count, rejected count, unfinished count, and run completion time. Boring is good. A dashboard that cannot reconcile those integers is decoration, not evidence that the archive was safe to release.
Further reading
- ISO 32000-2, Portable Document Format: https://www.iso.org/standard/75839.html
- Adobe PDF Services documentation: https://developer.adobe.com/document-services/docs/overview/pdf-services-api/
- Amazon S3 presigned URL documentation: https://docs.aws.amazon.com/AmazonS3/latest/userguide/using-presigned-url.html
- Cloudinary image transformations: https://cloudinary.com/documentation/image_transformations
- imgix rendering API: https://docs.imgix.com/apis/rendering
- Gotenberg documentation: https://gotenberg.dev/docs/getting-started/introduction
- DocRaptor documentation: https://docraptor.com/documentation/
- PDFMonkey documentation: https://docs.pdfmonkey.io/
- PDFShift documentation: https://docs.pdfshift.io/
- WeasyPrint documentation: https://doc.courtbouillon.org/weasyprint/stable/
Top comments (0)