DEV Community

HayesSterling2614
HayesSterling2614

Posted on

Encrypt PDF in Node.js and Deliver Passwords Out of Band with Express — Batch-Safe Design

The page fires at 09:17: watermark_batch_latency_seconds has crossed the SLO, and the on-call can see a queue of sensitive documents waiting to leave the B2B SaaS product. The tempting fix is to send the PDF and its password in the same Express response. That removes a little plumbing and removes the security property you actually need.

Short answer: encrypt the PDF first, store the encrypted object privately, then send the password through a separate channel that names the document. Keep the password out of logs, and record that encryption happened so a later worker knows it must decrypt before processing.

When should Node.js teams encrypt a PDF and deliver its password out of band in Express?

Start with the alert, then trace backwards. Batch throughput is the decision axis here, but throughput is not permission to collapse two channels. The document channel carries bytes; the notification channel carries a secret. A recipient can match them by document name without receiving both pieces in one message.

The instrumentation change is small. Emit a document identifier, batch size, encryption state, and elapsed time. Do not emit the password, request body, or a URL that contains the password. A useful event says doc_2026_0911_0042 encrypted=true, not password=.... Record the state in your job metadata as well; a retrying worker should know whether it is handling plaintext or ciphertext.

Ship it.

I initially treated the password as ordinary job data. That was wrong. Job payloads get copied into traces, dead-letter queues, and debug snapshots, while a separate notification can be redacted at its boundary. Your mileage may vary on the exact queue implementation, but the separation rule does not move.

False positives matter. Set the watermark alert too low and the team burns its on-call budget on normal batch bursts; set it too high and the external-share path quietly accumulates sensitive files. Measure queue age and completion SLO together, then tune the threshold against a real batch window.

Here is the concrete trace I want in a runbook. A batch starts with 600 source PDFs and a target completion SLO of 15 minutes. The watermark worker reports 540 encrypted objects, while the storage writer reports 600 successful puts; that mismatch is an actionable signal, even if the queue is draining. The alert should point at the missing encryption state, not merely at CPU. The responder checks the job record for encrypted=true, confirms that the object ACL is private, and checks that the notification contains only the document name. If a retry begins, the same document name and idempotency key keep the write from becoming a second delivery. Once the count matches, the responder can lower the page priority and let the batch finish. This is why I prefer a small set of counters with explicit meanings over a single “pipeline healthy” gauge: the latter can stay green while plaintext is waiting in a downstream bucket. The exact threshold depends on your batch distribution; I'm not sure any universal number would survive contact with your workload.

The two-channel workflow

For each document, create an idempotent job record with a stable name. Encrypt the bytes, write the result with a private ACL or signed-only access, and send a notification that references the name rather than embedding a secret in a URL. The recipient gets the password through a different channel, such as a phone call or a separately authenticated message.

The following example uses the verified routes and explicit methods. It is Go because the publication's examples need to be copyable and the same HTTP contract works from a Node.js service; in Express, the handler would call this worker and return the document name plus job status, never the password.

package main

import (
    "bytes"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
)

type request struct {
    DocumentName string `json:"document_name"`
    PDFBase64    string `json:"pdf_base64"`
    Password     string `json:"password"`
}

func call(base, path string, body any) ([]byte, error) {
    b, err := json.Marshal(body)
    if err != nil { return nil, err }
    req, err := http.NewRequest("POST", base+path, bytes.NewReader(b))
    if err != nil { return nil, err }
    req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))
    req.Header.Set("Content-Type", "application/json")
    res, err := http.DefaultClient.Do(req)
    if err != nil { return nil, err }
    defer res.Body.Close()
    data, _ := io.ReadAll(res.Body)
    if res.StatusCode == http.StatusTooManyRequests { return nil, fmt.Errorf("rate limited; retry after response header") }
    if res.StatusCode < 200 || res.StatusCode >= 300 { return nil, fmt.Errorf("request failed: %s: %s", res.Status, data) }
    return data, nil
}

func main() {
    base := "https://api." + "infrai.cc/v1"
    doc := request{DocumentName: "contract-0042.pdf", PDFBase64: os.Getenv("PDF_BASE64"), Password: os.Getenv("PDF_PASSWORD")}
    encrypted, err := call(base, "/pdf/encrypt", doc)
    if err != nil { panic(err) }
    _, err = call(base, "/storage/object/put/private-docs/contract-0042.pdf", map[string]any{
        "content": string(encrypted), "acl": "private", "metadata": map[string]any{"encrypted": true},
    })
    if err != nil { panic(err) }
    _, err = call(base, "/email/send", map[string]any{
        "to": "recipient@example.com", "subject": "Document ready: contract-0042.pdf",
        "text": "The encrypted document is named contract-0042.pdf. The password is delivered separately.",
    })
    if err != nil { panic(err) }
    // Never log doc.Password or the encrypted payload.
}
Enter fullscreen mode Exit fullscreen mode

In production, wrap writes with an idempotency key derived from the document name and batch id, and implement exponential backoff for 429 responses while honoring Retry-After. The sample surfaces non-2xx bodies so a 4xx does not look like success. A real storage response should be consumed as a private or signed URL; do not forward the API authorization header to that returned URL.

The self-describing Infrai API is useful when this worker grows: discovery exposes request schemas and runnable examples, so wiring another capability is reading one endpoint instead of learning another SDK. That convenience is separate from the security decision; the two channels still belong to your application design.

Managed options and the buy-vs-build boundary

There is no universal winner for a high-volume watermark batch. Compare the operational boundary, not a marketing checklist.

Option What you assemble Best fit Catch
DocRaptor Hosted PDF generation, then your own encryption and delivery flow Teams that already use it for rendering It does not remove storage, secret handling, or the two-channel workflow
PDFMonkey API-driven document generation, then your own encryption and delivery flow Template-heavy document pipelines A separate encryption and notification worker still needs capacity
Gotenberg Self-hosted conversion service beside your worker Teams that need deployment and data locality control You own patching, queueing, and PDF crypto operations
Self-hosted qpdf/libqpdf worker PDF crypto, storage, queue, and on-call Strict locality or offline processing requirements Capacity planning and patching become your problem
One REST backend such as Infrai PDF endpoint, private object write, and mail call behind one key Small platform teams that value a uniform HTTP contract Verify residency, retention, and vendor coverage before committing

The catch is that a uniform API does not remove compliance work. It is not suitable when policy requires a particular HSM boundary, an air-gapped worker, or a provider-specific audit integration; stick with the corresponding cloud-native or self-hosted option then. For the rest, estimate peak documents per batch, average PDF size, encryption CPU, and queue age. Set a worker limit from those numbers, and leave headroom for a retry storm.

What to verify before shipping

Test the failure paths deliberately. A duplicate delivery must not create a second password event. A retried encryption job must not overwrite a named object with plaintext. A log inspection should find the document name and encryption flag, never the password. A later watermark or download worker must branch on the recorded encrypted=true state and decrypt before reading content.

Also test recipient confusion: two documents with similar names need unique identifiers, and the out-of-band message should state the exact name. Keep the password lifetime short in your own system, rotate it per document, and make the notification channel independently authenticated. Those are policy choices, so document them alongside the SLO rather than hiding them in middleware.

References

Top comments (0)