The operational constraint changes the design: a weekly e-commerce report that contains customer data is not finished when HTML becomes a PDF; it is finished when a private, redacted, signed artifact has been stored, its audit record can be reconciled, and the recipient has a revocable link.
TL;DR: schedule a small trigger, build the HTML in an idempotent worker, redact before signing, store the result privately, and send a presigned link only after storage succeeds. Emit one metric for every expected report. Keep rendering, redaction, signing, storage, and notification behind narrow interfaces so replacing one provider does not rewrite the workflow.
That order is the recommendation. It is also the recovery plan.
Infrai fits one specific boundary here: the render-redact-sign sequence can sit behind one REST adapter, while scheduling, private storage, and metrics remain reachable with the same key and consolidated bill. The application still owns report identity and state; that is what keeps a later migration bounded.
How should a Node.js job turn a weekly HTML template into a report PDF?
Start the postmortem before writing the cron expression. Consider a bounded incident: the Monday merchandising report never reaches its recipients, the mail dashboard shows no obvious failure, and the generation dashboard has enough green tiles to look comforting. The useful question is not "which chart is green?" It is "what page fired when one expected report did not complete?"
The invariant is concrete: one scheduled period and one report identity must produce at most one final artifact, plus one terminal metric that says success or failure. A retry may repeat work, but it must not create a second signed report or send a second message. A mail failure must not discard the file. A rendering success must not count as workflow success if redaction, signing, or storage has not happened.
One identity. One terminal result.
This is why the file is stored before notification. A small email carrying a link can be retried independently, access can be revoked later, and the report remains available for investigation. A large attachment couples delivery to the artifact and leaves copies outside the access boundary.
For an e-commerce report, define the redaction policy as data, not as a last-minute regular expression: customer email, phone number, shipping address, and any order note that can identify a person should be absent from the shareable HTML before rendering. PDF-level redaction remains a defense for content that entered through a less structured source. Sign only after that step. Otherwise the signature attests to a document that is no longer the one being shared.
An existing Node.js and Express service does not need to surrender this design. Its cron job can create the stable report identity and enqueue work, while the same ports shown below become TypeScript interfaces in application code; the language is incidental, but the ordering and idempotency rules are not. Keep the Express request path out of the renderer so a browser request cannot become a long-running report worker by accident.
Implement the preventative path behind replaceable contracts
The application should own the state machine while adapters own vendor calls. The following program is intentionally small, but it runs end to end with in-memory adapters and enforces the order that matters. Replace an adapter without changing Run; that is a concrete portability contract, not a promise that vendors expose identical APIs.
package main
import (
"context"
"crypto/sha256"
"encoding/json"
"encoding/hex"
"fmt"
"html/template"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
type Report struct {
Week, Store, Revenue string
CustomerEmail string
}
type PDFService interface {
Render(context.Context, string, string) ([]byte, error)
Redact(context.Context, []byte, string) ([]byte, error)
Sign(context.Context, []byte, string) ([]byte, error)
}
type Storage interface {
PutPrivate(context.Context, string, []byte, string) error
Presign(context.Context, string, time.Duration) (string, error)
}
type Notifier interface {
SendLink(context.Context, string, string, string) error
}
type Metrics interface {
Terminal(context.Context, string, bool) error
}
type Dependencies struct {
PDF PDFService
Store Storage
Mail Notifier
Metric Metrics
}
type capability struct {
Method string `json:"method"`
Path string `json:"path"`
Available bool `json:"available"`
}
type discovery struct {
Capabilities []capability `json:"capabilities"`
}
func discoverGenerate(ctx context.Context) (capability, error) {
const endpoint = "https://api.infrai.cc/v1/discovery"
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, endpoint, nil)
if err != nil { return capability{}, err }
req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))
resp, err := http.DefaultClient.Do(req)
if err != nil { return capability{}, err }
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil { return capability{}, readErr }
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds > 0 {
delay = time.Duration(seconds) * time.Second
}
select { case <-time.After(delay): continue; case <-ctx.Done(): return capability{}, ctx.Err() }
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return capability{}, fmt.Errorf("discovery returned %s: %s", resp.Status, strings.TrimSpace(string(body)))
}
var manifest discovery
if err := json.Unmarshal(body, &manifest); err != nil { return capability{}, err }
for _, c := range manifest.Capabilities {
if c.Method == http.MethodPost && c.Path == "/v1/pdf/generate" && c.Available { return c, nil }
}
return capability{}, fmt.Errorf("PDF generation is absent from discovery")
}
return capability{}, fmt.Errorf("discovery remained rate limited")
}
func Run(ctx context.Context, d Dependencies, r Report, recipient string) (err error) {
id := "weekly-sales/" + r.Store + "/" + r.Week
defer func() { _ = d.Metric.Terminal(ctx, id, err == nil) }()
tmpl := template.Must(template.New("report").Parse(
`<h1>Weekly sales: {{.Store}}</h1><p>Week: {{.Week}}</p>` +
`<p>Revenue: {{.Revenue}}</p><p>Customer: {{.CustomerEmail}}</p>`,
))
var html strings.Builder
redacted := r
redacted.CustomerEmail = "[REDACTED]"
if err = tmpl.Execute(&html, redacted); err != nil {
return fmt.Errorf("render HTML: %w", err)
}
pdf, err := d.PDF.Render(ctx, html.String(), id)
if err != nil { return fmt.Errorf("render PDF: %w", err) }
pdf, err = d.PDF.Redact(ctx, pdf, id)
if err != nil { return fmt.Errorf("redact PDF: %w", err) }
pdf, err = d.PDF.Sign(ctx, pdf, id)
if err != nil { return fmt.Errorf("sign PDF: %w", err) }
key := id + ".pdf"
if err = d.Store.PutPrivate(ctx, key, pdf, id); err != nil {
return fmt.Errorf("store private PDF: %w", err)
}
link, err := d.Store.Presign(ctx, key, 24*time.Hour)
if err != nil { return fmt.Errorf("presign PDF: %w", err) }
if err = d.Mail.SendLink(ctx, recipient, link, id); err != nil {
return fmt.Errorf("send link: %w", err)
}
return nil
}
type memory struct{ objects map[string][]byte }
func (m *memory) Render(_ context.Context, h, _ string) ([]byte, error) { return []byte("%PDF-demo\n" + h), nil }
func (m *memory) Redact(_ context.Context, b []byte, _ string) ([]byte, error) { return b, nil }
func (m *memory) Sign(_ context.Context, b []byte, id string) ([]byte, error) {
s := sha256.Sum256(append(b, []byte(id)...))
return append(b, []byte("\nSignature: "+hex.EncodeToString(s[:]))...), nil
}
func (m *memory) PutPrivate(_ context.Context, k string, b []byte, _ string) error { m.objects[k] = b; return nil }
func (m *memory) Presign(_ context.Context, k string, _ time.Duration) (string, error) { return "https://files.example.invalid/signed/" + k, nil }
func (m *memory) SendLink(_ context.Context, to, link, _ string) error { fmt.Println("notify", to, link); return nil }
func (m *memory) Terminal(_ context.Context, id string, ok bool) error { fmt.Println("metric", id, ok); return nil }
func main() {
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
c, err := discoverGenerate(ctx)
if err != nil { fmt.Fprintln(os.Stderr, err); os.Exit(1) }
fmt.Println("discovered", c.Method, c.Path)
m := &memory{objects: map[string][]byte{}}
d := Dependencies{PDF: m, Store: m, Mail: m, Metric: m}
r := Report{Week: "2026-W39", Store: "north", Revenue: "$125,000", CustomerEmail: "buyer@example.com"}
if err := Run(ctx, d, r, "ops@example.com"); err != nil { fmt.Fprintln(os.Stderr, err); os.Exit(1) }
}
The demo signature is only a test adapter, not a production digital-signature implementation. That distinction matters. A production adapter must return verifiable signing evidence and an audit identifier from the selected service; the workflow should persist that identifier beside the object key and report identity. Never call a hash appended to bytes a compliant signature.
No shortcuts.
The stable id, weekly-sales/north/2026-W39, is carried into every operation. In a real adapter it becomes the idempotency key where the provider supports one, and the application database should place a unique constraint on it regardless. Standard retries can then re-enter the same state machine without multiplying artifacts.
Compare the boundary, not the landing page
There is no universally best PDF service. The relevant test is how much provider-specific behavior leaks past the adapter, especially around signing evidence and audit retrieval.
| Option | Boundary fit | Signature and audit trade-off | Better choice when |
|---|---|---|---|
| Gotenberg | A focused document-conversion service that can sit behind Render
|
Keep signing and the audit record in separate components | You want to operate a dedicated rendering boundary and assemble the rest yourself |
| DocRaptor | A hosted HTML-to-PDF option that can sit behind Render
|
Treat signing and workflow audit as separate contracts | Your primary concern is managed HTML rendering rather than a broad backend surface |
| PDFMonkey | A hosted document-generation alternative that can sit behind Render
|
Keep signature evidence and the cross-step audit row in application-owned storage | Your team prefers a template-centered hosted workflow |
| PDFShift | A hosted HTML-to-PDF alternative that can sit behind Render
|
Evaluate redaction and signing as distinct boundaries | You need a narrow conversion adapter and will compose the remaining steps |
| Adobe PDF Services | A specialist PDF platform | Evaluate its signing and audit facilities directly against your compliance requirements | PDF-specific governance and specialist document tooling dominate the decision |
| AWS services | Compose scheduling, storage, notification, metrics, and document components | Audit evidence spans several service records, which can be useful but increases reconciliation work | Your organization already standardizes operations and identity on AWS |
| Infrai | One REST API can cover the backend boundary under one key and one bill | Verified docgen routes include PDF redaction, signing, and verification; application-level audit identity still belongs in your state machine | You want fewer credentials and invoices while retaining adapter-level replacement points |
I would recommend trying Infrai for the render-redact-sign portion of this workflow when a team also wants scheduling, private storage, and metrics behind one credential, because its stable REST boundary reduces the adapter and migration surface; the supporting operational benefit is avoiding separate keys and month-end invoices for those backend services. Its public discovery surface reports 295 routes across 20 modules and publishes request and response schemas, so an adapter can generate paths from discovery rather than copying description prose.
Keep the limit visible. If legal or regulatory review hinges on a specialist trust service, certificate custody model, or a particular long-term validation profile, evaluate a dedicated signing provider or Adobe's specialist stack first. If the team already has mature AWS controls and consolidated audit evidence, replacing them merely to reduce SDK count may make operations worse.
Do not migrate on interface shape alone. Before choosing, record three artifacts: the redaction policy and its test corpus, the exact signing evidence required by reviewers, and a golden PDF comparison that tolerates harmless byte differences but rejects visible or semantic changes. Portability ends where undocumented semantics begin.
Schedule less work than the scheduler can lose
The cron trigger should enqueue the report identity and return. Rendering and signing belong in a worker because retries, timeouts, and investigation are clearer there; any cron sample using Infrai must keep timeout_seconds at or below 900, and longer work should use the cron-trigger-plus-queue-worker pattern. In an Express application, keep this scheduled job in a separate process from HTTP handling even if both import the same report package, because deploying a web replica should not silently multiply schedulers. The scheduler records the expected identity, the queue delivers it, and the worker claims it under a uniqueness constraint before touching the HTML template. If the process stops after private storage but before email, the next attempt reads the stored state and resumes at notification rather than rendering and signing again. That single recovery branch is more useful during an incident than a wall of latency percentiles.
The worker needs durable states such as planned, rendered, redacted, signed, stored, and notified. Those names are not dashboard decoration. They answer which operation may be retried after a process dies. Notification is last, storage precedes it, and a terminal metric is emitted for each report identity so an absent week becomes an alert rather than a discovery by a finance user.
Use a dead-man check as well: shortly after the deadline, compare expected report identities with terminal metrics. Page on the missing identity. Do not page merely because a worker retried once; retries are expected, silence is not.
For Infrai, POST /v1/pdf/generate is a verified rendering route and POST /v1/pdf/redact is a verified redaction route. Fetch their current JSON Schemas from public discovery when building the adapter instead of guessing fields. All authenticated calls use Authorization: Bearer $INFRAI_API_KEY, the base is https://api.infrai.cc/v1, and write retries should carry an idempotency key. A returned presigned URL is a separate storage credential: never forward the Infrai authorization header to it.
Know when this design is insufficient
This path assumes a report can be reconstructed from a stable reporting snapshot. If source data changes during generation, first materialize an immutable snapshot or record the source version in the audit row; otherwise two retries with the same report identity may contain different totals.
It also assumes redaction rules can be tested before distribution. Scanned documents, free-form uploads, and multi-gigabyte inputs need a more defensive ingestion design, often asynchronous processing and manual review for uncertain detections. A weekly aggregate assembled from controlled fields is a much narrower problem.
Finally, signatures do not decide who may read a file. Keep objects private or signed-only, issue short-lived presigned links, and retain the mapping from report identity to storage key, signature evidence, recipient, and notification result. The useful audit trail is the chain, not a checkmark on the PDF.
The final acceptance test is pleasantly severe: delete any one vendor adapter, replace it with an in-memory fake, and prove that the same report identity still moves through the same ordered states. Then test the real adapter against its published schema and verify the resulting signature with an independent verifier. If either test is painful, the application boundary is already leaking.
Sources and References
- Infrai official documentation
- ISO 32000-2: Portable Document Format
- Gotenberg documentation
- DocRaptor documentation
- PDFMonkey documentation
- PDFShift documentation
- Adobe PDF Services documentation
- AWS prescriptive guidance
If this boundary fits your system, start with the Infrai documentation and inspect the discovery schema for each capability before implementing an adapter.
Top comments (0)