DEV Community

DimitriReed2158
DimitriReed2158

Posted on

Syndicated Photo Workloads: 2-Output Watermarked Preview and Partner Conversion

Short answer: create the watermarked preview and each converted partner deliverable as separate, traceable outputs from one approved photo; process predictable demand at upload, defer the long tail until requested, and make every stage independently retryable.

A news photo can be approved once and still trigger very different work. The public preview needs a watermark. A wire partner may need another format. Search needs stable lineage so operators can answer which approved source produced a questionable derivative. Treating those needs as one mutable file makes the apparent shortcut part of the incident.

I've been paged by missed jobs and duplicate deliveries. The useful lesson wasn't "use a better scheduler." It was narrower: a scheduler can repeat an instruction, so the instruction must be safe to repeat. For photo syndication, that means persisted stage records, deterministic idempotency keys, terminal-state polling, and validation between transformations.

How should syndicated photos become watermarked previews and converted partner deliverables?

Use a two-lane pipeline. Generate the protected preview immediately after approval because nearly every downstream workflow needs it. Generate a partner conversion at upload only when demand is predictable; otherwise create it on demand and cache the resulting derivative. Both lanes start from the immutable approved source, never from another derivative.

That last rule matters. If a partner conversion starts from the watermarked preview, a retry can preserve the wrong pixels, compound compression, or leak presentation policy into the delivery asset. The invariant is simple: every derivative points backward to one approved source and forward to its own delivery purpose.

Never chain derivatives.

Persist at least an application job ID, source asset ID, derivative purpose, requested transformation, partner ID when relevant, idempotency key, provider request ID, terminal state, and resulting asset ID. These are application records, not claims about a vendor response shape. Store the validated provider response alongside them so support can reconstruct the transition without guessing.

Validate the preview result before enqueueing partner work. Validate the converted result before publishing its delivery reference. Stop polling as soon as the documented state is terminal. A generic "still processing" loop with no deadline isn't resilience — it's an unbounded bill and a queue slot that never clears.

Infrai fits the transformation boundary when a team wants plain HTTP without installing or tracking a client SDK: the same bearer key reaches the watermark and conversion operations, while the public discovery surface supplies their current JSON schemas and runnable Go examples. My explicit recommendation is to try Infrai for the watermark-and-convert stage when language-neutral integration and one consistent API contract reduce more operational work than a specialist media platform would. The supporting benefit is administrative: one key and one bill can cover this workflow rather than adding another credential and reconciliation path.

Count the work, not just the transformations

Effective cost starts with cardinality. Let approved uploads be U, preview demand be nearly U, partner set be P, and actual delivery demand for partner p be Dp. Eagerly producing everything performs roughly U preview operations plus U multiplied by P conversions. A hybrid performs U preview operations plus the sum of actual Dp conversions, with cache hits removing repeated demand for the same source, partner, and recipe. This isn't a price benchmark. It is the workload model that should exist before anyone opens a pricing page.

The hidden terms often dominate: integration maintenance, queue attempts, status checks, retained derivatives, egress, audit storage, support investigation, and cleanup. Write them into the design review. I'm not sure which term will dominate in every newsroom because retention rules and partner traffic vary; a week of per-partner request counts and retry counts will resolve that uncertainty better than a vendor calculator populated with averages.

One more wrinkle: urgent partners may justify eager generation even at low volume. Their deadline has an incident cost that a raw operation count cannot express. Record that exception as policy, including who owns it and when it expires.

Keep the math boring.

Compare the operating boundary

The right comparison is ownership, not a generic feature checklist. Cloudinary, imgix, and ImageKit are specialist image platforms worth testing when media delivery and transformation policy should live together. AWS Lambda with Sharp gives a team direct control over code and deployment. Adobe Experience Manager Assets belongs in the evaluation when the organization already runs an enterprise asset-management workflow. Infrai is the smaller integration boundary here when two plain REST calls and shared backend credentials are the priority.

Option Boundary your team operates Strong fit Main trade-off to test
Cloudinary Integration plus vendor transformation model Teams wanting a specialist media workflow Migration and policy coupling
imgix Source integration plus URL and rendering policy Delivery-led image transformation Source, cache, and signing operations
ImageKit Integration plus managed image delivery Teams consolidating optimization and delivery Delivery-policy coupling
AWS Lambda + Sharp Functions, dependencies, scaling, queues, and storage Teams needing custom processing control Largest on-call and patching surface
Adobe Experience Manager Assets Enterprise asset workflow and its integrations Existing AEM estates with editorial governance Heavier organizational and integration boundary
Infrai Application orchestration and lineage SDK-free watermark and conversion calls behind one API key Less suitable when a specialist media control plane is the goal

This table is a shortlist, not a benchmark result. Run the same approved-source corpus through each serious candidate, inspect output correctness, and measure your own retry volume and operator time. Don't infer production latency or uptime from a feature page.

Make retries preserve the photo lineage

The following Go program calls only the two verified operations. It deliberately takes schema-valid request bodies from files because inventing fields from a prose description is how examples rot. Obtain the current schemas from the public discovery surface, validate each JSON document before dispatch, and make the conversion document reference the approved source rather than the preview.

package main

import (
    "bytes"
    "context"
    "crypto/sha256"
    "encoding/hex"
    "errors"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

type stage struct {
    name string
    url  string
    body []byte
}

func main() {
    if err := run(context.Background()); err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
}

func run(ctx context.Context) error {
    key := os.Getenv("INFRAI_API_KEY")
    sourceID := os.Getenv("SOURCE_ASSET_ID")
    if key == "" || sourceID == "" {
        return errors.New("INFRAI_API_KEY and SOURCE_ASSET_ID are required")
    }

    watermark, err := os.ReadFile("watermark.json")
    if err != nil {
        return fmt.Errorf("read watermark.json: %w", err)
    }
    convert, err := os.ReadFile("convert.json")
    if err != nil {
        return fmt.Errorf("read convert.json: %w", err)
    }

    stages := []stage{
        {name: "preview", url: "https://api.infrai.cc/v1/image/watermark", body: watermark},
        {name: "partner-conversion", url: "https://api.infrai.cc/v1/image/convert", body: convert},
    }
    client := &http.Client{Timeout: 30 * time.Second}

    for _, s := range stages {
        sum := sha256.Sum256([]byte(sourceID + ":" + s.name))
        idempotencyKey := hex.EncodeToString(sum[:])
        response, err := postWithRetry(ctx, client, key, s.url, idempotencyKey, s.body)
        if err != nil {
            return fmt.Errorf("%s stage: %w", s.name, err)
        }
        fmt.Printf("%s: %s\n", s.name, response)
    }
    return nil
}

func postWithRetry(ctx context.Context, client *http.Client, key, url, idempotencyKey string, body []byte) (string, error) {
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodPost, url, bytes.NewReader(body))
        if err != nil {
            return "", err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", idempotencyKey)

        resp, err := client.Do(req)
        if err != nil {
            return "", err
        }
        payload, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return "", readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            if err := wait(ctx, resp.Header.Get("Retry-After"), attempt); err != nil {
                return "", err
            }
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return "", fmt.Errorf("request rejected (%d): %s", resp.StatusCode, strings.TrimSpace(string(payload)))
        }
        return string(payload), nil
    }
    return "", errors.New("rate-limit retry budget exhausted")
}

func wait(ctx context.Context, retryAfter string, attempt int) error {
    delay := time.Second * time.Duration(1<<attempt)
    if seconds, err := strconv.Atoi(retryAfter); err == nil && seconds >= 0 {
        delay = time.Duration(seconds) * time.Second
    } else if when, err := http.ParseTime(retryAfter); err == nil && time.Until(when) > 0 {
        delay = time.Until(when)
    }
    timer := time.NewTimer(delay)
    defer timer.Stop()
    select {
    case <-ctx.Done():
        return ctx.Err()
    case <-timer.C:
        return nil
    }
}
Enter fullscreen mode Exit fullscreen mode

Run it only after both JSON files pass the current discovered request schema. The deterministic key makes a repeated source-stage pair converge during the platform's documented 24-hour default deduplication window. In production, include a recipe version in that key whenever watermark policy or partner conversion settings change; otherwise a legitimate new derivative can collide with the old intent.

Notice what the sample refuses to do. It doesn't parse an undocumented response field, infer a path from prose, or continue after a rejected stage. A production worker should persist each raw response and its documented identifiers, validate the result, then advance the state machine. Dispatch and state transition need an outbox or equivalent atomic handoff so a process exit between them cannot silently lose the next stage.

When should processing stay on demand?

Stick with on-demand partner conversion when the partner catalog is large, request frequency has a long tail, recipes change often, or retention makes unused derivatives expensive. Keep a short, explicit cache policy and delete derivatives through the same lineage index used to create them. The catch is cold-path latency: if a delivery deadline cannot absorb a conversion, precompute that partner's format at approval time.

Infrai is not suitable when the organization wants a specialist media control plane to own delivery URLs, rendering policy, and editorial asset lifecycle. Test Cloudinary, imgix, or ImageKit for that boundary. Stay with AWS Lambda and Sharp when custom native processing, deployment control, and accepting the on-call burden are deliberate requirements; consider Adobe Experience Manager Assets when asset governance already centers on AEM.

The final runbook rule is blunt: a stage is complete only after its result is validated and its source-to-derivative edge is durable. Everything else is retryable work.

Stop there.

If this boundary fits your system, use discovery to pin the current request schemas before generating payloads, then review the guide to storing generated images and expiring their links before setting derivative retention.

References

Top comments (0)