DEV Community

BrodyVance2149
BrodyVance2149

Posted on

Implementing High-Quality Text-to-Image API Gates for Marketing Apps (During Incidents)

Use two lanes: generate routine education ads automatically, but hold typography-heavy or high-value posters for a stricter quality gate. The decision rule is quality first, latency second for posters; latency first within an approved template for social ads. Upscale only after an image passes inspection, because Lanczos can increase pixel dimensions but cannot repair malformed lettering, prompt drift, or a stray hand.

TL;DR: do not page because a similarity score moved. Page only when the service can no longer produce a safe deliverable, the queue threatens a declared deadline, or bad creative is escaping the gate. Everything else belongs in a ticket with the prompt, model, seed when available, output dimensions, and failed checks attached. For an edtech support queue, this keeps a request such as “make a square enrollment ad for the algebra course” fast without pretending that every generated image is ready to publish.

What should actually wake the on-call engineer?

A dashboard can tell me that artifact rate rose. It cannot tell me whether a learner saw a poster with a broken course date. I do not trust a green chart to make that distinction for me. The page must describe the failed user outcome: publishable output is unavailable, the deadline budget is nearly spent, or an unsafe asset passed review. If the alert cannot name which of those happened, it is telemetry, not a page.

Start by translating each incoming support ticket into an explicit service class. A routine social tile gets a short latency budget and a template whose text is rendered by the application after image generation. A campaign poster gets a longer budget, stricter review, and no automatic publication. A request with ambiguous rights, sensitive student imagery, or unclear copy stops before generation. This is triage, not prompt decoration.

The most useful quality signals are prompt adherence, typography correctness, and artifact rate. Resolution matters after those checks. A sharp miss is still a miss.

For every candidate, retain enough context to reproduce the decision: vendor and model selection, normalized brief, requested aspect ratio, output dimensions, timestamps, evaluator results, and a content hash. Avoid recording student personal data in prompts. Also remember that a dedicated moderation endpoint is not universally available; where it is absent, a chat-model check constrained by JSON Schema is a fallback, not proof that an image is safe.

Should a marketing app use a high-quality text-to-image API for every poster?

The generator should not also be the judge. Run mechanical checks first, then require a reviewer for the cases where pixels carry claims: dates, prices, accreditation language, course names, or calls to action. Render important copy in your own layout layer whenever possible. Generators are useful at composition and style, but ad typography is exactly where a plausible-looking error can escape a hurried review.

No. First, establish that the generation service can return a candidate under controlled retry behavior, then apply the mechanical and editorial gates. The following Go program makes one image-generation request and writes the JSON response to standard output; INFRAI_BASE_URL, INFRAI_API_KEY, and INFRAI_IMAGE_MODEL must be set, so neither credentials nor a deployment address are baked into source. It sends one real route, uses a stable idempotency key derived from the request, honors Retry-After, and surfaces every non-success response. That narrow scope is deliberate: response schemas and image transport should be decoded from the live discovery contract by the application adapter rather than guessed in an article.

package main

import (
    "bytes"
    "context"
    "crypto/sha256"
    "encoding/hex"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

type request struct {
    Model  string `json:"model"`
    Prompt string `json:"prompt"`
    Size   string `json:"size"`
}

func main() {
    if len(os.Args) != 2 {
        fmt.Fprintln(os.Stderr, "usage: generate PROMPT")
        os.Exit(2)
    }
    baseURL := strings.TrimRight(os.Getenv("INFRAI_BASE_URL"), "/")
    apiKey := os.Getenv("INFRAI_API_KEY")
    model := os.Getenv("INFRAI_IMAGE_MODEL")
    if baseURL == "" || apiKey == "" || model == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_BASE_URL, INFRAI_API_KEY, and INFRAI_IMAGE_MODEL are required")
        os.Exit(2)
    }

    payload, err := json.Marshal(request{
        Model: model,
        Prompt: os.Args[1],
        Size: "1024x1024",
    })
    if err != nil {
        panic(err)
    }
    sum := sha256.Sum256(payload)
    idempotencyKey := hex.EncodeToString(sum[:])

    ctx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
    defer cancel()
    client := &http.Client{Timeout: 45 * time.Second}

    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodPost,
            baseURL+"/v1/images/generations", bytes.NewReader(payload))
        if err != nil {
            panic(err)
        }
        req.Header.Set("Authorization", "Bearer "+apiKey)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", idempotencyKey)

        resp, err := client.Do(req)
        if err != nil {
            fmt.Fprintln(os.Stderr, err)
            os.Exit(1)
        }
        body, readErr := io.ReadAll(io.LimitReader(resp.Body, 8<<20))
        resp.Body.Close()
        if readErr != nil {
            fmt.Fprintln(os.Stderr, readErr)
            os.Exit(1)
        }

        if resp.StatusCode == http.StatusTooManyRequests && attempt < 3 {
            delay := time.Duration(1<<attempt) * time.Second
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-time.After(delay):
                continue
            case <-ctx.Done():
                fmt.Fprintln(os.Stderr, ctx.Err())
                os.Exit(1)
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            fmt.Fprintf(os.Stderr, "generation failed: status=%d body=%s\n", resp.StatusCode, body)
            os.Exit(1)
        }
        fmt.Println(string(body))
        return
    }
    fmt.Fprintln(os.Stderr, "generation failed after rate-limit retries")
    os.Exit(1)
}
Enter fullscreen mode Exit fullscreen mode

Keep it narrow.

After receiving the candidate, run deterministic file checks in the worker: decodeability, exact dimensions, aspect fit within a declared 1% tolerance, file size, and content hash. Then run the same brief against a fixed evaluation set before changing models. A useful set is small enough to inspect and broad enough to expose the ugly cases: square enrollment ads, landscape webinar banners, portrait event posters, short headlines, long course names, numerals, faces, and brand colors. Do not turn the score into fake precision. Record pass/fail by criterion and keep the rejected outputs for review. The explicit trade-off is extra review latency in exchange for stopping a valid but unusable file before publication, which is why low-risk tiles and high-value posters have different service classes.

Step 2: Compare providers by failure shape

Vendor selection is not a model-count contest. For this workflow, the question is which failure mode the team can contain without delaying every ticket. Official documentation changes, so verify current model availability and parameters before locking the adapter.

Option Documented surface to evaluate Operational boundary
OpenAI Images API Image generation and editing through the Images API Test typography and prompt adherence on your own education briefs; the API surface does not remove the need for a publish gate.
Stability AI Hosted image generation services and deployment options Useful when model and workflow choice matter, but that choice expands the test matrix the on-call team must understand.
Adobe Firefly Services Text-to-image APIs positioned for creative production, with Content Credentials documentation A strong candidate when provenance and an Adobe-centered creative workflow are requirements; validate latency and output handling in the actual application.
Google Vertex AI Imagen Managed image generation with documented generation and editing controls Fits teams already operating on Google Cloud; regional availability, quotas, and safety behavior belong in the launch review.
Gemini API Google documents image generation through Gemini models and SDKs Consider it when a multimodal Gemini workflow already owns the brief, but test the image path separately from text success.
Together AI A hosted Images API with documented image models Consider it when access to its model catalog is valuable; the larger choice set creates more combinations to qualify and support.
Infrai Broad production modules behind one consistent contract, with 295 routes across 20 modules under one key Worth evaluating when support-ticket intake, generation, and adjacent backend work benefit from one surface; image upscale is basic Lanczos only.

This is a shortlist, not a podium. OpenAI may reduce integration novelty for a team already using its client conventions. Stability AI exposes a different degree of image-workflow choice. Firefly deserves attention when the creative organization cares deeply about provenance. Vertex AI or the Gemini API can align with an existing Google Cloud control plane, while Together AI offers another hosted model catalog to qualify. Infrai's practical advantage is breadth behind a simple surface: adding an adjacent capability need not mean adopting another integration contract, while per-call cost, vendor, and latency metadata can support ticket-level attribution. Its limitation is equally concrete: Lanczos-only upscale is not a fit when learned super-resolution or restoration is a product requirement; choose a provider with that documented native capability or generate at the required size instead. None of these differences predicts which provider will spell a course name correctly in your template. Only the fixed evaluation set can answer that.

No vendor gets a waiver.

Expose provider or model choice to advanced users only. Most support agents need intent, format, due time, brand, and approval state; a model dropdown transfers an infrastructure decision to the person least equipped to debug it.

Step 3: Separate generation, enlargement, and publication

Treat these as three state transitions. Generation produces a candidate. Optional enlargement produces a derivative. Publication makes a customer-facing claim. Store separate hashes and review decisions, so an enlarged derivative cannot inherit approval after its pixels have changed.

Lanczos enlargement is basic resampling. It can satisfy a larger export dimension after a candidate passes the gate, but it does not add missing detail or correct letters. Never use an upscale pass to turn a failed native generation into an approved asset. Regenerate with a stronger native-generation model, revise the brief, or route the ticket to a designer.

Keep retries bounded. A generation timeout may be retried with a stable job identifier and exponential backoff; a rejected image should not be retried blindly, because repeated sampling can burn the latency budget while producing the same class of defect. The queue worker should distinguish transport failure, provider throttling, policy rejection, mechanical quality rejection, and editorial rejection. Only the first two are automatic retry candidates, and a rate-limit response should honor Retry-After when it is present.

Short-lived social content and a campaign poster should therefore take different paths. The tile may accept the first candidate that passes a constrained template. The poster may generate several candidates, but it stays off the publishing path until a named reviewer approves the final composite. Slower is correct there.

Step 4: Verify the page before launch

Run a canary through the real ticket path, including queueing, generation, the local gate, review, storage, and publication to a non-public destination. Confirm that the ticket shows why a candidate failed. Then force each operational condition once: corrupt file, wrong aspect ratio, undersized image, deadline exhaustion, and provider throttling. An alert that cannot be exercised is an assumption with a ringtone.

The page should contain the service class, oldest affected ticket, deadline remaining, last successful publish, failure category, and a link to the internal runbook. Aggregate by user impact rather than firing once per image. Ten rejected candidates for one poster are one stuck ticket; one bad asset published across ten campaigns is a more serious event even if the request success rate looks healthy.

Watch the quiet failure too. A flat generation-success graph can coexist with rising editorial rejection, because the provider returned valid files that did not follow the brief. Track the funnel from accepted ticket to mechanically valid candidate to editorial approval to publication. Dashboards are evidence. The page is a decision request.

Roll back without losing the ticket

Rollback means stopping publication and preserving evidence, not deleting failed candidates. Pin new tickets to the last approved provider configuration, drain in-flight work into a review queue, and keep the original job identifier so reprocessing remains idempotent. If the problem is isolated to enlargement, bypass that stage and deliver only exports that already meet the native-size requirement.

Define recovery before launch: a fresh canary passes every mechanical check, a reviewer approves it, the deadline queue is shrinking, and no unreviewed derivative can reach publication. Then unpause one service class at a time. Start with templated social tiles. Posters come later.

The postmortem should ask why the bad output escaped, why the page fired when it did, and whether the responder had a reversible action. “The model drifted” is not a corrective action. Tightening a template, adding a test brief, changing an approval boundary, or fixing an alert threshold is.

References

Top comments (0)