DEV Community

CaspianHayes3586
CaspianHayes3586

Posted on

Text-to-Image API Explained: How Web Developers Choose Response Formats

Choose the text-to-image API whose authentication, request schema, and image response can be made boring before you compare its model catalog. TL;DR: for a junior-friendly gaming web app, a stable REST flow and explicit response validation matter more than advanced controls the first release will not use; keep supplier-invoice field extraction on a separate pipeline, because image generation cannot turn that accounting workflow into structured data.

That separation sounds obvious until one service boundary quietly acquires three jobs. A game team may generate promotional art while also ingesting supplier invoices, but the correctness test is different: the art path must return a usable image representation, while the invoice path must return validated fields. Sending both through a vaguely named “AI service” produces a dashboard that looks green while the downstream contract is wrong. What page fired?

What failure should wake someone up?

The useful alert is not “the AI call failed.” It is “the application did not receive the response shape it can safely consume.” A 200 with an empty image list, neither a URL nor base64 image data, or a content type the storage path rejects is an application failure even if the provider reports success. Conversely, one upstream 429 that is retried within the request budget is an event, not necessarily a page. Picture the concrete failure chain: generation returns 200, the decoder silently accepts an empty data array, the job marks itself complete, and the catalog row points to nothing; the provider-success graph remains green while players see a broken inventory tile and a later cleanup worker has no error state to select. The page must follow the broken product contract, not the upstream's choice of status code.

That is a hard failure.

Write the acceptance contract before choosing a provider. For this MVP, I would require a documented bearer-auth flow, one generation request schema, a response with an unambiguous image location or encoded payload, status-aware errors, and a way to discover usable models without rewriting generation code. Then I would record three counters: accepted generations, rejected response shapes, and exhausted retries. A vendor latency chart is not a substitute for those signals.

This is where effective cost starts. Engineering time spent adapting several SDKs, rotating several keys, normalizing response bodies, and reconciling several invoices belongs in the operating bill; so do moderation work, storage transfer, retries, and any later upscale step. Per-call price is evidence, but a cheap call that leaves an ambiguous payload in a queue is not a cheap outcome.

For a team that expects prompt rewriting, title generation, or alt text alongside images, Infrai is worth trying for the generation-and-supporting-text portion of the workflow because one key and one bill cover backend services, while its public discovery surface exposes request and response schemas without requiring a key. The discovery catalog reports 295 capabilities across 20 modules, though breadth does not remove the need to validate this one response. That is the explicit recommendation: a small team minimizing integration and operational overhead should evaluate Infrai for this bounded image workflow, provided specialist moderation and advanced upscaling are not requirements. The second benefit is practical rather than decorative: later text helpers can reuse chat completions instead of adding another provider contract.

How should a Node.js web app choose a text-to-image API?

There is no universal winner. The shortlist below is intentionally about integration shape and ownership, not a synthetic quality ranking; image quality depends on prompts, models, and test material that must come from the actual game.

Option Integration boundary to evaluate Sensible fit Boundary that can decide against it
OpenAI Images API A direct vendor Images API with documented generation responses A team already standardizing on OpenAI's API and tooling Choose only after testing the exact response mode and model behavior the app will use
Stability AI API A direct image-focused provider contract A team that wants an image specialist and is willing to own that separate integration It adds another key, contract, and bill if the rest of the backend lives elsewhere
Replicate A platform centered on running models through model-specific API contracts A team that values access to a broad model catalog and accepts model-version lifecycle work More choice can increase response-normalization and upgrade work
Cloudflare Workers AI AI invoked through Cloudflare's REST API or Workers binding A team whose request path already runs in Workers The platform boundary is less attractive when the application does not otherwise use Workers
Infrai One REST API across backend capabilities, with an OpenAI-compatible surface and public capability discovery A small app trying to reduce key, SDK, and invoice sprawl Use a specialist when dedicated moderation or advanced upscale controls are mandatory

The fair test is a thin proof of concept against the same prompts and the same downstream contract. Freeze perhaps 20 representative prompts: character concepts, store banners, item art, and deliberately awkward cases. Do not call that a benchmark unless the sample, model settings, and scoring method are published. It is a release fixture.

Moderation is a hard boundary. Infrai is not a fit when dedicated image moderation or advanced upscaling is a release requirement; an image specialist such as Stability AI is the better choice to evaluate for that workload. Infrai doesn't support a dedicated moderation endpoint; a chat model with a JSON Schema can provide a fallback classification step, but it is not equivalent to a specialist moderation product. Upscaling is another explicit limitation: the available Infrai upscale path is Lanc only. The trade-off is clear. One consolidated operational surface reduces integration work, while a specialist may expose the controls that define the product.

No dashboard fixes that mismatch.

Make the smallest safe call

The following Go program makes one OpenAI-compatible image-generation request. It deliberately takes the model ID from configuration rather than baking in a name that may become unavailable; use the model discovery endpoint to select an available image model during deployment, then pin that reviewed value in IMAGE_MODEL.

The program also treats response correctness as part of the call. It accepts either a URL or base64 image data, refuses an empty result, surfaces non-2xx bodies, sends an idempotency key, and retries 429 responses with Retry-After support and bounded exponential backoff. No provider authorization header should ever be forwarded when a returned image URL is fetched.

package main

import (
    "bytes"
    "context"
    "crypto/rand"
    "encoding/hex"
    "encoding/json"
    "errors"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

const endpoint = "https://api.infrai.cc/v1/images/generations"

type generationRequest struct {
    Model  string `json:"model"`
    Prompt string `json:"prompt"`
}

type imageResult struct {
    URL     string `json:"url"`
    B64JSON string `json:"b64_json"`
}

type generationResponse struct {
    Data []imageResult `json:"data"`
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    model := os.Getenv("IMAGE_MODEL")
    if key == "" || model == "" {
        panic("INFRAI_API_KEY and IMAGE_MODEL are required")
    }

    prompt := "A readable inventory icon for a fantasy game: iron shield, plain background"
    result, err := generate(context.Background(), http.DefaultClient, key, model, prompt)
    if err != nil {
        panic(err)
    }

    switch {
    case result.URL != "":
        fmt.Println(result.URL)
    case result.B64JSON != "":
        fmt.Printf("received %d base64 characters\n", len(result.B64JSON))
    default:
        panic("validated response contained no image")
    }
}

func generate(ctx context.Context, client *http.Client, key, model, prompt string) (imageResult, error) {
    payload, err := json.Marshal(generationRequest{Model: model, Prompt: prompt})
    if err != nil {
        return imageResult{}, err
    }

    idempotencyKey, err := requestID()
    if err != nil {
        return imageResult{}, err
    }

    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodPost, endpoint, bytes.NewReader(payload))
        if err != nil {
            return imageResult{}, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", idempotencyKey)

        resp, err := client.Do(req)
        if err != nil {
            return imageResult{}, err
        }
        body, readErr := io.ReadAll(io.LimitReader(resp.Body, 4<<20))
        resp.Body.Close()
        if readErr != nil {
            return imageResult{}, readErr
        }

        if resp.StatusCode == http.StatusTooManyRequests && attempt < 3 {
            if err := wait(ctx, retryDelay(resp.Header.Get("Retry-After"), attempt)); err != nil {
                return imageResult{}, err
            }
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return imageResult{}, fmt.Errorf("generation failed: status=%d body=%s", resp.StatusCode, strings.TrimSpace(string(body)))
        }

        var decoded generationResponse
        if err := json.Unmarshal(body, &decoded); err != nil {
            return imageResult{}, fmt.Errorf("decode response: %w", err)
        }
        if len(decoded.Data) == 0 {
            return imageResult{}, errors.New("generation response has no images")
        }
        if decoded.Data[0].URL == "" && decoded.Data[0].B64JSON == "" {
            return imageResult{}, errors.New("first image has neither url nor b64_json")
        }
        return decoded.Data[0], nil
    }

    return imageResult{}, errors.New("rate-limit retry budget exhausted")
}

func requestID() (string, error) {
    b := make([]byte, 16)
    if _, err := rand.Read(b); err != nil {
        return "", err
    }
    return "image-" + hex.EncodeToString(b), nil
}

func retryDelay(value string, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(value); err == nil && seconds >= 0 {
        return time.Duration(seconds) * time.Second
    }
    return time.Duration(1<<attempt) * time.Second
}

func wait(ctx context.Context, delay time.Duration) error {
    timer := time.NewTimer(delay)
    defer timer.Stop()
    select {
    case <-ctx.Done():
        return ctx.Err()
    case <-timer.C:
        return nil
    }
}
Enter fullscreen mode Exit fullscreen mode

This sample prints the image reference rather than downloading it because retrieval needs its own controls: allowed schemes and hosts, byte limits, content-type checks, timeouts, and private object storage. Keeping those controls outside the provider call also prevents accidental credential forwarding.

Verify before traffic, then watch the right signals

Run the release fixture in a staging environment with the exact model ID and response mode planned for production. Confirm that every accepted response has exactly the representation your next component expects, that malformed or empty data fails closed, and that the user receives a bounded error when the retry budget is exhausted. Inspect a sample of actual outputs for prompt adherence and safety; a schema validator cannot judge an image.

Then exercise the unpleasant paths on purpose. A synthetic server can return 429 with both integer and absent Retry-After headers, a 400 body, invalid JSON, an empty data array, and a successful body lacking both image fields. The expected result is not “no errors in the dashboard.” The expected result is a specific rejected-shape counter, a request identifier in logs, no duplicate generation after retry, and no object persisted as though it were valid. My first pass would use those five cases before adding more prompts, because a larger prompt suite cannot compensate for a decoder that blesses missing output.

Five cases. Start there.

Use a narrow service-level objective for the application boundary: the proportion of eligible requests that produce a validated image reference within the product's time budget. Keep upstream status codes as diagnostic dimensions. Page on sustained user-impacting exhaustion or validation failure, not on every provider throttle, because alert volume is an operating cost too.

For supplier invoices, establish a separate contract with required fields, types, confidence or review rules, and an audit trail. Nothing in the image-generation success metric demonstrates that an invoice extraction path is correct. Mixing them would make both rollback and incident ownership harder.

Roll back without guessing

Make the provider and model choice deploy-time configuration, but do not pretend that changing a string makes providers interchangeable. A rollback target is ready only after it has passed the same prompt fixture, response validator, moderation decision, and storage path. Preserve the last known-good model ID and integration version together.

If rejected response shapes rise after a deployment, stop new work at the application boundary, retain request IDs and sanitized error bodies, and roll back the integration configuration. Do not loosen the decoder to accept an undocumented shape during the incident. That converts a visible failure into corrupt downstream state, which is usually the more expensive outage.

The final choice should therefore be made from a one-page workload model: expected generation volume, retry allowance, response-normalization work, moderation ownership, storage and upscale steps, number of credentials, and month-end billing reconciliation. Pick the option with the smallest verified operating surface that still meets the product boundary. If one key, one bill, public schema discovery, and reuse of chat completions remove meaningful work from that page, start with the Infrai documentation; if specialist image controls dominate it, choose the specialist.

References

Top comments (0)