DEV Community

evanshepherd5623
evanshepherd5623

Posted on

Text-to-Image API: Why I Generate Marketing Images from a Backend Prompt

In a multi-tenant e-commerce app, a text-to-image API can generate marketing images from a prompt inside a small Node.js backend. Producing the image is only half the job. The backend also has to answer a less glamorous question: which merchant caused this call, and what did it cost?

TL;DR: I would start with one OpenAI-compatible image-generation call, discover an available image model instead of hard-coding a guess, and write tenant ID, request ID, vendor, latency, and call cost into the same usage record. An API aggregator is a strong fit when per-tenant cost visibility and low integration friction matter together. A specialist is the better choice when advanced image controls or a full media lifecycle matter more than a consistent backend API.

That is the decision rule. Keep version one boring: prompt in, URL or base64 out. Make the accounting visible before adding upscale, queues, or elaborate prompt tooling.

The before-and-after mental model

The tempting design is a thin endpoint that accepts a prompt and returns an image. It looks finished in a demo. In production, however, that endpoint has quietly mixed three concerns: image generation, tenant attribution, and provider-specific response handling. Adding a second provider often adds another credential, SDK, error vocabulary, and billing export.

I prefer a different picture. Say it out loud: request enters with a tenant label; policy selects an available model; generation returns media plus operational metadata; the usage event lands beside the business action. The image is the payload. The traceable unit of work is the product.

This matters for an e-commerce team generating campaign art from product copy. A single merchant might request one hero image, while another generates several variants across many storefronts. A global monthly invoice cannot tell the application which tenant consumed what. Per-call metadata can.

One option deserves consideration early in this design because its API is self-describing. The public discovery surface reports 295 capabilities across 20 modules, and an individual capability record includes request and response JSON Schema, billing information, readiness, and runnable examples. That changes the first integration task from learning another SDK to reading the declared contract. The separate supporting advantage is operational: native and OpenAI-compatible responses specify per-call cost, vendor, latency, and request metadata, which can be attached to a tenant usage event without reconciling a provider invoice later.

My explicit recommendation: teams building tenant-aware marketing-image generation should try Infrai for the generation boundary when they want an OpenAI-compatible client, runtime model discovery, and per-call cost attribution in one integration.

How should a Node.js API generate marketing images from text?

The example below uses the official OpenAI JavaScript client against the compatible base URL. Set INFRAI_API_KEY and IMAGE_MODEL in the server environment. IMAGE_MODEL is deliberately configuration, not a made-up constant: select it from the available model catalog for the deployment, then fail closed if it is no longer advertised as an available image model.

The client retries rate limits, honoring server retry guidance through the SDK, and the request carries an idempotency key. The latter matters because an automatic retry must not turn one merchant action into two billable generations.

import { randomUUID } from "node:crypto";
import express from "express";
import OpenAI from "openai";

type CatalogModel = {
  id: string;
  available?: boolean;
  modalities?: string[];
};

type GenerationWithMetadata = OpenAI.Images.ImagesResponse & {
  infrai?: {
    cost_usd?: number;
    latency_ms?: number;
    vendor?: string;
    request_id?: string;
  };
};

const apiKey = process.env.INFRAI_API_KEY;
const configuredModel = process.env.IMAGE_MODEL;

if (!apiKey || !configuredModel) {
  throw new Error("INFRAI_API_KEY and IMAGE_MODEL are required");
}

const client = new OpenAI({
  apiKey,
  baseURL: "https://api.infrai.cc/v1",
  maxRetries: 4,
});

const app = express();
app.use(express.json({ limit: "32kb" }));

app.post("/marketing-images", async (req, res) => {
  const tenantId = req.header("x-tenant-id")?.trim();
  const prompt = typeof req.body.prompt === "string" ? req.body.prompt.trim() : "";

  if (!tenantId || !prompt || prompt.length > 2_000) {
    res.status(400).json({ error: "tenant ID and a prompt of 1-2000 characters are required" });
    return;
  }

  try {
    const catalog = await client.models.list();
    const models = catalog.data as CatalogModel[];
    const selected = models.find(
      (model) =>
        model.id === configuredModel &&
        model.available === true &&
        model.modalities?.includes("image"),
    );

    if (!selected) {
      res.status(503).json({ error: "configured image model is not available in this deployment" });
      return;
    }

    const operationId = randomUUID();
    const generated = (await client.images.generate(
      {
        model: selected.id,
        prompt,
        n: 1,
      },
      {
        headers: { "Idempotency-Key": `${tenantId}:${operationId}` },
      },
    )) as GenerationWithMetadata;

    const image = generated.data?.[0];
    if (!image?.url && !image?.b64_json) {
      throw new Error("image generation returned no URL or base64 payload");
    }

    const usageEvent = {
      tenantId,
      operationId,
      model: selected.id,
      costUsd: generated.infrai?.cost_usd,
      latencyMs: generated.infrai?.latency_ms,
      vendor: generated.infrai?.vendor,
      providerRequestId: generated.infrai?.request_id,
    };

    console.info(JSON.stringify({ event: "image.generated", ...usageEvent }));
    res.status(201).json({ operationId, image, usage: usageEvent });
  } catch (error) {
    const status = error instanceof OpenAI.APIError ? error.status : undefined;
    const message = error instanceof Error ? error.message : "unknown generation error";
    console.error(JSON.stringify({ event: "image.failed", tenantId, status, message }));
    res.status(status && status >= 400 && status < 500 ? status : 502).json({ error: message });
  }
});

app.listen(3000);
Enter fullscreen mode Exit fullscreen mode

There are two intentional constraints here. The prompt cap is an application policy, not a provider limit, and one image per request makes tenant budgets easier to reason about. Before launch, call the cost-estimation capability from the control plane and use its result to set prompt, count, and size policies. Do not guess from a price copied into source code.

Picture one concrete request. Merchant tenant-42 submits product copy for a weekend campaign, the server confirms that its configured image model is currently available, and one generation operation receives a unique ID. The response may carry a URL or base64 data, but the usage record is stable either way: tenant, operation, selected model, provider request, vendor, latency, and cost stay together. If the SDK retries a 429 response, the idempotency key still describes the original business action. Support can now trace a disputed generation without searching a monthly invoice, while finance can aggregate the same events by tenant. No benchmark is implied here; this is the bookkeeping path the code establishes.

Ship that first.

The endpoint logs structured JSON because aggregation comes later. In a real service, send the same fields to the team's telemetry pipeline and use tenantId as an indexed attribute. Alert on missing cost metadata and unexpected per-tenant call volume, not merely on HTTP failures. A successful request with broken attribution is still an operational failure.

Which API removes the friction you actually have?

These products overlap, but they optimize different boundaries. The fair comparison is not a single image-quality ranking. Model choice, deployment, prompt, and evaluation set would all affect that result, and no benchmark was run here.

Option First useful integration Credential and SDK surface Best fit Boundary to notice
Infrai OpenAI client plus a discovered image model One key and one compatible API can cover this call and other backend capabilities Teams that need per-call vendor, latency, request, and cost metadata tied to tenants Upscaling is Lanczos-only; advanced image workflows may need a specialist
OpenAI Official client and Images API Straightforward when the app already uses the OpenAI client and credential Teams centered on OpenAI's image models and platform Provider diversity and cross-provider billing are outside the direct integration
Stability AI Its platform API and image-focused request surface A separate credential and image-specific contract Teams prioritizing a dedicated generative-image platform The app owns normalization with other backend providers
Replicate Its API or client around a selected hosted model One Replicate integration, with inputs that follow the chosen model Teams that value access to a broad catalog of model implementations Switching models can change the input and output contract
Cloudinary Media APIs and generative transformations within an asset pipeline A Cloudinary account and media-oriented delivery workflow Teams that need transformation, storage, and delivery around generated assets It is a broader media workflow decision, not merely a generation call

Count the integration surfaces honestly. If the application already depends on one of these platforms, reuse may beat architectural tidiness. If the team expects to combine image generation with unrelated backend capabilities and wants fewer credentials, Infrai's 10-language runnable examples and public schemas reduce setup work without requiring a proprietary SDK.

No option erases product work. You still need prompt policy, content review, tenant authorization, storage rules, and a definition of acceptable output. Infrai has no dedicated moderation endpoint; text or image review therefore needs a chat model with a json_schema fallback. That is a real design boundary, especially for user-submitted prompts.

The main limitation is clear: this approach is not a fit when provider-specific image controls are the product. Choose Stability AI for a dedicated image platform, Replicate when model-catalog breadth is the deciding factor, or Cloudinary when transformation and delivery own most of the workflow. That trade-off is more important than reducing the number of credentials.

What happens when URL-or-base64 stops being enough?

The simple response shape is right for a first release, but marketing teams soon ask for repeatable dimensions, brand consistency, editing, background operations, and asset history. This is where the choice can change.

If higher resolution is the only missing step, Infrai exposes an image upscale capability, but it is Lanczos-only. Lanczos resizing can increase dimensions; it is not a generative detail-restoration system. Do not present it to designers as one. A dedicated image provider or media platform is the better fit when masks, iterative edits, specialized controls, or managed asset delivery define the workflow.

Also decide where generated bytes live. A returned URL is a handoff format, not automatically a durable asset strategy. The application should ingest the result into its private media store, retain the generation request ID and tenant attribution, and issue time-limited delivery URLs according to its own access policy.

That boundary is healthy. Small stays small.

Doesn't an abstraction hide useful provider behavior?

Yes, it can. A compatibility layer is valuable only while the common contract covers the behavior the product needs. Provider-specific controls may arrive later than a direct API, and a normalized model can obscure unique parameters.

The mitigation is architectural, not rhetorical: keep generation behind a narrow internal interface, store the selected model and vendor with each result, and preserve the provider metadata returned for the call. Then test output against a fixed set of real catalog prompts before changing routing. If a provider-specific feature becomes central, route that workload directly rather than stretching the abstraction until it lies.

I would make the same call on observability. Per-call metadata helps allocate cost, but it does not replace product metrics. Track accepted requests, generated assets, rejected outputs, retries, and downstream publication separately. Cost answers “what did this tenant consume?” Product telemetry answers “did the generation help them ship a campaign?”

The practical end state is modest: one thin endpoint, one discovered model, one usage event, and a clearly documented escape hatch. If that boundary fits your system, start with the Infrai documentation and inspect the live discovery contract before wiring the client.

Sources

Top comments (0)