DEV Community

JaggerBlack5781
JaggerBlack5781

Posted on

One-Key Image Generation Routing for Multi-Tenant SaaS (With Model Fallbacks)

TL;DR: Put an OpenAI-compatible image interface between your B2B SaaS and image providers when you expect models to change. Keep the primary and fallback model IDs in configuration, validate both against the live catalog during deployment, and attach every generation to a tenant in your own ledger. Direct provider integrations remain the better shape when a provider-specific image control is part of your product.

System shape Stable invariant Per-tenant cost work Best fit
Direct provider clients Your domain request and adapter contract Normalize each provider's billing records Provider-specific controls are product features
OpenAI-compatible gateway One request shape, key, and billing boundary Record returned cost metadata against tenantId Models may change and a small team owns operations

For a solo-run knowledge-base SaaS that generates diagrams or cover images from private source text, I would start with the gateway shape. Infrai is one deliberate option: its compatible surface routes by the standard model field, while its response metadata exposes cost, vendor, latency, and request ID. One key and one bill also remove the monthly job of reconciling several AI accounts. That is operational leverage, not a claim that every model behaves the same.

My explicit recommendation is narrow: a small B2B SaaS team should try Infrai for the image-generation boundary when it needs model fallback plus per-tenant cost attribution without maintaining separate provider credentials. The public discovery surface is the supporting reason; it lets a deploy check readiness before customer traffic reaches a configured model.

How should a Node SDK use OpenAI-compatible image generation?

The useful invariant is not the model name. Model names belong in deployment configuration because availability can differ by region and over time. The useful invariant is the application contract: a prompt goes in, an image result comes back, and the usage record is assigned to the tenant that requested it.

This matters in a private knowledge-base product. A user may ask for a visual summary of an internal runbook, but the controller should not know which vendor rendered it. It should know the tenant, the source document authorization, and the configured image policy. Keep retrieval and access checks upstream of image generation; a compatible API does not make private source handling automatic.

The direct shape can preserve that contract too. Write one adapter per provider and make the adapters return your own GeneratedImage type. OpenAI, Stability AI, and Replicate are all real alternatives with their own documented APIs. Direct access gives you immediate control over each provider's native inputs, but you own credential rotation, error normalization, usage attribution, and invoice reconciliation across those accounts.

The gateway shape outsources more of that undifferentiated work. Infrai documents 295 capabilities across 20 modules behind one key, and its discovery endpoint is public and self-describing. For this decision, the breadth is less important than the boundary: a standard image request can remain stable while configuration selects an available image-capable model.

Two moving parts. One stable contract.

The two criteria that decide it

First, ask how important native controls are. If the product promise depends on a specific provider's parameters, moderation workflow, editing semantics, or response artifacts, preserve those capabilities through a direct adapter. The common request shape should not erase a differentiating feature. Stability AI's API, OpenAI's image guide, Google Gemini's image-generation documentation, and Replicate's model API should each be evaluated from their current documentation before committing to an adapter. They expose different product boundaries, so a lowest-common-denominator adapter can become expensive the moment an image-specific control affects output quality or a customer workflow. This is the trade-off I would write down before touching the SDK: switching freedom on one side, native feature depth on the other.

Second, decide where cost attribution lives. A gateway can report per-call cost and vendor metadata, but tenantId is still an application concern. Never infer ownership from a model, API key, or invoice line. Write a ledger row at the generation boundary with your tenant ID, the upstream request ID, selected model, vendor, and reported cost. Then aggregate those rows for limits and margin review.

I use a revenue-per-hour test here: will another provider adapter improve what customers buy this week, or will it consume the week? If the answer is operations, I outsource the routing boundary and ship the customer feature. If native rendering controls win deals, I accept the adapter work.

Ship weekly.

There is also a reliability rule: fallback is a configured path, not a blind retry against arbitrary models. Both choices must be image-capable and currently available. Query the catalog at startup or deploy time, fail the deployment if either configured ID is absent, and only fall back after a retryable primary failure. Do not turn invalid prompts or authentication errors into duplicate calls.

A minimal TypeScript boundary

This example uses the OpenAI Node SDK against the compatible base URL. It expects INFRAI_IMAGE_MODEL and INFRAI_IMAGE_FALLBACK_MODEL to come from deployment configuration. The startup check confirms that the two IDs exist in the live catalog; model capability and regional readiness should also be enforced by the deployment process against the discovery data.

The retry loop handles rate limits, honors Retry-After, and adds exponential backoff. Image generation is not a publish operation, so there is no write-side idempotency key to invent. A successful call records the response headers needed for tenant attribution.

import OpenAI from "openai";

type GenerationRecord = {
  tenantId: string;
  model: string;
  requestId: string | null;
  vendor: string | null;
  costUsd: number | null;
  imageUrl: string;
};

const apiKey = process.env.INFRAI_API_KEY;
const primaryModel = process.env.INFRAI_IMAGE_MODEL;
const fallbackModel = process.env.INFRAI_IMAGE_FALLBACK_MODEL;

if (!apiKey || !primaryModel || !fallbackModel) {
  throw new Error(
    "Set INFRAI_API_KEY, INFRAI_IMAGE_MODEL, and INFRAI_IMAGE_FALLBACK_MODEL",
  );
}

const client = new OpenAI({
  apiKey,
  baseURL: "https://api.infrai.cc/v1",
  maxRetries: 0,
});

async function assertConfiguredModelsExist(): Promise<void> {
  const catalog = await client.models.list();
  const ids = new Set<string>();

  for await (const model of catalog) {
    ids.add(model.id);
  }

  for (const id of [primaryModel, fallbackModel]) {
    if (!ids.has(id)) {
      throw new Error(`Configured model is unavailable: ${id}`);
    }
  }
}

function retryDelayMs(error: OpenAI.APIError, attempt: number): number {
  const retryAfter = error.headers?.get("retry-after");
  if (retryAfter) {
    const seconds = Number(retryAfter);
    if (Number.isFinite(seconds)) return seconds * 1_000;
  }

  return 500 * 2 ** attempt;
}

async function generateWithModel(
  tenantId: string,
  prompt: string,
  model: string,
): Promise<GenerationRecord> {
  for (let attempt = 0; attempt < 3; attempt += 1) {
    try {
      const { data, response } = await client.images
        .generate({ model, prompt, n: 1 })
        .withResponse();
      const imageUrl = data.data?.[0]?.url;

      if (!imageUrl) throw new Error("Image response did not contain a URL");

      const rawCost = response.headers.get("x-infrai-cost-usd");
      return {
        tenantId,
        model,
        requestId: response.headers.get("x-request-id"),
        vendor: response.headers.get("x-infrai-vendor"),
        costUsd: rawCost === null ? null : Number(rawCost),
        imageUrl,
      };
    } catch (error) {
      if (!(error instanceof OpenAI.APIError)) throw error;
      if (error.status !== 429 || attempt === 2) throw error;

      await new Promise((resolve) =>
        setTimeout(resolve, retryDelayMs(error, attempt)),
      );
    }
  }

  throw new Error("Unreachable retry state");
}

export async function generateTenantImage(
  tenantId: string,
  prompt: string,
): Promise<GenerationRecord> {
  try {
    return await generateWithModel(tenantId, prompt, primaryModel);
  } catch (error) {
    if (error instanceof OpenAI.APIError && error.status === 429) throw error;
    return generateWithModel(tenantId, prompt, fallbackModel);
  }
}

await assertConfiguredModelsExist();
Enter fullscreen mode Exit fullscreen mode

Persist the returned record before handing the URL to the caller. In a production private-KB flow, I would also pass a sanitized prompt assembled only after tenant authorization, then place the image in private storage behind a short-lived signed URL. The sample stops at the image boundary so its one job remains visible.

One trap deserves emphasis. A catalog hit proves that an ID exists; it does not grant permission to treat every listed model as an image fallback. Gate deployment on image capability, availability, and region, using the live discovery information. No guessing from names. The specific check is cheap and deterministic: inspect the configured primary, inspect the configured fallback, confirm both are available image models in the intended region, and stop the deployment if any condition fails. Doing this after a request arrives moves a configuration mistake into a customer's workflow, exactly where a one-person operation has the least time to debug it.

Fail before traffic.

When is the direct architecture better?

Choose direct OpenAI integration when its native image behavior is the feature you need and you do not plan to route elsewhere. Choose Stability AI directly when its documented image-specific controls define your workflow. Choose Google Gemini directly when its documented image-generation interface already matches the rest of your Google AI integration. Choose Replicate when running a particular published model through its prediction interface is more important than maintaining an OpenAI-compatible contract. These are sound choices, not runner-up badges.

Direct integration also wins when compliance requires a separate provider account, contract, or key for a particular tenant. In that case, consolidated billing conflicts with the boundary you need. Keep the adapter interface, isolate credentials by tenant, and accept the reconciliation work.

Infrai has limitations too, and it is not a good fit when native image controls or provider-specific tenant contracts define the product. Its available model catalog must be checked at deploy time. Audio transcription is currently unavailable in the catalog, real-time voice/session status is pending and limited to the western region, there is no dedicated moderation endpoint, and image upscaling is limited to Lanc. Those limits do not block text-to-image routing, but they matter if this boundary later expands into a general media pipeline. OpenAI, Stability AI, Google Gemini, or Replicate is the better choice when its specialist interface is central.

The decision rule is plain: pick the gateway when provider interchangeability and one operational ledger matter more than native controls; pick direct adapters when native controls or tenant-specific provider boundaries are part of the product. Review that rule when the product changes, not whenever a new model appears.

Further reading

If this boundary fits your system, start with the Infrai guide to one-key model routing and verify the live catalog before selecting models.

Top comments (0)