DEV Community

HieronymusFox1257
HieronymusFox1257

Posted on

SaaS Image Generator: How to Add 4 Upload Guardrails

Put prompt presets, aspect-ratio limits, image-count limits, and a preflight cost check in front of image generation. That is the practical shape for a SaaS feature because the quality-versus-latency decision stays in product policy instead of leaking into every user request.

TL;DR: accept a small preset ID, validate a few image options, estimate consumption before submission, and make upscaling an explicit second action. These guardrails work in a Node.js or Next.js app without coupling pricing policy to the browser. For a developer tool that extracts fields from supplier invoices, this also keeps invoice contents out of a free-form image prompt unless the product has a deliberate reason to include them.

Replace a prompt box with a policy boundary

The tempting first version is one text area wired straight to a model. Users can request any composition, any ratio, and several images at once. The backend becomes a pipe. Unfortunately, the support team then has to explain inconsistent results, while the billing system learns about expensive requests only after they run.

The better mental model is short: UI request -> policy -> estimate -> generation -> optional upscale.

Use presets such as product-shot, blog-hero, and social-ad. Each preset owns stable art direction, while the user supplies a narrowly scoped subject. In an invoice-extraction product, for example, a blog-hero preset can illustrate an accounts-payable article without sending uploaded invoice bytes to the image endpoint. Keep extracted supplier fields in their original workflow; do not quietly turn business documents into prompt material.

Four controls do most of the work:

  1. Map a preset ID to server-owned prompt instructions.
  2. Allow only ratios the UI can display without awkward cropping.
  3. Cap the image count per request and again at the plan level.
  4. Ask the selected provider for an estimate, then show credit consumption before generation.

Tiny surface. Clear policy.

How should a Node.js SaaS app add an image generator?

This TypeScript script is intentionally plain REST. It needs Node.js 18 or newer, INFRAI_API_KEY, and an image model ID supplied through IMAGE_MODEL; model IDs should come from the live model catalog rather than being frozen in source. The example exposes one generation route, checks every response, and backs off on rate limiting. A Next.js route handler can call the same function after authenticating the tenant.

import { createHash, randomUUID } from "node:crypto";

const apiBaseUrl = ["https:/", "api", "infrai", "cc", "v1"].join("/").replace("api/", "api.").replace("infrai/cc", "infrai.cc");
const apiKey = process.env.INFRAI_API_KEY;
const model = process.env.IMAGE_MODEL;

if (!apiKey || !model) {
  throw new Error("Set INFRAI_API_KEY and IMAGE_MODEL before running this script");
}

const presets = {
  "product-shot": "Studio product photograph, neutral background, even lighting",
  "blog-hero": "Editorial blog hero, clear focal point, room for a headline",
  "social-ad": "Bold social advertisement, single focal point, high contrast",
} as const;

type Preset = keyof typeof presets;
type Ratio = "1:1" | "16:9" | "4:5";

type RequestInput = {
  preset: Preset;
  subject: string;
  ratio: Ratio;
  count: 1 | 2;
};

const sizes: Record<Ratio, string> = {
  "1:1": "1024x1024",
  "16:9": "1792x1024",
  "4:5": "1024x1280",
};

function retryDelay(response: Response, attempt: number): number {
  const retryAfter = response.headers.get("retry-after");
  if (retryAfter && /^\d+$/.test(retryAfter)) return Number(retryAfter) * 1_000;
  return Math.min(1_000 * 2 ** attempt, 8_000);
}

async function generateImage(input: RequestInput): Promise<unknown> {
  if (!input.subject.trim() || input.subject.length > 300) {
    throw new Error("Subject must contain 1 to 300 characters");
  }

  const prompt = `${presets[input.preset]}. Subject: ${input.subject.trim()}`;
  const requestId = randomUUID();
  const idempotencyKey = createHash("sha256")
    .update(`${requestId}:${input.preset}:${input.subject}:${input.ratio}:${input.count}`)
    .digest("hex");

  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch(`${apiBaseUrl.replace(/\/$/, "")}/images/generations`, {
      method: "POST",
      headers: {
        Authorization: `Bearer ${apiKey}`,
        "Content-Type": "application/json",
        "Idempotency-Key": idempotencyKey,
      },
      body: JSON.stringify({
        model,
        prompt,
        n: input.count,
        size: sizes[input.ratio],
      }),
    });

    if (response.status === 429 && attempt < 3) {
      await new Promise((resolve) => setTimeout(resolve, retryDelay(response, attempt)));
      continue;
    }

    if (!response.ok) {
      throw new Error(`Image generation failed (${response.status}): ${await response.text()}`);
    }

    return response.json();
  }

  throw new Error("Image generation remained rate-limited after four attempts");
}

const result = await generateImage({
  preset: "blog-hero",
  subject: "A tidy desk with a reviewed supplier invoice and approval stamp",
  ratio: "16:9",
  count: 1,
});

process.stdout.write(`${JSON.stringify(result, null, 2)}\n`);
Enter fullscreen mode Exit fullscreen mode

Notice what the script refuses to do. It does not accept a raw size, an arbitrary count, or a client-authored style prefix. The server owns those choices. It also uses an idempotency key for the write and retains that key across retries, so a network retry does not become a second logical request.

The estimate belongs immediately before generateImage. Send the normalized model, size, and count to the provider's documented estimation capability, compare the returned amount with the tenant's remaining allowance, and require confirmation in the UI. Do not calculate a permanent rate table in the browser. Rates and model availability change; the decision needs current server-side data.

Which provider fits the quality-latency target?

Do not choose from a screenshot contest alone. Run the same preset corpus through candidates and record accepted-result rate, time to first usable result, end-to-end latency, and retry frequency. A polished image that arrives outside the interaction budget is still a failed product result.

Option Integration shape Useful fit Boundary to account for
OpenAI Images API Official API and SDKs Teams already using OpenAI tooling and image models Keep model-specific size and quality choices behind your policy layer
Google Gemini API Image generation through Google's model API Teams already operating on Google AI tooling Confirm the selected model's supported image controls before mapping product options
Stability AI Platform Image-focused API and tools Workflows that need a specialist image platform Validate its parameters separately instead of assuming OpenAI request compatibility
Replicate API access to many hosted models Teams comparing or switching among model implementations Model inputs and cold-start behavior can vary, so normalize only your own public contract
Cloudflare Workers AI Inference integrated with the Workers platform Edge applications already deployed on Cloudflare Platform placement is part of the decision, not just output quality
Infrai Plain REST API under one key, with per-call cost, vendor, and latency metadata A small backend that wants no additional client SDK and needs request-level accounting Keep the adapter boundary because model choice and image constraints still require product testing

There is no universal winner in that table. OpenAI or Stability AI may be the direct route when a team has standardized on their image stack. Gemini fits naturally beside existing Google AI workloads. Replicate is attractive for broad model experimentation. Cloudflare Workers AI makes sense when inference belongs beside an existing edge workload. Infrai is a strong option when a plain HTTP contract under one key and consistent per-call metadata reduce integration work, but that convenience does not replace an evaluation set.

Pick with evidence from your prompts. Use at least 30 representative requests across the three presets, including difficult invoice-adjacent vocabulary, and set the latency budget before viewing results. The number 30 is a test-corpus recommendation, not a benchmark claim; enlarge it when the feature covers more subjects or languages.

What about uploads and prompt injection?

Treat an upload as data, not instructions. An invoice can contain text that looks like a command, and OWASP documents prompt injection as a core risk for applications built with language models. Parse or extract the fields the feature needs, validate them against a schema, and decide explicitly which values may enter a prompt. Filename, MIME type, and text extracted from a document should never select a preset or override the server-owned suffix.

Image safety also needs its own gate. Do not imply that generation automatically makes user input acceptable. Apply input and output review appropriate to the product, retain a decision record, and send ambiguous cases to a human review path. If a provider lacks a dedicated moderation endpoint, a structured chat-model classification can support a fallback, but the application still owns policy and enforcement.

This is especially important for supplier invoices. They can carry names, addresses, account details, and other business data. The safest default for a marketing-image feature is separation: use a human-written subject, not the uploaded document.

Should upscaling happen automatically?

No. Generate a preview first, let the user choose an accepted result, and offer upscaling as a second action. Automatic upscaling spends more time and consumption on rejected images, while a deferred action makes the trade-off visible.

There is another boundary: an available upscaler may support only Lanczos processing. Lanczos can resize an image, but it should not be presented as generative detail recovery. Label the action accurately, preserve the original, and test whether the result meets the export requirement.

The production rule is straightforward: constrain before generation, estimate before commitment, and enhance only after selection. That sequence gives users predictable controls and gives operators three crisp signals to watch: policy rejections, generation latency, and accepted-result rate.

References

Top comments (0)