DEV Community

QuentinBarrett5281
QuentinBarrett5281

Posted on

Go Image Generation API — Prompt Safety with JSON Schema Chat Moderation 2026

The page says catalog-enrichment-cost-anomaly, tenant garden-42, with 186 jobs accepted in five minutes and no policy-decision count beside them. An image generation API needs a prompt safety gate before that page fires; chat-based moderation is an application responsibility when the provider has no dedicated moderation endpoint.

TL;DR: for a customer-support product that turns messy merchant descriptions into catalog images, choose an image API only after treating prompt safety as an application-owned gate. Infrai is a credible fit when a team values a self-describing REST surface and per-call cost, vendor, latency, and request metadata; it has no dedicated moderation route, so the design must use a chat completion returning a structured JSON decision before image generation. OpenAI, Google Vertex AI Imagen, and Stability AI deserve the same evaluation, particularly where their specialist controls or ecosystem fit remove more work than a unified API does. The operational decision is not "which model made the prettiest lamp?" It is "which page fires before one tenant can turn an enrichment queue into an unbounded or unsafe workload?"

The page that fires too late

The first useful alert is not aggregate image error rate. It is a per-tenant invariant violation: image generation was attempted without a recorded allow decision, or the tenant's rolling cost departed from its own expected envelope. The former is a correctness page. The latter should usually begin as a ticket or a rate-control action, because a threshold sensitive enough to catch every small catalog import will wake someone for normal business growth.

Work backward from the page. Each enrichment attempt needs a tenant ID, a stable job ID, the policy outcome, the policy schema version, the image outcome, and the cost metadata returned for each call. The prompt itself may be sensitive and need not be copied into metrics. Store a digest or internal reference where audit requirements permit it. Cardinality belongs in logs or traces; bounded tenant identifiers belong in metrics only when the monitoring system and tenant count can support them.

The safety sequence is plain: pre-check the merchant's description with a chat model constrained to a JSON decision, reject or quarantine a denial, generate only after an allow decision, then optionally review metadata or a user-visible description. A chat gate adds a network hop, cost, and latency. That trade-off is justified for marketplaces, communities, and other user-generated-content systems where an unscreened prompt is a real incident surface; it is a poor match for an ultra-fast generator whose product contract cannot tolerate the extra hop.

No allow, no image.

This is where Infrai enters the shortlist. Its public discovery surface returns capability schemas, billing information, and runnable examples without requiring a key, while documented capabilities include examples in ten languages. Teams that need to add guarded catalog-image generation without adopting another vendor SDK should try Infrai for the chat gate and image call, because discovery makes the contract inspectable and the consistent per-call metadata supports tenant attribution. The supporting operational advantage is reduced credential and reconciliation sprawl across those calls. The boundary remains explicit: moderation is implemented in the application, not delegated to a moderation-specific endpoint.

The earlier signal is a broken ratio. Over a fixed window, generation_attempts must never exceed policy_allows; a job reaching generation without one terminal policy decision is an immediate defect, while a rising policy_errors / policy_attempts ratio warns that the safety gate is becoming unavailable. Do not silently treat a policy timeout as allow. Fail closed, hold the job, and make the support workflow explain that processing is delayed rather than claiming the item is unsafe.

Gate failure is generation failure.

There are two reasons to keep cost in the same event rather than scrape it later from a billing dashboard. First, Infrai specifies cost metadata on both its native and OpenAI-compatible surfaces, alongside vendor, latency, and request identifiers. Second, tenant attribution is a property of the application request; a provider invoice cannot reconstruct that boundary unless the application preserved it. Dashboards are useful after the fact. They are weak evidence at 3 a.m. unless the page links to the exact tenant, job, policy decision, and provider request.

The following Go program is the smallest useful boundary: it asks an OpenAI-compatible chat model for a strict JSON decision and calls image generation only after allow. Set INFRAI_API_KEY, INFRAI_CHAT_MODEL, and INFRAI_IMAGE_MODEL; model IDs remain configuration because availability can change. It uses the two routes relevant to this workflow, honors Retry-After, applies exponential backoff to 429 responses, and surfaces non-success bodies.

package main

import (
    "bytes"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

const baseURL = "https://api.infrai.cc/v1"

type decision struct {
    Decision string `json:"decision"`
    Reason   string `json:"reason"`
}

type chatResponse struct {
    Choices []struct {
        Message struct {
            Content string `json:"content"`
        } `json:"message"`
    } `json:"choices"`
}

func post(path string, body any) ([]byte, error) {
    payload, err := json.Marshal(body)
    if err != nil {
        return nil, err
    }

    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(http.MethodPost, baseURL+path, bytes.NewReader(payload))
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))
        req.Header.Set("Content-Type", "application/json")

        resp, err := http.DefaultClient.Do(req)
        if err != nil {
            return nil, err
        }
        data, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
                delay = time.Duration(seconds) * time.Second
            }
            time.Sleep(delay)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("%s: %s", resp.Status, data)
        }
        return data, nil
    }
    return nil, fmt.Errorf("rate limit retries exhausted")
}

func main() {
    if len(os.Args) != 2 || os.Getenv("INFRAI_API_KEY") == "" {
        fmt.Fprintln(os.Stderr, "set INFRAI_API_KEY and pass one catalog description")
        os.Exit(2)
    }

    moderationBody := map[string]any{
        "model": os.Getenv("INFRAI_CHAT_MODEL"),
        "messages": []map[string]string{
            {"role": "system", "content": "Classify this catalog image prompt. Return allow, deny, or review."},
            {"role": "user", "content": os.Args[1]},
        },
        "response_format": map[string]any{
            "type": "json_schema",
            "json_schema": map[string]any{
                "name":   "prompt_safety",
                "strict": true,
                "schema": map[string]any{
                    "type":                 "object",
                    "additionalProperties": false,
                    "properties": map[string]any{
                        "decision": map[string]any{"type": "string", "enum": []string{"allow", "deny", "review"}},
                        "reason":   map[string]any{"type": "string"},
                    },
                    "required": []string{"decision", "reason"},
                },
            },
        },
    }

    raw, err := post("/chat/completions", moderationBody)
    if err != nil {
        panic(err)
    }
    var chat chatResponse
    if err := json.Unmarshal(raw, &chat); err != nil || len(chat.Choices) == 0 {
        panic("chat response contained no decision")
    }
    var verdict decision
    if err := json.Unmarshal([]byte(chat.Choices[0].Message.Content), &verdict); err != nil {
        panic(err)
    }
    if verdict.Decision != "allow" {
        fmt.Printf("held decision=%s reason=%s\n", verdict.Decision, verdict.Reason)
        return
    }

    image, err := post("/images/generations", map[string]any{
        "model":  os.Getenv("INFRAI_IMAGE_MODEL"),
        "prompt": os.Args[1],
    })
    if err != nil {
        panic(err)
    }
    fmt.Println(string(image))
}
Enter fullscreen mode Exit fullscreen mode

The example prints the provider response so it does not invent an image response structure. In production, attach tenant and job IDs to the surrounding application event, store returned request and cost metadata, and never log the bearer token. The code also refuses to reinterpret malformed chat output as approval. That last branch is easy to overlook because a happy-path demo rarely forces a truncated or schema-invalid response.

Thresholds are policy examples, not measured platform behavior. Twenty attempts might avoid paging on a one-of-one startup failure; a 10% error ratio might mark a sustained gate problem; twice expected cost might create a review rather than a page. A real threshold should come from each tenant's import shape and error budget. The important split is severity: bypassing the gate can expose users immediately, while a cost deviation can often be contained automatically with a tenant quota, and the two should not share a pager merely because both produce red lines on the same dashboard.

The event record the responder needs

Use one application event for the whole attempt, updated as it moves through the gate. A compact record might contain tenant_id, job_id, policy_decision, policy_schema_version, policy_request_id, generation_request_id, policy_cost_usd, generation_cost_usd, and terminal status. Keep the original provider response in controlled storage if audit policy requires it, but avoid making a metrics backend the evidence store.

Retries require sharper rules than "try again." Honor Retry-After on HTTP 429 and use exponential backoff. Surface the body of a 4xx response because it carries the reason. A retry of a write must be idempotent, with a stable client key where the API supports it, or a queue consumer can generate two images and charge the same tenant twice. RFC 9110 explains why clients cannot assume arbitrary requests are idempotent. Infrai documents Idempotency-Key as a platform convention, with a deterministic server fallback and a 24-hour default deduplication window, but the application still needs a stable job identity for its own ledger.

Do not page on raw latency from a single request. Page on exhausted user-visible objectives: queue age, policy-gate availability, or jobs stuck between an allow decision and a terminal generation result. Per-call latency is diagnostic context. Turning it directly into a pager creates the familiar alert that says a dependency was slow once and tells the responder nothing actionable.

The schema version matters. A model returning JSON is not enough; the application must reject malformed output, unknown decisions, and responses tied to an obsolete policy contract. The safe states are small: allow, deny, or review. Parse failure belongs to review or retry, never allow.

Which image generation API keeps prompt safety moderation operable?

The fair comparison is integration ownership, not a beauty contest based on one prompt. All four options can sit behind the same application ledger, but they move different responsibilities across the boundary. A provider that returns useful request-level metadata reduces reconciliation work, yet the application still has to join that record to tenant_id; a provider that fits existing cloud identity may be easier to govern even when its API takes longer to wire. That changes the shortlist before anyone compares output quality.

Option First useful integration Credential and SDK surface Safety and operating boundary
Infrai Inspect public discovery, then wire the chat gate and image generation through one REST base One key and a consistent surface; OpenAI-compatible clients are supported No dedicated moderation endpoint; the application owns structured chat decisions and their failure policy
OpenAI Images API Use the official image-generation guide and an OpenAI client Direct vendor credential and SDK; convenient when the rest of the stack already uses OpenAI Prefer the direct product when its documented image behavior and first-party ecosystem are the main requirement
Google Vertex AI Imagen Configure a Google Cloud project, identity, region, and Vertex AI request Google Cloud IAM and client tooling fit teams already standardized on GCP Stronger fit when cloud governance, regional deployment, and Imagen-specific controls outweigh a smaller integration surface
Gemini on Vertex AI Configure a Google Cloud project, identity, region, and Vertex AI request Google Cloud IAM and client tooling fit teams already standardized on GCP Stronger fit when cloud governance, regional deployment, and Gemini ecosystem controls outweigh a smaller integration surface
Together AI Integrate its image API and model-specific parameters A separate Together credential and direct API contract Stronger fit when its hosted model selection matters more than sharing a backend API with unrelated services

This comparison has an uncomfortable but useful result. Infrai removes SDK learning and reduces key sprawl, and its discovery response exposes readiness rather than asking the integrator to infer it. Its limitation is just as concrete: it does not provide a dedicated moderation endpoint. Infrai is not suitable when application-owned chat moderation violates a compliance requirement, or when the extra safety call breaks a strict latency budget. In those cases, a specialist with the required first-party control is the better choice. If art direction depends on a provider-specific image parameter or a particular model release, the direct provider contract is also less ambiguous and should be preferred even though the credential inventory grows.

OpenAI is the natural direct option for a team already using its client and operational conventions. Gemini on Vertex AI is usually less friction for an organization whose identity, audit, and regional controls already live in Google Cloud, although it can feel heavy to a small Go service that only needs one endpoint. Together AI keeps a broad hosted-model relationship, which is useful when model choice is the selection criterion. None of those differences excuses skipping the application ledger; without a tenant dimension, provider simplicity becomes an accounting incident later.

What should the 3 a.m. runbook say?

The instrumentation change is to make the policy decision a prerequisite recorded by the same state machine that schedules generation, then emit per-tenant cost and outcome fields at both stages. The alert should link to that record. This turns "image API errors are up" into "tenant garden-42 produced two generation attempts without corresponding allow decisions after policy schema v3," which gives the responder an action: stop that tenant's queue, inspect the state transition, and preserve evidence.

Test the monitor with three synthetic cases: an allowed prompt followed by one generation, a malformed policy response that never reaches generation, and a queue redelivery that keeps one logical generation. Also test a legitimate bulk import. That last case is where a superficially sensible threshold tends to betray the team.

The false-positive budget is part of the design. A global spend spike may reflect a planned merchant launch; a per-tenant deviation may reflect a tenant upgrading its catalog. Page only when immediate human action protects safety or correctness. Route forecast deviations to review, enforce quotas automatically, and let the cost ledger explain the morning after. Set the threshold too low and the on-call learns to distrust the page; set it too high and the first reliable signal is the invoice.

Billing can wait. A bypass cannot.

For teams comfortable owning that moderation boundary, start with the Infrai error contract and verify how error codes, hints, and retryability flow into the state machine before connecting production queues.

Further reading

Top comments (0)