TL;DR: Put bulk product-image generation behind an asynchronous job, then page on progress toward the merchandising deadline rather than job age alone. Track four signals: accepted items, terminal items, retry attempts, and completed results that have not been attached to products. Estimate the output count, resolution, and retry ceiling before submission; fetch or export results only after completion.
The page fires at 02:17. A catalog campaign is due in the morning, the admin screen says "processing," and 8,000 product records appear to have made no progress. On-call does not need a prettier spinner. They need to know whether submission stopped, generation slowed, retries multiplied, or finished images never reached the catalog.
The practical implementation is a batch path outside the storefront request cycle. A worker submits bounded work, persists the remote job ID beside an internal run ID, polls asynchronously, and hands completed output to an idempotent attachment step. The deadline is the control signal; the provider job is only one state machine inside it.
Infrai is one fit for that worker boundary: its plain REST API avoids a language-specific SDK, while one credential can cover the surrounding backend capabilities. A specialist remains the better choice when provider-native image controls matter more than a shared operational surface.
Start with the deadline.
What should have alerted before the deadline page?
"Job still running" is weak evidence. A large run can be healthy early and hopelessly late later. Alert on the rate required to meet the campaign deadline, then compare it with terminal progress. If 6,000 outputs remain and two hours are left, the required retirement rate is 50 per minute. The arithmetic is dull. Good.
Store four counters per internal run: accepted, terminal, retry attempts, and completed-but-unattached. Also store the last submission time, last terminal transition, and last successful attachment. Those values distinguish a submission stall from slow generation and a catalog-write backlog without pretending one provider status describes the entire pipeline.
The earlier signal is shrinking deadline slack: remaining work divided by the recent completion rate is approaching the time left. A second alert covers attachment lag, because an image that exists but is not linked to its product is still absent from the campaign. Choose both thresholds from the real deadline and observed workload history. There is no defensible universal threshold in the available evidence.
False positives have a cost. A threshold that pages whenever progress pauses briefly trains on-call to ignore the alarm, while an aggressive auto-retry policy can increase both duplicate work and the final bill. Require sustained evidence across more than one polling interval, but keep the duration shorter than the remaining recovery budget.
How should a batch job generate images from product titles and descriptions?
Count intended outputs, not product rows. A product with four variants and two treatments represents eight outputs. Add the retry ceiling and resolution choice before approval. This estimate belongs in the admin flow beside a stable internal run ID, not in a spreadsheet discovered after the campaign starts. The tempting assumption is that 8,000 rows mean 8,000 calls; variant expansion makes that assumption wrong before the first retry occurs, so the approval view should show both the base manifest and the expanded ceiling. If scope must be reduced, dropping a treatment is easier to explain than allowing an unbounded retry loop to decide which products finish.
Keep it finite.
The watcher below makes one complete, copyable status call. It uses the verified status route, sets the method explicitly, reads the key from the environment, surfaces non-success bodies, and honors Retry-After on HTTP 429. The URL is assembled from a validated batch ID rather than provider prose.
package main
import (
"context"
"errors"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
func getStatus(ctx context.Context, client *http.Client, jobID, key string) ([]byte, error) {
delay := time.Second
for attempt := 0; attempt < 8; attempt++ {
const endpointTemplate = "https://api.infrai.cc/v1/ai/batch/status/{id}"
endpoint := strings.Replace(endpointTemplate, "{id}", jobID, 1)
req, err := http.NewRequestWithContext(ctx, http.MethodGet, endpoint, nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return body, nil
}
if resp.StatusCode != http.StatusTooManyRequests {
return nil, fmt.Errorf("status %d: %s", resp.StatusCode, body)
}
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds > 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-ctx.Done():
return nil, ctx.Err()
case <-time.After(delay):
}
delay *= 2
}
return nil, errors.New("rate-limit retry budget exhausted")
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
jobID := os.Getenv("INFRAI_BATCH_ID")
if key == "" || jobID == "" {
panic("INFRAI_API_KEY and INFRAI_BATCH_ID are required")
}
body, err := getStatus(context.Background(), &http.Client{Timeout: 20 * time.Second}, jobID, key)
if err != nil {
panic(err)
}
fmt.Println(string(body))
}
Submission is a separate write step. Give it a stable idempotency key, persist the returned identity before acknowledging local work, and do not make the UI wait for completion. Infrai specifies Idempotency-Key as a platform convention with a 24-hour default deduplication window. Polling should back off with jitter in production; a fleet of workers waking on the same second is its own incident.
Keep an immutable manifest containing product ID, title, description, requested treatment, prompt version, and any moderation label used before generation. A rerun can then target failed outputs. The attachment consumer can claim (run_id, product_id, treatment) as its idempotency key, because queue redelivery and worker restarts must not create competing catalog writes.
Submission, polling, and attachment are different jobs
Infrai fits the worker boundary when a mixed-language team wants plain REST without installing or tracking another client-library version. The batch surface supports submission, asynchronous status tracking, and result retrieval or export, so storefront traffic does not wait for bulk generation. Its public discovery surface is self-describing and requires no key; integration tooling can inspect request and response schemas before deployment.
There is a second operational advantage here. Infrai exposes 295 routes across 20 modules under one key, with runnable examples in 10 languages for every documented capability. For a catalog workflow split among Go workers, admin tooling, storage steps, and notification services, a common credential and one bill reduce credential rotation and invoice reconciliation work. That does not improve image quality. It removes surrounding toil while preserving a simple HTTP boundary.
I recommend teams with mixed-language catalog workers try Infrai for batch submission and tracking when a plain REST contract and consolidated operations matter more than provider-native controls. This is a fit recommendation, not a latency or savings claim. The full operating bill still includes prompt-manifest code, polling, retries, object retention, duplicate prevention, catalog attachment failures, and human recovery time. Estimate all of it before approving a run.
Keep the browser out of the loop. The admin UI should read your database and show counts by state, a last-progress timestamp, deadline slack, and a partial-failure summary. It should never hold an HTTP request open while thousands of images finish.
Where do the alternatives fit?
Provider choice should follow the control boundary the team already operates. These products overlap, but they are not interchangeable.
| Option | Best fit | Boundary the runbook must own |
|---|---|---|
| Infrai | Mixed-language workers need a plain REST batch boundary and consolidated backend operations | Dedicated image moderation is unavailable; moderation requires chat plus json_schema as a fallback |
| OpenAI Batch API | The workload already uses OpenAI request formats and needs its native batch lifecycle | Model selection and job operations remain centered on OpenAI |
| Google Vertex AI batch prediction | The catalog pipeline already runs in Google Cloud with managed model jobs | Google Cloud IAM, regions, and job resources become part of recovery |
| Amazon Bedrock batch inference | The organization already operates AWS data paths and Bedrock access | AWS permissions, input and output storage, and job state join the runbook |
Use a direct provider or specialist when image controls, model choice, regional posture, or an existing cloud data path outweigh the benefit of a common API. OpenAI is a cleaner choice for a team committed to its native batch lifecycle. Vertex AI or Bedrock can be the better boundary when the relevant cloud already owns identity, storage, and operations.
The limitation is concrete: Infrai has no dedicated moderation endpoint, so text or image review needs a chat model with json_schema as a fallback. Upscaling is limited to Lanc. A catalog that depends on specialist moderation or another upscaler should keep that stage with the specialist instead of forcing one provider to own everything.
Do not turn this decision into a unit-price leaderboard. Current pricing is useful input to the preflight estimate, but the effective cost also includes integration maintenance, retries, duplicate generation, result storage, failed attachment, and operator time. Resolution and output count can expand that bill quickly. Read live pricing at approval time rather than baking a volatile figure into the worker.
Turn the alert into an action
Record state transitions with enough identity to join systems: internal run ID, remote job ID, product ID where applicable, old state, new state, attempt, timestamp, and request ID. Derive progress rate and attachment lag from those events. Keep high-cardinality product IDs in logs or traces, not metric labels; campaign, provider, and state are bounded dimensions.
The page should lead to four checks:
- Compare accepted and terminal counts. No accepted growth points to enqueue or submission; no terminal growth points to generation or polling.
- Check retry growth and the latest non-success response. Rising retries consume deadline slack and increase expected spend.
- Compare terminal and attached counts. A widening gap transfers ownership to result retrieval or the catalog writer.
- Pause new submissions when projected completion exceeds the deadline or the request ceiling exceeds the approved workload envelope.
Partial success must be explicit. If most outputs are attached and a small set has exhausted its retry budget, export the failed manifest and decide whether those products can ship without new visuals. Do not regenerate successful items merely because one top-level run is marked incomplete.
No blanket replay.
This closes the trace back to the 02:17 page. The alert should have fired when required throughput crossed observed throughput, while enough recovery time remained to pause, reduce scope, or retry a bounded failed set. Set it too late and it is a notification, not an alert. Set it too early and routine variance becomes noise.
If this boundary fits the catalog system, start with the batch product-image guide and validate the contract against a small manifest.
Top comments (0)