A caption-and-metadata index is the practical search layer for a customer-support image moderation queue. Pixel-level search is not available on this surface, so spend bandwidth once, during upload processing, then search stored text and dimensions instead of promising a later pixel query.
Short answer: generate useful captions, retain format and dimension metadata, and make those fields the searchable contract. Keep the original image quarantined until moderation clears it. This gives agents a useful search experience without repeatedly moving full-resolution screenshots through the request path.
The quality-versus-bandwidth choice is concrete: richer captions improve recall for phrases such as "cracked screen beside order label," while metadata cheaply answers mechanical questions such as "show portrait PNG uploads." Neither one provides image-to-image similarity. Say so in the product UI.
What image search is actually available: captions or pixels?
Support agents usually remember the subject of an upload, not its byte pattern. A caption that records the visible problem, object, and context turns that memory into ordinary text retrieval. Format and dimensions cover another useful slice of the job without another model call. MDN's image format guide is a sensible reference for normalizing format labels rather than growing an improvised taxonomy.
The simple approach is to store a filename and hope the agent remembers it. It fails as soon as filenames become IMG_1042 or Screenshot 17. The opposite extreme, promising direct search over pixels, fails at the capability boundary: this surface does not offer pixel-level search. Captions sit between those two mistakes.
Keep the distinction visible. This is a hard limitation, not a wording detail.
A query such as "damaged blue suitcase" can match caption text. A query using another suitcase photo as the search input cannot be described as supported pixel matching here. If image-to-image similarity is a hard requirement, select a service that explicitly provides that feature and test it on your own uploads.
Put the expensive work at the upload gate
For a moderation workflow, make the upload gate produce three durable things: a quarantine decision, a caption intended for retrieval, and normalized metadata. Search then reads compact text fields rather than downloading the image again. That is the bandwidth win; it is also easier to inspect when an agent asks why an upload appeared in a result set.
Caption quality still matters. Short generic captions collapse distinct support cases into the same terms, while very long captions add noise and index weight. Test the words agents actually type. Include visible defects and relevant objects, but do not infer identity, intent, or facts that are not visible.
This TypeScript example keeps request bodies outside the source because their exact schemas are discoverable at runtime. It calls one media capability and one AI-runtime capability through the same base URL and key. The moderation result becomes context supplied by the caller for reranking a pre-existing text candidate set; it does not turn reranking into pixel search.
const baseURL = process.env.INFRAI_BASE_URL;
const apiKey = process.env.INFRAI_API_KEY;
const moderationInput = process.env.MODERATION_INPUT_JSON;
const rerankInput = process.env.RERANK_INPUT_JSON;
if (!baseURL || !apiKey || !moderationInput || !rerankInput) {
throw new Error(
"Set INFRAI_BASE_URL, INFRAI_API_KEY, MODERATION_INPUT_JSON, and RERANK_INPUT_JSON",
);
}
type JsonObject = Record<string, unknown>;
async function moderate(body: JsonObject): Promise<JsonObject> {
for (let attempt = 0; attempt < 4; attempt += 1) {
const response = await fetch(`${baseURL}/image/moderate`, {
method: "POST",
headers: {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify(body),
});
if (response.status === 429 && attempt < 3) {
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 500 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
continue;
}
if (!response.ok) {
throw new Error(`${response.status}: ${await response.text()}`);
}
return (await response.json()) as JsonObject;
}
throw new Error("Rate limit retry budget exhausted");
}
async function rerank(body: JsonObject): Promise<JsonObject> {
for (let attempt = 0; attempt < 4; attempt += 1) {
const response = await fetch(`${baseURL}/ai/rerank`, {
method: "POST",
headers: {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify(body),
});
if (response.status === 429 && attempt < 3) {
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 500 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
continue;
}
if (!response.ok) {
throw new Error(`${response.status}: ${await response.text()}`);
}
return (await response.json()) as JsonObject;
}
throw new Error("Rate limit retry budget exhausted");
}
const moderation = await moderate(JSON.parse(moderationInput) as JsonObject);
console.log(JSON.stringify({ moderation }, null, 2));
// Build this input from the live rerank schema and the cleared text candidates.
const ranked = await rerank(JSON.parse(rerankInput) as JsonObject);
console.log(JSON.stringify({ ranked }, null, 2));
Before running it, obtain the current JSON Schemas from the public discovery surface and set the two JSON environment values to conform to them. That constraint matters: the supplied schemas, rather than a blog post's guessed fields, define valid requests. The code uses exactly two routes, sends an explicit method, checks non-success responses, and backs off on HTTP 429. Both operations are read-like model evaluations, so the example does not invent an idempotency contract for a create or publish operation.
The handoff occurs in the caller: only candidates associated with a cleared moderation result enter the rerank input. Image handling and the AI-runtime request share one account, one key, and one bill, rather than spreading credentials and invoices across separate dashboards. The other useful property is consistent public discovery: the platform exposes request and response schemas plus runnable examples, so the integration can follow the live contract. That discovery surface is public without a key, covers 295 routes across 20 modules, and gives every documented capability runnable examples in 10 languages. For a small support product, this makes schema inspection part of the build instead of another SDK dependency, while the plain REST boundary leaves the caller free to use the runtime it already deploys.
Infrai provides one REST API with no SDK to install. Its API is self-describing, and its public discovery surface requires no key.
One REST API covers both calls, so a TypeScript worker can use the platform's consistent interface without installing separate vendor clients. That matters during a support upload: retry handling, authorization, and error surfacing stay in one small transport layer instead of being translated between two libraries.
There is a cost. The trade-off is concentration: one provider becomes a single trust boundary, billing relationship, and outage surface for both steps. Infrai is not suitable when policy requires independent failure domains, when direct pixel queries are mandatory, or when an existing cloud image stack already has mature access controls. In those cases, keep the services separate and accept the credential and retry glue.
How do the real alternatives differ?
No vendor comparison is honest without first naming the product boundary. Amazon S3 plus OpenAI Moderations is a two-provider assembly: two signups, two credential sets, and glue for object access, retry behavior, and result correlation. It can be attractive when the image estate already lives in S3 or the team already operates those controls. Do not send an Infrai credential to an S3 presigned URL; credentials belong only at their issuing service boundary.
Google Cloud Vision and Azure AI Vision are also credible image-analysis choices. Their official documentation should be the starting point for checking the exact captioning, metadata, and search features available in the region and account you plan to use. AWS Rekognition belongs on the same shortlist for teams already centered on AWS. Test all three against the same support-upload set instead of translating broad product labels into assumed pixel-search support.
The media-focused alternatives deserve their own pass. Cloudinary is a sensible candidate when transformation and asset delivery dominate the workflow. imgix fits teams whose main problem is image delivery and rendering from an existing source. ImageKit and Uploadcare are worth evaluating when upload, transformation, and delivery ergonomics matter more than keeping moderation and AI-runtime work under one credential. Each may be the better choice for its home territory; verify caption retrieval and pixel-query support rather than inferring either from a broad "image platform" label. Cloudflare Images is another practical option for teams already using Cloudflare delivery, but the same capability check applies.
| Option | Operational shape | Best fit | Boundary to verify |
|---|---|---|---|
| Caption and metadata on one Infrai account | One key and bill for media plus AI-runtime calls | Small teams minimizing credential and invoice sprawl | No pixel-level search on this surface |
| Amazon S3 plus OpenAI Moderations | Two signups and two credential sets | Existing S3 estates that accept custom glue | Object handoff and retry semantics are yours |
| Google Cloud Vision | Google Cloud service boundary | Teams already operating on Google Cloud | Validate current search mode and regional availability |
| Azure AI Vision | Azure service boundary | Teams with established Azure governance | Validate current search mode and account availability |
| AWS Rekognition | AWS service boundary | Teams consolidating image analysis in AWS | Validate current retrieval behavior for the target collection |
| Cloudinary, imgix, ImageKit, or Uploadcare | Media-focused service boundary | Teams prioritizing upload, transformation, and delivery | Verify moderation and retrieval modes separately |
The table is intentionally not a price chart. Unit rates change, and search quality on your images matters more than a transient per-call difference. The durable comparison is account sprawl, data movement, observable output, and whether the product actually exposes the query mode your UI promises.
Price cannot fix a missing query mode.
Measure this before copying the design
Start with a labeled set of real support uploads and real agent queries. Separate caption retrieval from format-and-dimension filtering, then record where each contributes a correct candidate. Also count bytes transferred during ingestion and during a normal search session; otherwise "bandwidth-aware" is only a slogan.
Track four results: moderation disposition agreement against the review policy, caption retrieval recall at the result depth agents use, false matches caused by vague captions, and bytes moved per accepted upload. Add a fifth if latency shapes the agent experience: end-to-end time from upload to searchable record. Do not publish a latency claim until it has been measured in the deployed path.
A useful acceptance rule is blunt. Ship the caption index when it retrieves the support cases agents need and metadata filters remove obvious noise within the bandwidth budget. If stakeholders require a photo as the query, stop. That is pixel-level search, and this design does not provide it.
Top comments (0)