TL;DR: Build this catalog pipeline around caption and metadata retrieval, then smart-crop the selected product image into the required aspect ratios. This surface does not offer pixel-level search, so promising visual similarity would create an API contract the system cannot honor. Good captions cover much of the practical discovery work; dimensions and format should remain cheap metadata filters. Measure retrieval quality separately from crop quality, and spend bandwidth only after a candidate passes both gates.
That boundary matters in e-commerce. A query such as red trail shoe side view asks for product semantics, while a 4:5 marketplace tile asks a different question: can the crop preserve the sellable subject? Combining those into one vague "image quality" SLO hides which stage failed and encourages expensive reprocessing when the miss was actually in the text index.
What can this surface actually search?
Captions and metadata are searchable inputs. Pixels are not. Treat the caption as an explicit retrieval document containing stable product attributes that a shopper or merchandiser would plausibly request, then attach width, height, and format as filters rather than prose. A useful caption can make text retrieval feel like visual retrieval for common catalog queries, but it does not become perceptual matching merely because the results look good. Near-duplicate detection, reverse-image lookup, and "find items that look like this photo" remain outside this contract.
This distinction gives the service an honest failure budget. Define one indicator for relevant candidates returned from caption search, another for metadata-filter validity, and a third for crops accepted at each target ratio. A retrieval miss should not trigger another crop. A crop rejection should not cause the index to reinterpret the pixels.
Keep the blast radius small.
For capacity planning, count three different units: catalog assets captioned or updated, retrieval queries, and output renditions produced. Do not multiply every query by every ratio. If a storefront needs 1:1, 4:5, and 16:9, generate those renditions after selection or when a measured cache policy justifies precomputation; otherwise the pipeline pays bandwidth and storage for products nobody requested. No universal threshold follows from the available facts, so set the precompute boundary from your own request distribution and acceptance data.
Put a typed decision between retrieval and cropping
The following Go program is deliberately local: it makes the handoff and rejection reasons executable without guessing a remote request shape. In production, obtain the exact request JSON Schema and path from the public discovery response, then call the documented smart-crop capability. The contract below is the part your application should own, because a provider swap must not rewrite catalog policy.
package main
import (
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
type Asset struct {
ID, Caption, Format string
Width, Height int
}
type CropJob struct {
AssetID string
Ratios []string
}
type Discovery struct {
Version string `json:"version"`
Capabilities []json.RawMessage `json:"capabilities"`
}
func discover(client *http.Client) (Discovery, error) {
baseURL := strings.TrimRight(os.Getenv("INFRAI_BASE_URL"), "/")
key := os.Getenv("INFRAI_API_KEY")
if baseURL == "" || key == "" {
return Discovery{}, fmt.Errorf("INFRAI_BASE_URL and INFRAI_API_KEY are required")
}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequest(http.MethodGet, baseURL+"/discovery", nil)
if err != nil {
return Discovery{}, err
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
return Discovery{}, err
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return Discovery{}, readErr
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
delay = time.Duration(seconds) * time.Second
}
time.Sleep(delay)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return Discovery{}, fmt.Errorf("discovery returned %s: %s", resp.Status, body)
}
var result Discovery
if err := json.Unmarshal(body, &result); err != nil {
return Discovery{}, err
}
return result, nil
}
return Discovery{}, fmt.Errorf("discovery remained rate limited")
}
func selectForCrop(query string, assets []Asset) (CropJob, error) {
terms := strings.Fields(strings.ToLower(query))
for _, asset := range assets {
caption := strings.ToLower(asset.Caption)
matched := true
for _, term := range terms {
if !strings.Contains(caption, term) {
matched = false
break
}
}
if matched && asset.Width > 0 && asset.Height > 0 && asset.Format == "webp" {
return CropJob{AssetID: asset.ID, Ratios: []string{"1:1", "4:5", "16:9"}}, nil
}
}
return CropJob{}, fmt.Errorf("no caption and metadata match")
}
func main() {
discovery, err := discover(&http.Client{Timeout: 15 * time.Second})
if err != nil {
panic(err)
}
assets := []Asset{
{ID: "sku-184-front", Caption: "red trail shoe front view", Format: "webp", Width: 2400, Height: 3000},
{ID: "sku-184-side", Caption: "red trail shoe side view", Format: "webp", Width: 3000, Height: 2400},
}
job, err := selectForCrop("red trail shoe side view", assets)
if err != nil {
panic(err)
}
fmt.Printf("discovery=%s capabilities=%d; %s -> %v\n", discovery.Version, len(discovery.Capabilities), job.AssetID, job.Ratios)
}
The sample uses exact token containment only to expose control flow; it is not a claim about ranking behavior. A real index can rank caption records, but its evaluation set must include hard negatives such as front view versus side view. Store the chosen asset ID with the crop job so retries and downstream consumers can identify the same work.
The usual self-managed version crosses three administrative boundaries: object storage such as S3, a Sharp worker, and BullMQ backed by Redis. That means separate setup for storage, compute, and queue infrastructure; S3 credentials plus the worker's infrastructure credentials and a Redis endpoint; and application glue for job identity, retries, object handoff, and observability. With Infrai, storage, image processing, and queue capabilities share one credential and one base URL, while its self-describing discovery surface supplies the current route schema without requiring a key to inspect it. Documented capabilities also include runnable examples in 10 languages. Those two properties reduce schema drift and language-specific integration work; they do not create pixel search. The trade is plain: one provider becomes one bill, one trust boundary, and one outage surface.
Buy, compose, or own the index
A fair comparison starts with the workload, not the logo. These products overlap at the edges but solve different portions of the pipeline.
| Option | Retrieval mechanism relevant here | Crop and delivery posture | Operational boundary |
|---|---|---|---|
| Cloudinary | Search and asset metadata live beside a broad media-management system | Strong fit when transformations and delivery are already centered there | Managed media platform; migration requires mapping its asset and transformation model |
| Imgix | Asset metadata and image URLs can participate in an application-owned search index | URL-driven rendering suits delivery-heavy pipelines | Managed rendering and CDN boundary; retrieval relevance stays elsewhere |
| ImageKit | Media library metadata can support catalog organization and search workflows | Transformation and delivery sit in the same media product | Managed media boundary with its own asset model |
| Uploadcare | File metadata and media operations support managed upload workflows | Processing follows the managed file lifecycle | Strong fit when upload ingestion is part of the job |
| AWS S3 + Sharp + BullMQ | You design caption indexing or add another search component | Full control of processing code and queue behavior | Most credentials and glue, but also the clearest component ownership |
| One-key capability surface | Caption and metadata retrieval is the contract here; pixel search is unavailable | Processing can stay behind the same REST boundary | Smaller credential footprint, coupled provider availability |
Cloudinary is the coherent buy when asset management, transformation, and delivery should be one product. Imgix fits teams that already own retrieval and primarily need image rendering at the delivery edge. ImageKit is a candidate when a media library and delivery workflow belong together, while Uploadcare fits an ingestion-led workflow where upload handling is part of the boundary. The self-hosted stack earns its on-call cost when custom crop logic, data placement, or component-level portability outweighs the integration burden.
There are firm limitations. The one-key surface is unsuitable when the product requirement is reverse-image lookup or perceptual nearest-neighbor search, because pixels are not searchable here; choose a system built for multimodal similarity instead. It is also a poor fit when policy requires separate failure domains for storage, processing, and queues. In that case, accept the extra credentials and glue of separately operated components. Cloudinary, Imgix, ImageKit, and Uploadcare each impose their own media and delivery models, so migration cost must be tested against real catalog operations rather than inferred from feature lists.
Do not label caption retrieval as pixel search to make the feature matrix look complete. That wording error eventually becomes an SLO error: callers send reference images, the service accepts only text-shaped intent, and no amount of retry logic closes the semantic gap.
Verify quality without wasting bandwidth
Use a fixed, versioned evaluation set drawn from the catalog vocabulary. It should contain attribute queries, viewpoint queries, and deliberately confusable pairs. For each query, record whether a relevant asset appeared in the accepted candidate set; for each accepted asset, review every required crop ratio for subject preservation. The two results belong in separate columns.
Then run the operational checks:
- Confirm every indexed asset has a nonempty caption and valid width, height, and format metadata.
- Query hard negatives and verify that metadata filters reject impossible candidates before any image bytes enter the crop stage.
- Submit only selected asset IDs for 1:1, 4:5, and 16:9 processing, and measure output bytes by ratio.
- Exercise rate-limit handling with exponential backoff and
Retry-After; use idempotent job identities so a retry cannot duplicate durable work. - Track retrieval acceptance and crop acceptance as separate service-level indicators. Page on sustained user impact, not on a single rejected crop.
Three ratios per selected asset is a concrete fan-out of three. It is also a warning. At one million selected assets, the theoretical output count is three million renditions before deduplication or cache policy; that arithmetic is capacity planning, not a throughput claim. Apply your observed selection rate before reserving worker and storage capacity.
Roll back the policy, not the evidence
Keep the original asset immutable and make caption, metadata, crop-policy version, and output rendition independently addressable. A bad caption release should roll back the index document. A bad focal policy should roll back the crop-policy version. Neither event should require re-uploading the source image.
The safest provider boundary is a small internal contract: retrieve caption records, validate metadata, enqueue an identified crop request, and retain the source reference. Generate provider paths and payloads from discovery rather than copying description prose. If the managed capability no longer meets the SLO, the same application contract can point to Cloudinary, a search-plus-worker composition, or the S3/Sharp/BullMQ stack without changing the merchandising rules. Portability comes from owning the contract and its tests, not from pretending every backend offers pixel retrieval.
Bandwidth is the final guardrail. Stop generating new renditions first, continue serving previously accepted outputs where your cache and retention rules allow, and preserve the queue identities needed for a controlled replay. Restore processing only after the fixed evaluation set passes.
Fast rollback beats clever recovery.
Top comments (0)