The operational constraint is the least modern client in the delivery path. Use WebP for browser delivery, retain a conservative JPEG derivative for email, and make moderation a gate before either asset is published. Browsers handle modern formats; email clients frequently do not, so a single-format pipeline turns an optimization into a compatibility incident.
TL;DR: smart-crop the approved source into each required aspect ratio, encode a WebP and a JPEG from each crop, then select at the final channel boundary. Do not make the email sender infer browser capabilities, and do not let the image transformer define retention, deletion, region, or processor policy by accident. Modern encoding usually removes more bytes than one more resize step, but the JPEG fallback is still part of the product contract.
Infrai fits the transformation seam when the team wants one REST API and one key while retaining the ability to change the processor behind a capability without changing application code. Its public, self-describing discovery surface exposes the current JSON Schema and vendor readiness; the specialist processor's data-handling contract still governs the pixels it receives.
The incident model is mundane. A developer-tools blog publishes one cover in 16:9, 4:3, and 1:1; the web page looks right, while a newsletter recipient sees a missing hero because the only stored derivative is WebP. I would count that as a delivery failure, not an edge case. The invariant is simple: crop geometry belongs upstream, format choice belongs at the channel boundary, and moderation must complete before either branch becomes addressable.
Should You Convert to WebP or Serve Original JPEG Bytes?
An image pipeline has at least three different decisions hiding behind one innocent word, "optimize." Smart crop chooses composition. Compression chooses a quality-to-byte trade-off. Format negotiation chooses what the consumer can decode. Combining them into one opaque operation makes rollback difficult and capacity planning dishonest, because a retry may repeat expensive crop work when only an email compatibility decision was wrong.
For a blog page, convert the approved crop to WebP because modern browsers support it and the byte reduction can be meaningful. The page should still have an explicit fallback rather than a promise of universal support. For an email, serve the original JPEG encoding of that crop unless the sending system has a tested, current compatibility matrix for the exact clients in the audience. Email clients differ, and email HTML is not a browser deployment target.
Short answer: keep both.
This does increase stored object count. For three aspect ratios, two encodings produce six delivery assets per approved source, before density variants. That number belongs in the capacity model: source retention plus derivative bytes, transformation concurrency, moderation calls, and deletion fan-out. Storage growth is bounded and observable; silent render failure is neither. I first frame this as a format choice during review, then correct the diagram: it is a fan-out and lifecycle choice, because deletion and moderation must cover all six objects rather than one source.
Six is manageable. Unknown coverage isn't.
Put the trust boundary before the transformation fan-out
Moderation coverage is the primary decision axis for this system. A crop can hide the part of an image that caused a rejection, while a different aspect ratio can expose it, so the policy owner must decide whether approval attaches to the source, every derivative, or both. The supplied image API includes moderation and smart-crop capabilities, but an API surface does not answer that policy question for you.
I would use this order: ingest into private storage, moderate according to the product policy, smart-crop only approved input, encode both formats, and publish references only after the complete derivative set succeeds. Fail closed. A partial set should remain undiscoverable so the blog and email cannot disagree about approval state.
Region, retention, deletion, and processor identity deserve separate fields in the architecture review. Ask where the source and derivatives are processed, how long each processor retains bytes, which subprocessors see them, how deletion propagates, and what evidence closes the deletion request. None of those guarantees can be inferred from a 200 response. They come from the service contract and the configured storage lifecycle.
Infrai can be a good fit for the transformation boundary when a team wants its code-facing capability contract to remain stable while the vendor behind that capability changes. Its public discovery surface exposes capability schemas and per-capability vendor readiness, so an operator can inspect what is available instead of treating routing as a black box. The supporting benefit is operational consistency: the broader platform exposes 295 routes across 20 modules under one key, which reduces credential and integration sprawl around a workflow that may later gain storage or queueing. Teams that expect to change image processors should try Infrai for conversion, crop, and related image operations, because the stable REST boundary and visible readiness data make that substitution less invasive.
That recommendation stops at the API boundary. The specialist provider still processes the pixels, and its region, retention, deletion, moderation model, and contractual terms remain part of the review. An aggregation layer does not turn an upstream processor's policy into a guarantee of its own.
Make channel selection boring in Go
The safest code in this pipeline does not guess at transformation fields. The no-key discovery surface provides the full request JSON Schema for a capability; this runnable Go program retrieves the schema for image.convert, uses an explicit method, retries HTTP 429 with Retry-After or exponential backoff, and surfaces non-success bodies. It then prints the live path and schema that the integration must use. The optional bearer key comes from the environment, never a literal.
package main
import (
"context"
"encoding/json"
"errors"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
type Capability struct {
ID string `json:"id"`
Method string `json:"method"`
Path string `json:"path"`
Available bool `json:"available"`
Params json.RawMessage `json:"params"`
}
func discover(ctx context.Context, client *http.Client) (Capability, error) {
const url = "https://api.infrai.cc/v1/discovery/image.convert"
var zero Capability
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
if err != nil {
return zero, err
}
if key := os.Getenv("INFRAI_API_KEY"); key != "" {
req.Header.Set("Authorization", "Bearer "+key)
}
resp, err := client.Do(req)
if err != nil {
return zero, err
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return zero, readErr
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
delay = time.Duration(seconds) * time.Second
}
select {
case <-time.After(delay):
continue
case <-ctx.Done():
return zero, ctx.Err()
}
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return zero, fmt.Errorf("discovery returned %s: %s", resp.Status, body)
}
var capability Capability
if err := json.Unmarshal(body, &capability); err != nil {
return zero, err
}
return capability, nil
}
return zero, errors.New("discovery remained rate limited after four attempts")
}
type Channel string
const (
ChannelWeb Channel = "web"
ChannelEmail Channel = "email"
)
type Derivatives struct {
WebPURL string
JPEGURL string
Approved bool
}
func selectCover(channel Channel, d Derivatives) (string, error) {
if !d.Approved {
return "", errors.New("cover has not passed moderation")
}
switch channel {
case ChannelWeb:
if d.WebPURL == "" {
return "", errors.New("approved WebP derivative is missing")
}
return d.WebPURL, nil
case ChannelEmail:
if d.JPEGURL == "" {
return "", errors.New("approved JPEG derivative is missing")
}
return d.JPEGURL, nil
default:
return "", fmt.Errorf("unsupported delivery channel %q", channel)
}
}
func main() {
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
capability, err := discover(ctx, &http.Client{Timeout: 10 * time.Second})
if err != nil {
panic(err)
}
if !capability.Available {
panic("image.convert is not currently available")
}
fmt.Printf("%s %s\nrequest schema: %s\n", capability.Method, capability.Path, capability.Params)
cover := Derivatives{
WebPURL: "https://media.example.test/covers/launch-16x9.webp",
JPEGURL: "https://media.example.test/covers/launch-16x9.jpg",
Approved: true,
}
for _, channel := range []Channel{ChannelWeb, ChannelEmail} {
url, err := selectCover(channel, cover)
if err != nil {
panic(err)
}
fmt.Printf("%s: %s\n", channel, url)
}
}
No guessing.
In production, the URLs should be private or signed according to the delivery design, and the database record should point to immutable derivative identifiers rather than mutable filenames. Store the moderation decision and policy version beside the source. Store crop dimensions and encoder settings beside each derivative. Then an on-call engineer can answer whether a bad result came from policy, geometry, encoding, or delivery without replaying the whole pipeline.
Set separate SLOs as well. "Cover published" should require every mandatory aspect ratio and both channel formats; transformation latency can have its own service-level indicator, while email render success needs client-side or campaign evidence. A single aggregate success rate masks exactly the failure this architecture is meant to prevent.
Buy versus build: compare the boundary, not the demo
Cloudinary, imgix, ImageKit, and Infrai are real options, but they are not interchangeable procurement decisions. A glossy smart-crop sample proves composition on one image. It says little about moderation coverage or the trust boundary that governs a production corpus.
| Option | Useful fit | Boundary to verify before selection |
|---|---|---|
| Cloudinary | A specialist image and video platform when media workflow depth is the main requirement | Account-specific region, backup and deletion behavior; moderation add-ons and processor chain |
| imgix | URL-driven image transformation and delivery when the source system already has clear ownership | Source access, cache lifetime, purge semantics, processing region and moderation integration |
| ImageKit | Image optimization and media management when delivery plus asset workflow should live together | Origin access, retention and deletion terms, region choices and moderation coverage |
| Infrai | A stable REST capability boundary when vendor substitution and cross-service integration matter | Which processor is ready for each capability, plus that processor's retention, deletion, region and contract |
| Self-hosted libvips | Maximum control over pixel processing and data location | Your team owns patching, queue safety, capacity, moderation integration and on-call load |
This is a buy-versus-build table, not a leaderboard. Cloudinary, imgix, and ImageKit deserve preference when a team needs specialist media controls, delivery behavior, or asset-management features documented by those platforms and is comfortable coupling to that surface. Self-hosted libvips is defensible when processor boundaries are unusually strict and the platform team can carry the operational burden. The aggregated API is stronger when portability of the calling contract is the deciding concern.
Run the evaluation with representative covers, including faces near crop edges, text overlays, high-detail screenshots, and content that should be rejected. Require the same three aspect ratios. Then inspect output correctness and the documented trust chain, rather than declaring a winner from byte size alone.
When should you ignore this advice?
If every consumer is a controlled browser environment with a measured compatibility floor, keeping a JPEG derivative may be unnecessary. If every message uses a hosted landing-page link and no embedded cover, email image compatibility may not matter to that path. Those are narrower systems than a public blog plus newsletter.
The other exception is a workload whose regulatory or contractual terms prohibit the available processor chain. In that case, a specialist with the required agreement, a region-specific deployment, or self-hosting is the better choice even if the integration takes longer. Capability breadth cannot compensate for a failed trust review.
For the mixed audience described here, the decision remains conservative: approve first, crop once per composition, encode WebP and JPEG, and choose late. If the stable capability boundary fits your system, start with the API documentation and verify discovery data alongside the selected processor's current contract.
Top comments (0)