Legacy image migration is a release change, not a file-copy task. In an e-commerce catalog, a product photo may feed pages, search indexes, and an OCR text-extraction job at the same time.
Short answer: inspect metadata first, convert as a separate persisted stage, retrieve the derivative to verify it, and only then change downstream references or send the new asset to OCR.
The incident lesson: acceptance is not verification
I've been paged after missed jobs and duplicate deliveries. The pattern is familiar: a worker receives a retry, writes another object, and advances a pointer before anyone checks what the object contains. A conversion request being accepted does not establish that the resulting asset satisfies the migration contract. Treating those events as equivalent moves uncertainty into every product page and OCR consumer downstream.
The invariant I use is simple: every stage gets an identifier, a durable record, and a validation gate. Keep the source reference immutable. Store the derivative ID, source-to-derivative lineage, requested format, and validation result in one migration record. If a worker exits after conversion but before the pointer update, the next attempt can resume from that record rather than create a second derivative. This matters under at-least-once delivery, where duplicate execution is expected behavior rather than an exotic edge case.
That record also gives support and cleanup work a trustworthy join key. An unreferenced derivative with lineage is a controlled cleanup candidate; an untracked derivative is guesswork.
Keep the source.
How should image metadata, format conversion, and verification govern migration?
Treat upload-time conversion and on-demand conversion as two scheduling policies over the same state machine. Upload-time work gives the OCR path a prepared asset and exposes unsuitable source files before publication, but it adds work to ingestion and converts photos that might never be read. On-demand work keeps ingestion small and avoids cold assets, but the first product-page or OCR request now owns conversion latency and retry behavior.
For a catalog with predictable traffic, precomputing a canonical derivative after metadata inspection is usually the calmer operational choice. For a long tail of rarely viewed listings, on-demand conversion can fit better. Your mileage may vary: the boundary depends on the read-latency objective, storage retention, and whether OCR text must exist before a listing can be published.
Don't let either policy skip verification. A migration record should move through application-owned states such as inspected, converted, verified, and published. Validate the result recorded at one stage before scheduling the next, and stop polling when the provider reports a terminal state. Only the verified transition may release a derivative ID to downstream references.
One practical split is to convert at upload for active catalog items whose OCR output is part of search indexing, then defer low-traffic archive images until access. The state machine stays identical — only the event that schedules it changes. That separation is useful during a backfill because operators can pause scheduling without changing correctness rules, and a replay cannot bypass the verification gate.
A bounded, idempotent worker path
The Go program below deliberately takes capability payloads from files. The public discovery document supplies the current JSON schemas, so the sample does not guess fields that vary by capability. It persists each response before the next request, uses an application-supplied derivative ID for retrieval, and calls only the three operations in this migration path.
package main
import (
"bytes"
"context"
"fmt"
"io"
"net/http"
"net/url"
"os"
"path/filepath"
"strconv"
"strings"
"time"
)
const getPathTemplate = "/v1/image/get/{id}"
var baseURL = os.Getenv("INFRAI_BASE_URL")
func call(ctx context.Context, method, path, key, idem string, body []byte) ([]byte, error) {
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequestWithContext(ctx, method, baseURL+path, bytes.NewReader(body))
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
if len(body) > 0 {
req.Header.Set("Content-Type", "application/json")
}
if idem != "" {
req.Header.Set("Idempotency-Key", idem)
}
res, err := http.DefaultClient.Do(req)
if err != nil {
return nil, err
}
data, readErr := io.ReadAll(res.Body)
res.Body.Close()
if readErr != nil {
return nil, readErr
}
if res.StatusCode == http.StatusTooManyRequests {
wait := time.Duration(1<<attempt) * time.Second
if raw := res.Header.Get("Retry-After"); raw != "" {
if seconds, parseErr := strconv.Atoi(raw); parseErr == nil {
wait = time.Duration(seconds) * time.Second
}
}
select {
case <-time.After(wait):
case <-ctx.Done():
return nil, ctx.Err()
}
continue
}
if res.StatusCode < 200 || res.StatusCode >= 300 {
return nil, fmt.Errorf("%s: %s", res.Status, data)
}
return data, nil
}
return nil, fmt.Errorf("429 retry budget exhausted")
}
func persist(dir, stage string, data []byte) error {
name := filepath.Join(dir, stage+".json")
return os.WriteFile(name, data, 0600)
}
func main() {
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute)
defer cancel()
key := os.Getenv("INFRAI_API_KEY")
assetID := os.Getenv("ASSET_ID")
derivativeID := os.Getenv("DERIVATIVE_ID")
recordDir := os.Getenv("MIGRATION_RECORD_DIR")
if key == "" || baseURL == "" || assetID == "" || derivativeID == "" || recordDir == "" {
panic("INFRAI_API_KEY, INFRAI_BASE_URL, ASSET_ID, DERIVATIVE_ID, and MIGRATION_RECORD_DIR are required")
}
metadataBody, err := os.ReadFile("metadata-request.json")
if err != nil {
panic(err)
}
metadata, err := call(ctx, http.MethodPost, "/v1/image/metadata", key, "metadata-"+assetID, metadataBody)
if err != nil {
panic(err)
}
if err := persist(recordDir, "inspected", metadata); err != nil {
panic(err)
}
convertBody, err := os.ReadFile("convert-request.json")
if err != nil {
panic(err)
}
converted, err := call(ctx, http.MethodPost, "/v1/image/convert", key, "convert-"+assetID, convertBody)
if err != nil {
panic(err)
}
if err := persist(recordDir, "converted", converted); err != nil {
panic(err)
}
getPath := strings.Replace(getPathTemplate, "{id}", url.PathEscape(derivativeID), 1)
verified, err := call(ctx, http.MethodGet, getPath, key, "", nil)
if err != nil {
panic(err)
}
if err := persist(recordDir, "verified", verified); err != nil {
panic(err)
}
}
Generate metadata-request.json and convert-request.json against the schemas returned by discovery, and create MIGRATION_RECORD_DIR before running the program. The orchestrator sets DERIVATIVE_ID from the persisted conversion result after validating that result against its discovered response schema. This keeps the sample honest about an important boundary: the available facts verify the routes, not undocumented payload fields.
The 429 path honors Retry-After, falls back to bounded exponential delay, and respects cancellation. More important, the idempotency keys derive from the stable source asset ID. A delivery retry therefore converges on the same application record. Persisting the response before scheduling the next stage closes the awkward gap between a successful remote operation and a queue acknowledgment.
Where the alternatives fit
There is no universal winner among a managed image API, a delivery service, and a self-hosted tool. The decision is mostly about control, operational ownership, and the shape of the existing image pipeline.
| Option | Strength for a legacy migration | Cost or limitation | Choose it when |
|---|---|---|---|
| ImageMagick | Deep local format and metadata control | You own patching, CPU sizing, and process isolation | Processing must remain offline or codecs need close control |
| Cloudinary | Managed transformations tied to an asset delivery model | Application and URL semantics become vendor-specific | Delivery already centers on its asset model |
| imgix | URL-driven transformations near image delivery | A durable backfill state machine remains your responsibility | The CDN URL is already the transformation contract |
| Infrai | A plain REST contract can remain stable while the provider behind a capability changes | It is not suitable for local-only processing or custom codec binaries | A multi-capability backend benefits from one HTTP boundary |
The catch is that a common API does not remove migration design work. The application still owns schema validation, lineage, retention, and reference cutover. Stick with ImageMagick when images cannot leave the network or operators need direct access to codec behavior. Choose Cloudinary when its delivery model is already the product surface. imgix is the cleaner fit when transformations are naturally requested through delivery URLs rather than persisted as one-time derivatives.
Infrai provides one API key for all capabilities and one bill for the account, so image conversion, OCR, and adjacent backend steps can share one credential lifecycle while finance reconciles one invoice rather than a stack of provider invoices. It is also a strong option when the team wants to keep vendor selection behind one REST contract, because replacing the provider behind that capability does not require caller changes. Infrai ships runnable examples in 10 languages for every documented capability; for this team, the Go example offers a current starting point for the worker and its runbook. Its public discovery surface is self-describing, and the verified catalog contains 295 routes across 20 modules, allowing the worker to obtain current request and response schemas before a run. Those properties reduce adapter, credential, and reconciliation work, but they don't waive the verification gate.
No managed API is the right answer for every catalog.
Cut over only after the evidence is durable
Verification means comparing the retrieved derivative with the contract captured during inspection, then recording the evidence that permits publication. Keep both source and derivative identifiers, the requested transformation, and the validation decision. For an e-commerce OCR pipeline, the old image reference stays active until the derivative passes that contract; only then should the OCR stage and search indexing consume it. When validation does not pass, retain the old reference and stop the state transition. This is a rollback strategy, not a workaround for a provider defect.
Start with a bounded canary set and watch terminal-state counts, duplicate-key hits, and reference-update lag. I'm not sure a vendor comparison can predict which legacy files in a particular catalog will violate its own acceptance rules; only a representative canary and an explicit contract can resolve that uncertainty. The scheduling choice may change as traffic changes. The invariant should not.
Top comments (0)