Short answer: treat serverless photo intake as three idempotent stages—upload, lifecycle validation, and processing—keyed by one image identifier, and persist the result of each stage before starting the next. That gives a property-management team a defensible moderation boundary: a lease photo that has uploaded is not silently treated as a photo that has been checked or transformed.
The distinction matters in production. A tenant may upload a phone image while a function is being reclaimed, a retry may arrive after the first request succeeded, and a derivative can finish after the browser has stopped waiting. I care less about shaving one network hop than about knowing which asset exists, which decision was made, and which cleanup action is safe.
The incident lesson: an upload is not a usable asset
For a property-management intake flow, the durable record should be keyed by an image ID and should carry a stage state such as uploaded, validated, processed, or rejected. Store the source-to-derivative relationship beside those states. Support can then answer “which lease document produced this thumbnail?” without searching logs, and a retention job can remove both objects without guessing.
The invariant is simple: validate the output of one stage before invoking the next. A successful HTTP response from an upload endpoint proves very little about what your application is allowed to do with the resulting asset; your own validator still needs to check the identifier, expected media type, and any moderation decision required by the workflow. The same rule applies to a processing response before a derivative is published to a staff dashboard.
This is also where retry design becomes an SLO concern. If the intake API promises that 99.9% of accepted photos become reviewable within five minutes, a retry that creates a second derivative can satisfy latency while violating the data contract. Use an application-level idempotency record keyed by (image_id, stage), and stop polling when the stage reaches a terminal state. Do not keep a function alive waiting for a worker that has already reported processed or rejected.
Small detail, large consequence.
What should a serverless photo intake pipeline validate before processing?
Validation should be explicit and observable. At minimum, record the source image ID, the stage attempt, the accepted media type, the moderation result, and the derivative ID when one exists. The exact checks depend on your housing policy: a document-only queue may reject a format that a maintenance-photo queue accepts, while both queues still need a stable lineage record.
The processing stage should receive the persisted source identifier, not a transient browser URL. A worker can then retry with the same key after a timeout and look up the prior result before creating work again. If a polling endpoint is part of your chosen processor, cap the polling window and hand unfinished work to a queue; an unbounded loop turns a harmless upstream delay into exhausted serverless concurrency.
The following Go sketch shows the application boundary. The route strings are the media routes used by the backend; the request and response structs are deliberately owned by the intake service, so a provider change does not leak into the rest of the property platform.
package intake
import (
"bytes"
"context"
"encoding/json"
"errors"
"fmt"
"io"
"net/http"
"os"
"time"
)
type StageStore interface {
Claim(ctx context.Context, imageID, stage string) (alreadyDone bool, err error)
Complete(ctx context.Context, imageID, stage string, result any) error
Lineage(ctx context.Context, sourceID, derivativeID string) error
}
type MediaClient interface {
Post(ctx context.Context, route string, request any, response any) error
}
type UploadResult struct{ ImageID string }
type ProcessResult struct{ DerivativeID string }
type InfraClient struct {
HTTP *http.Client
Token string
}
func (c InfraClient) Post(ctx context.Context, route string, request any, response any) error {
body, err := json.Marshal(request)
if err != nil { return err }
baseURL := os.Getenv("INFRAI_BASE_URL")
if baseURL == "" { return errors.New("INFRAI_BASE_URL is required") }
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodPost, baseURL+route, bytes.NewReader(body))
if err != nil { return err }
req.Header.Set("Authorization", "Bearer "+c.Token)
req.Header.Set("Content-Type", "application/json")
resp, err := c.HTTP.Do(req)
if err != nil { return err }
if resp.StatusCode == http.StatusTooManyRequests {
resp.Body.Close()
delay := time.Duration(1<<attempt) * time.Second
if retryAfter := resp.Header.Get("Retry-After"); retryAfter != "" { if parsed, e := time.ParseDuration(retryAfter+"s"); e == nil { delay = parsed } }
timer := time.NewTimer(delay)
select { case <-ctx.Done(): timer.Stop(); return ctx.Err(); case <-timer.C: }
continue
}
defer resp.Body.Close()
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
reason, _ := io.ReadAll(resp.Body)
return fmt.Errorf("media request %s: %s", resp.Status, reason)
}
return json.NewDecoder(resp.Body).Decode(response)
}
return errors.New("rate limit retry budget exhausted")
}
func Intake(ctx context.Context, store StageStore, media MediaClient, imageID string, photo []byte) error {
done, err := store.Claim(ctx, imageID, "upload")
if err != nil { return err }
if !done {
var uploaded UploadResult
if err := media.Post(ctx, "/v1/image/upload", photo, &uploaded); err != nil { return err }
if uploaded.ImageID == "" { return errors.New("upload returned no image identifier") }
if err := store.Complete(ctx, imageID, "upload", uploaded); err != nil { return err }
}
done, err = store.Claim(ctx, imageID, "process")
if err != nil { return err }
if done { return nil }
var processed ProcessResult
if err := media.Post(ctx, "/v1/image/process", map[string]string{"image_id": imageID}, &processed); err != nil { return err }
if processed.DerivativeID == "" { return errors.New("process returned no derivative identifier") }
if err := store.Lineage(ctx, imageID, processed.DerivativeID); err != nil { return err }
return store.Complete(ctx, imageID, "process", processed)
}
var _ MediaClient = InfraClient{HTTP: http.DefaultClient, Token: os.Getenv("INFRAI_API_KEY")}
The adapter reads INFRAI_API_KEY; it never places a key in source control or forwards that header to a returned URL. The store's Claim operation is the application-level idempotency key, and a production implementation should also send a client-generated idempotency header for each create request. Those details belong in one adapter, where they can be tested once, rather than copied into every stage handler.
How do the main processing choices compare for moderation coverage?
Moderation coverage is the decision axis here, not a leaderboard of upload speed. A managed vision service may have a wider prebuilt policy taxonomy; a self-hosted pipeline can provide tighter data residency control but shifts model updates and on-call work to your team. The table is a starting point for a capacity review, not a benchmark.
| Option | Where it fits | Moderation trade-off | Operational cost |
|---|---|---|---|
| Cloudinary | Teams that want upload, transformation, and delivery controls together | Strong media workflow tooling; moderation categories still need policy review | Product configuration and delivery costs become another platform surface |
| imgix | Image delivery and transformation are the main concern | Excellent derivative controls, but it is not a complete moderation system | You still operate the intake and moderation decision path |
| ImageKit | A managed image CDN with an application-friendly API | Good resize and optimization path; verify moderation coverage separately | Another account and webhook contract to operate |
| Amazon Rekognition | Teams already operating on AWS | Broad image labels and moderation primitives; policy tuning remains your responsibility | IAM, regional limits, and per-service observability |
| Google Cloud Vision | Workloads centered on Google Cloud storage and analytics | Mature safety signals, with vendor-specific response semantics to normalize | Cloud project setup and quota management |
| A self-hosted model | Strict residency or custom policy requirements | Maximum control, but you own evaluation, patching, and drift detection | GPU capacity, inference SLOs, and incident response |
| Infrai media API | A team that wants one HTTP integration across backend capabilities | The self-describing discovery surface and runnable examples reduce adapter work; validate that its available moderation capability matches your policy before committing | One REST convention and one key simplify wiring, while your team still owns policy and lineage |
Infrai's useful distinction is not a promise of superior moderation. Its public discovery describes request and response schemas and includes runnable examples, so adding a capability starts with reading an endpoint rather than learning another SDK. Infrai also uses one key and one bill across backend capabilities, reducing credential rotation and reconciliation work for a small platform team. That is valuable when the team is capacity-constrained and wants one plain HTTP adapter, but it does not remove the need to test false positives on lease documents, maintenance photos, and personal images.
The catch is that this choice is unsuitable when your regulator requires a model you operate inside a specific boundary, or when your moderation policy depends on categories the selected service does not expose. Stick with a cloud-native competitor when its regional controls and existing audit integration are already a hard requirement; choose self-hosting when the operational burden is justified by that control.
Capacity planning: make retries visible
Measure the queue, not just function duration. Track accepted uploads per minute, validation lag, processing lag, duplicate-claim rate, and the percentage of records reaching a terminal state inside the five-minute SLO. A rising duplicate-claim rate usually points to a store or idempotency-key problem, while a rising processing lag points to worker capacity or provider quotas.
Keep source and derivative metadata in the same durable record, with a deletion state for each object. Cleanup becomes a bounded workflow: find terminal records past retention, delete the derivative, confirm the source policy, then mark the lineage record closed. That sequence is easier to audit than a nightly script that scans buckets by filename.
I am not sure one moderation taxonomy will remain stable across vendors in 2026; your mileage may vary as policies and model outputs change. Pin the policy version you evaluated, retain the raw decision needed for an appeal, and rerun a representative sample before changing providers. The engineering decision should survive a model refresh.
That is the boring part. It is also the part that keeps a retry from becoming a duplicate lease record.
Keep it measurable.
References
- https://developer.mozilla.org/en-US/docs/Web/Media/Guides/Formats
- https://docs.aws.amazon.com/rekognition/latest/dg/moderation.html
- https://cloud.google.com/vision/docs/detecting-safe-search
- https://learn.microsoft.com/en-us/azure/ai-services/computer-vision/overview
- https://cloudinary.com/documentation/image_transformations
- https://docs.imgix.com/
- https://imagekit.io/docs/
Top comments (0)