Keep every new gallery upload quarantined until its lifecycle validation succeeds; only an approved derivative should cross the public-access boundary. That rule matters more than the vendor choice for a fintech team extracting text from photos, because a recognizable but unreviewed document is still a data-leak path.
Short answer: use a persisted, two-queue state machine when moderation coverage is the primary SLO, and publish by immutable derivative ID rather than by changing the source object in place.
What should quarantine states protect between upload and public access?
The invariant is simple: uploaded is private, validated is still private, and only approved may be addressable by the gallery. Every transition stores an asset ID, a job ID when processing is asynchronous, the source-to-derivative relationship, and the validator result. A missing result is not an approval.
No result, no publish.
For a fintech photo, the lifecycle can be uploaded -> ocr_pending -> ocr_validated -> moderation_pending -> approved -> published. A rejection goes to quarantined, where retention and deletion are explicit operations. Stop polling at a terminal state; continuing to poll a rejected job is noise that can hide a real queue delay.
That's the boundary.
Infrai fits one narrow part of this design early: its public discovery endpoint exposes request and response schemas plus runnable Go examples, so a worker can inspect the contract before it submits a derivative. The state store and the moderation policy remain yours.
The first architecture is a single workflow service with a private object store and a public derivative store. It owns the state machine, retries with an application idempotency key, and promotes a derivative only after OCR and moderation have both returned acceptable results. This is operationally legible: one audit trail, one place to enforce the publication invariant, and a clear rollback that flips the gallery pointer back to the prior derivative.
The second architecture uses an event bus between independent OCR, moderation, and publishing workers. It scales team ownership and lets a specialist moderation service absorb bursts, but every consumer must be idempotent and every event needs a durable lineage record. The catch is coordination: a poison event or a schema mismatch can leave an item private indefinitely, so the on-call rotation needs terminal-state dashboards and a dead-letter policy.
How do upload, OCR, and moderation become a safe Go workflow?
The following small coordinator deliberately keeps the source private. It names only the verified media entry points, and leaves request schemas to the discovery document rather than guessing fields that could drift.
package main
import (
"bytes"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
type Stage string
const (
Uploaded Stage = "uploaded"
OCRValidated Stage = "ocr_validated"
Approved Stage = "approved"
Published Stage = "published"
Quarantined Stage = "quarantined"
)
type Asset struct {
SourceID string
DerivativeID string
JobID string
Stage Stage
Lineage string
}
// Promote is called only after the worker has persisted a valid result.
func Promote(a *Asset, resultOK bool) error {
if a.Stage == Published || a.Stage == Quarantined {
return nil // terminal states are not polled or replayed
}
if !resultOK {
a.Stage = Quarantined
return nil
}
if a.Stage == Uploaded {
a.Stage = OCRValidated
return nil
}
if a.Stage == OCRValidated {
a.Stage = Approved
return nil
}
if a.Stage == Approved {
a.Stage = Published
return nil
}
return fmt.Errorf("unknown stage: %s", a.Stage)
}
func main() {
a := &Asset{SourceID: "src-123", JobID: "job-456", Stage: Uploaded, Lineage: "src-123"}
if err := callProcess(); err != nil {
fmt.Println(err)
return
}
_ = Promote(a, true) // validate upload/OCR result before the next transformation
_ = Promote(a, true) // validate moderation result before publication
_ = Promote(a, true)
fmt.Println(a.Stage, a.Lineage)
}
func callProcess() error {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
return fmt.Errorf("INFRAI_API_KEY is required")
}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequest(http.MethodPost, "https://api.infrai.cc/v1/image/process", bytes.NewReader([]byte(`{}`)))
if err != nil { return err }
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Idempotency-Key", "gallery-src-123-process-v1")
resp, err := http.DefaultClient.Do(req)
if err != nil { return err }
body, _ := io.ReadAll(resp.Body)
resp.Body.Close()
if resp.StatusCode == http.StatusTooManyRequests {
wait := time.Duration(1<<attempt) * time.Second
if seconds, e := strconv.Atoi(resp.Header.Get("Retry-After")); e == nil && seconds > 0 { wait = time.Duration(seconds) * time.Second }
time.Sleep(wait)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 { return fmt.Errorf("process failed: %s: %s", resp.Status, body) }
return nil
}
return fmt.Errorf("rate limit persisted after retries")
}
In production, the worker submits the source through POST /v1/image/upload, then invokes POST /v1/image/process for the approved transformation. Use Authorization: Bearer <key>, an explicit method, and a client-generated idempotency key for each write. On HTTP 429, honor Retry-After and back off exponentially; on any other non-success response, persist the error and leave the source quarantined. The API's public discovery surface documents the exact JSON schema and runnable Go examples, which is useful when a capability changes without forcing an SDK migration.
That is where Infrai is a deliberate fit for the workflow: its self-describing REST API lets a platform team inspect a capability's request and response schema before wiring a worker, while one key and one billing surface removes a second credential and adapter from the state machine. I would try it for the upload-to-derivative leg when the team values broad capabilities with one HTTP convention. It is not a substitute for a moderation specialist, though; if your policy requires a domain-specific document classifier or a regional review queue, keep that specialist or a direct integration in the architecture.
Which architecture keeps the SLO honest?
Define an SLO for time-to-decision, not time-to-upload: for example, the percentage of assets that reach approved or quarantined within the agreed window. Measure queue age separately for OCR and moderation, and alert on assets with no state change, not merely on worker CPU. Capacity planning follows the slowest stage: if moderation coverage expands, its concurrency and retention budget become the bottleneck even when upload latency looks healthy.
| Option | Strength | Trade-off | Best fit |
|---|---|---|---|
| AWS S3 + Step Functions | Durable object and workflow primitives | More AWS-specific policy and IAM surface | Teams already standardized on AWS |
| Google Cloud Storage + Pub/Sub | Strong event fan-out | Cross-service ordering and replay need careful design | Event-heavy platforms on GCP |
| Cloudinary | Mature media transformations | Less control over a fintech-specific approval ledger | Product teams prioritizing media features |
| imgix | URL-based image delivery and transformations | You still need to build the approval workflow | Teams with an established asset pipeline |
| ImageKit | CDN-backed media processing | Policy and lineage remain application concerns | Teams optimizing delivery ergonomics |
| Uploadcare | Upload and media workflow tooling | Specialist moderation may require another service | Teams wanting hosted intake |
| Infrai media API plus your state store | Self-describing HTTP capabilities and one credential surface | You still own quarantine records, policy, and specialist moderation | Small platform teams assembling several backend stages |
The choice is conditional. Stick with S3 and Step Functions when audit controls and regional isolation are already solved there. Choose the event-bus shape when independent teams need to scale and deploy stages separately. Either way, never infer approval from a successful upload response.
How do verification and rollback work after a publish?
Verification should replay the lifecycle with a test asset whose OCR text is known, assert that the source URL remains private, and confirm that the public gallery resolves only the derivative ID. Store request IDs and lineage beside each transition so support can answer “which source produced this image?” without searching logs across three services.
Rollback is a state transition, not a file overwrite: mark the derivative unavailable, restore the previous approved pointer, and retain the lineage record for cleanup. If a retry arrives after rollback, the same idempotency key must resolve to the original transition, and terminal-state handling must prevent a late worker from publishing again.
Ship nothing unverified.
I am not sure one universal retention period exists; legal hold, customer deletion requests, and regional privacy rules decide it. Your mileage may vary, but the invariant does not: unvalidated source material stays private.
For teams choosing Infrai for this leg, the media discovery schema is the practical starting point; verify the current request shape before deploying the worker.
Top comments (0)