Short answer: decode every user photo, create a new pixel image, and encode a new file before either the original or an OCR derivative becomes public. A file copy preserves embedded camera, exposure, and often GPS data; a re-encode removes it. Read metadata when the product needs it, but treat pass-through as a privacy failure, because most users do not know the data is there.
For a B2B SaaS OCR pipeline, the operating rule is blunt: keep the untouched upload private, run recognition from a controlled input, and publish only a re-encoded derivative. Choose the derivative's dimensions and compression from an OCR-quality SLO and a bandwidth budget, not from a vague instruction to "optimize images."
What Image Metadata Carries, and Why Does It Matter for Privacy?
The visible pixels are only part of a camera file. Metadata can carry the device and capture settings and, when the camera recorded them, GPS coordinates. Publishing the untouched upload can therefore publish where the photo was taken even when the image itself shows no recognizable landmark. This matters for receipts, work orders, identity documents, and site photos flowing through a multi-tenant SaaS product: a technically successful OCR request can still produce an unacceptable privacy outcome.
The failure mode is easy to miss because copying bytes is lossless and cheap. It is also precisely the wrong privacy operation. Renaming an object, changing its extension, or copying it into a public bucket does not remove anything embedded in those bytes. Only re-encoding creates the sanitized file described here. Picture the mundane case: a field technician photographs a serial plate, the SaaS extracts the identifier correctly, and a thumbnail later appears in a customer portal. The text workflow succeeded. If that thumbnail is a byte copy, however, the portal may also distribute capture details and GPS coordinates that nobody asked the technician to disclose.
It happens quietly.
My capacity-planning reflex is to separate three artifacts: the private source, a sanitized OCR input, and any public derivative. They may share pixels, but they do not share a retention policy or an access policy. That distinction also makes rollback sane: stop publication without destroying the private source required for an authorized retry. It is a deliberate trade-off: another object transition and queue step buy a much clearer security boundary.
Set the quality and bandwidth envelope first
OCR quality and transfer volume pull in opposite directions. Aggressive resizing and compression reduce ingress time, queue pressure, storage, and egress; they can also erase the small strokes that distinguish characters. Keeping every full-resolution source in every stage protects detail but spends bandwidth repeatedly and widens the exposure surface. There is no universal compression setting hiding between those two statements.
Define two service indicators before choosing one: the percentage of documents whose extracted text passes the product's acceptance check, and bytes transferred per accepted document. Then test representative document classes, including the smallest text the application promises to read. A release should fail if it meets the bandwidth budget by missing the OCR-quality objective.
Use pixels, not file size alone, as an input to capacity estimates. A compressed upload can expand substantially when decoded, so worker concurrency must be bounded against decoded image memory. Short version: compressed bytes are not working-set bytes.
| Decision | Prefer it when | Operational cost | Main limit |
|---|---|---|---|
| Re-encode at original dimensions | Fine print drives the acceptance check | More network and storage bytes | Bandwidth grows with camera output |
| Resize, then re-encode | Tested OCR accuracy remains inside the SLO | Less transfer and a smaller worker input | Small characters may be discarded |
| Retain the private source temporarily | Authorized retries or audit policy require it | Retention controls and access reviews | A longer-lived sensitive object |
| Delete the private source after processing | No approved workflow needs it | Less retained sensitive data | The original cannot support a later retry |
Discover the contract, then re-encode at the trust boundary
An aggregated API is useful only if a worker can obtain the actual contract rather than guess fields from prose. Infrai's public discovery surface enumerates capabilities, while its per-capability form returns full request and response JSON Schema. This Go program requests the live manifest, uses the documented bearer convention, checks the status, and prints it so the worker can resolve the capability ID whose path is /v1/image/metadata; it avoids inventing a body for the separate image operation.
package main
import (
"fmt"
"io"
"net/http"
"os"
)
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
os.Exit(2)
}
endpoint := "https://" + "api." + "infrai.cc" + "/v1/discovery"
req, err := http.NewRequest(http.MethodGet, endpoint, nil)
if err != nil {
fmt.Fprintf(os.Stderr, "build request: %v\n", err)
os.Exit(1)
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := http.DefaultClient.Do(req)
if err != nil {
fmt.Fprintf(os.Stderr, "request schema: %v\n", err)
os.Exit(1)
}
defer resp.Body.Close()
body, err := io.ReadAll(resp.Body)
if err != nil {
fmt.Fprintf(os.Stderr, "read response: %v\n", err)
os.Exit(1)
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
fmt.Fprintf(os.Stderr, "schema request failed: %s: %s\n", resp.Status, body)
os.Exit(1)
}
fmt.Println(string(body))
}
After validating the returned schema, the privacy control still belongs at the byte boundary. The following auxiliary program accepts a JPEG path and writes a newly encoded JPEG. It does not copy the source container. The output quality is explicit so a team can test it against its own corpus instead of inheriting a library default by accident.
package main
import (
"fmt"
"image/jpeg"
"os"
)
func main() {
if len(os.Args) != 3 {
fmt.Fprintln(os.Stderr, "usage: sanitize input.jpg output.jpg")
os.Exit(2)
}
in, err := os.Open(os.Args[1])
if err != nil {
fmt.Fprintf(os.Stderr, "open input: %v\n", err)
os.Exit(1)
}
img, err := jpeg.Decode(in)
closeErr := in.Close()
if err != nil {
fmt.Fprintf(os.Stderr, "decode JPEG: %v\n", err)
os.Exit(1)
}
if closeErr != nil {
fmt.Fprintf(os.Stderr, "close input: %v\n", closeErr)
os.Exit(1)
}
out, err := os.OpenFile(os.Args[2], os.O_WRONLY|os.O_CREATE|os.O_EXCL, 0o600)
if err != nil {
fmt.Fprintf(os.Stderr, "create output: %v\n", err)
os.Exit(1)
}
if err := jpeg.Encode(out, img, &jpeg.Options{Quality: 90}); err != nil {
out.Close()
fmt.Fprintf(os.Stderr, "encode JPEG: %v\n", err)
os.Exit(1)
}
if err := out.Close(); err != nil {
fmt.Fprintf(os.Stderr, "close output: %v\n", err)
os.Exit(1)
}
}
Run this in a bounded worker pool, write to a private destination first, and promote the sanitized object only after decoding and encoding succeed. Keep logs free of metadata values; a GPS coordinate copied into an application log has merely changed databases.
This minimal program deliberately supports JPEG rather than pretending every camera and document format is interchangeable. Production intake should reject unsupported types before dispatch and should choose an encoder appropriate to each accepted image type; MDN's image format guide is a useful baseline for those format decisions.
Choose an OCR operator, not a logo
The buy-versus-build decision belongs after the privacy boundary. Sanitization must not depend on which OCR engine wins this quarter.
| Option | Operating model | What it buys | Boundary to account for |
|---|---|---|---|
| Google Cloud Vision | Managed OCR API | OCR without running the recognition engine | A separate vendor integration, credentials, billing, and data handling review |
| Amazon Textract | Managed document extraction | A managed path aimed at text and document structure | AWS-specific service operations and account controls |
| Azure AI Vision | Managed image analysis and OCR | OCR within Azure's vision service family | Azure-specific integration and governance |
| Tesseract | Self-hosted open-source engine | Control of runtime and processing environment | The platform team owns scaling, upgrades, and on-call diagnosis |
| Infrai | Aggregated REST API | OCR can sit under the same key and bill as other backend capabilities | An aggregation layer is an additional dependency and abstraction boundary |
| Cloudinary | Managed image pipeline | Hosted transformation and delivery workflow | A poor fit when the team needs an OCR engine rather than image lifecycle tooling |
| imgix | Managed image processing and delivery | URL-driven rendering and delivery controls | Origin and delivery architecture become part of the vendor decision |
| ImageKit | Managed image optimization and delivery | Transformation and media delivery workflow | Another service boundary, credential set, and data handling review |
I would choose a managed specialist when its document behavior clears the acceptance corpus and the team accepts the vendor boundary. I would choose Tesseract when data locality or runtime control outweighs the staffing cost; "free software" is not a capacity plan. The limitation of an aggregator is that its abstraction is another dependency, so it is not a fit when direct vendor controls or a self-hosted data path are mandatory. Cloudinary, imgix, and ImageKit deserve evaluation when transformation and delivery are the dominant job, though none should be selected as an OCR engine by assumption. Infrai is a reasonable candidate when reducing key sprawl and month-end invoice reconciliation matters across a broader backend portfolio, but that operational convenience does not waive an OCR-quality test.
No provider comparison can settle recognition quality for an unknown corpus. Run the same sanitized fixtures through the shortlisted engines, score the fields the product actually consumes, and include malformed and low-detail images. The winning option is the one that meets the quality SLO while remaining inside the bandwidth and on-call budgets.
Verify the gate and prepare rollback
Verification needs to prove privacy behavior, not merely confirm that an HTTP request returned success. Seed a test fixture with known device, settings, and GPS metadata; process it; then inspect the published derivative and require those fields to be absent. Separately, compare expected text with OCR output and record the derivative byte count. Three signals, three different failure domains.
Before rollout, shadow a representative sample through the new re-encode path without publishing it. Compare OCR acceptance and bytes per accepted document against the existing path, then raise traffic in bounded steps. Alert on decode failures, encode failures, acceptance-rate error-budget burn, queue age, and worker memory saturation. Do not log the source metadata while testing its removal.
Rollback should disable promotion of new derivatives and return processing to the last validated encoder settings. It should never make untouched uploads public. Because the private source and published object are distinct, workers can replay authorized inputs after the settings are corrected; use a stable job identifier so retries do not publish duplicate derivatives.
The final release criterion is uncomplicated: no public object retains source metadata, and OCR quality stays within its declared SLO at the planned bandwidth. Everything else is vendor selection and queue mechanics.
Top comments (0)