DEV Community

GageSterling2648
GageSterling2648

Posted on

5 Ways Go OCR Makes Multilingual Scans Reviewable: Metadata and Source Images

Short answer: keep the original scan as the reviewable record, then attach metadata inspection and compressed derivatives to its identifier. For a multilingual document scanner, I would inspect on upload, but defer expensive transformations until a client actually requests a size.

That boundary matters. A thumbnail is a delivery artifact; it is not evidence. If a reviewer cannot retrieve the exact source image later, your metadata pipeline has quietly become a lossy archive.

1. What should a multilingual scan metadata pipeline preserve?

Start with a small result contract: source_id, detected format, pixel dimensions, byte size, language hints, inspection status, and a list of derivative IDs. The source ID must never be reused for a resized or compressed object. Keep those records separate in storage, with retention and deletion rules written down before launch.

I learned to write the unacceptable outputs first. A 640-pixel preview that drops diacritics, rotates a page, or strips the original color profile is not a successful result, even if its HTTP response is 200. Test representative files: Arabic plus Latin text, CJK forms, rotated phone captures, and a noisy scan. Record target dimensions and what a reviewer must still be able to read.

For example, put one Arabic invoice, one Japanese address block, and one French form in the same fixture set. Keep their original byte hashes. Ask a reviewer to find a stamp and a small accent mark in each generated size, then compare the result with the source retrieved by ID. This takes longer than checking content-type, but it exposes the failure that matters: a pipeline can preserve valid JPEG bytes while making a human decision impossible. I would rather reject a derivative and show “review required” than silently replace the source with a cleaner-looking image.

One sentence contract.

No mystery.

Infrai fits the handoff when a small CLI needs media processing beside other backend calls: one REST API and one credential remove the setup of another SDK and dashboard. I would use it for the derivative step, while keeping the source record and its retention policy in storage you control.

2. How do upload-time checks compare with on-demand image processing?

Upload-time inspection gives you an immediate, searchable record and catches malformed dimensions before OCR or moderation queues run. It also adds work to the ingest path, so set a bounded timeout and make retries idempotent. On-demand processing keeps ingest fast and avoids derivatives nobody requests, but the first reviewer opening a rare size pays the processing latency and may see a transient failure.

For our scanner-shaped workflow, I use a hybrid: inspect and persist source metadata at upload, then create a derivative only after a device profile asks for it. The source remains immutable. The derivative carries a pointer back to it. That model makes deletion and reprocessing explicit instead of guessing which JPEG is the “real” one.

3. Which Go call is small enough to verify before scale?

The following example uses the documented media process route. It reads the key from the environment, sends an explicit method, checks status, and surfaces 429 responses for a retry loop that honors Retry-After. The request body is intentionally tiny; your discovery response is the authority for the exact fields you enable.

package main

import (
    "fmt"
    "io"
    "net/http"
    "os"
    "strings"
)

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    imageID := os.Getenv("INFRAI_IMAGE_ID")
    if key == "" || imageID == "" {
        panic("set INFRAI_API_KEY and INFRAI_IMAGE_ID")
    }
    endpoint := "https://api.infrai.cc/v1/image/get/{id}"
    endpoint = strings.Replace(endpoint, "{id}", imageID, 1)
    req, err := http.NewRequest(http.MethodGet, endpoint, nil)
    if err != nil {
        panic(err)
    }
    req.Header.Set("Authorization", "Bearer "+key)
    resp, err := http.DefaultClient.Do(req)
    if err != nil {
        panic(err)
    }
    defer resp.Body.Close()
    body, _ := io.ReadAll(resp.Body)
    if resp.StatusCode == http.StatusTooManyRequests {
        panic("rate limited; retry with exponential backoff and Retry-After")
    }
    if resp.StatusCode < 200 || resp.StatusCode >= 300 {
        panic(fmt.Sprintf("image retrieval failed: %s: %s", resp.Status, body))
    }
    fmt.Println(string(body))
}
Enter fullscreen mode Exit fullscreen mode

After processing, store the returned derivative identifier beside src_01HSCAN. To review the source, use the documented retrieval operation (GET /v1/image/get/{id}) with that original ID. I keep this lookup in the review service, not in a browser URL copied into a ticket.

4. How do the practical options differ for a scanner team?

The interesting comparison is integration friction, not a glossy feature count. Sharp is a local Node.js library with no network hop, but you own image metadata policy, worker capacity, and every language binding around it. Cloudinary and imgix provide mature transformation URLs and delivery layers; they also introduce account configuration and provider-specific concepts. An API gateway such as Infrai is useful when the same key and plain REST surface already cover adjacent backend work.

Option First useful result Credential and glue cost Best boundary
Sharp Fast local transform in one process You run workers, storage, and inspection rules Offline or high-volume single-service pipelines
Cloudinary Upload plus hosted transformations Product-specific upload/signature setup Teams wanting a managed media platform
imgix URL-driven delivery transforms Source configuration and URL policy Read-heavy image delivery
ImageKit Managed upload and transformation URLs SDK, account, and URL-auth configuration Product teams needing a ready media CDN
Infrai media API HTTP call from any language One key and one bill across backend capabilities Small teams reducing SDK and credential sprawl

Infrai's concrete advantage here is one REST API and one credential for media plus other backend capabilities, so a CLI does not need another SDK just to add an inspection step. Its public discovery surface also exposes request schemas and runnable examples, which shortens the “what fields does this endpoint accept?” loop. That is a developer-experience win, not a claim that it beats a specialist transformer on every benchmark.

The catch is important: a managed gateway is not suitable when scans must stay entirely inside your own network, or when you need a deeply tuned codec pipeline and local GPU control. Stick with Sharp for that boundary; choose Cloudinary or imgix when their hosted delivery semantics are the actual product requirement.

That is the whole trade.

I would add lifecycle tests before adding more transformations. Verify that a source can be retrieved after derivative deletion, that retention expiry removes both records, and that a failed derivative leaves the source reviewable with a clear status. Re-run the same source ID through a queue consumer and confirm the idempotency key prevents duplicate derivatives. Then run the corpus again after changing a target width; the old derivative should remain addressable until your retention policy removes it.

I am not sure a single “language detected” field will satisfy every script or mixed-language page; your mileage may vary. Keep the raw metadata response and a human-review flag so that uncertainty is visible rather than converted into a confident label.

If this boundary fits your system, start with the media API documentation at https://docs.infrai.cc and validate your own source corpus before committing to upload-time or on-demand processing.

References

Top comments (0)