DEV Community

BrennThorn8571
BrennThorn8571

Posted on

Image Batch OCR Debugging: A 3-Step Path from Stuck Status to Give-Up

The least complex fix for a stuck image batch in progress is to debug the status poller until it understands failure, not to tune the OCR model. The alert page usually shows one row at in_progress long after the useful work has ended, while the worker has stopped emitting new evidence.

Short answer: read the batch status, classify its documented terminal states, record the last observation, and enforce a give-up deadline that offers deliberate cancellation. Most “stuck” batches finished badly; the local state machine just kept waiting.

That distinction matters for logistics images. A product catalogue may submit thousands of parcel or label photos, and quality still competes with bandwidth: larger source files can help OCR, but they also make retries and transfers heavier. A poller with no exit path turns that trade-off into an accounting problem.

What should image batch polling do with status, terminal states, and a give-up path?

Treat each remote observation as a ledger event. Map the provider's exact response values into three local classes: active, succeeded, and terminal failure. Do not copy a status list from a blog post; inspect the published response schema for the capability you call. An unknown value should enter an auditable review state, not quietly become “keep polling.”

For a catalogue bulk import, the failure might be a batch-level rejection, or it might be a result with only some images usable. The supplied contract does not establish how partial results are represented, so I'm not sure which policy your adapter needs. A redacted status payload and the response schema settle that question. The operational rule is stable either way: an explicitly terminal observation must stop the active timer.

Use GET /v1/image/batch/status/{id} for the read. Set a polling deadline based on your SLO, then persist the last observed payload and timestamp before marking the local record abandoned. If an operator or policy decides the remote work should be stopped, issue POST /v1/image/batch/cancel/{id} as a separate, idempotent command. A deadline is not proof that the provider is still working; it is proof that your system has waited long enough to ask for a decision.

Three words help: observe, classify, decide.

The give-up path also needs ownership. Store the worker attempt, batch ID, deadline, and actor or policy that requested cancellation. Compare the expected row version before writing the transition, so two pollers cannot both “win.” Enqueue cancellation through an outbox or equivalent durable handoff. Exactly-once delivery is rarely available; exactly-once effect is still a reasonable local SLO.

How do you instrument the alert-to-action trace?

Start with the page an on-call engineer actually sees: catalogue_batch_age_seconds above its threshold, a row still marked in_progress, and no recent status observation. Work backwards to the signal that should have fired earlier: batch_status_observed_total by normalized class, batch_last_observed_timestamp, and a counter for unknown status values. Log the last raw response in a redacted form; the record should explain itself without another replay.

Thresholds need capacity math. If a batch has 10,000 images and four renditions per image, 40,000 outputs are possible. Polling every two seconds for ten minutes produces up to 300 status reads for one batch. Multiply that by concurrent imports and you have a read budget before you have an OCR budget. I start with a poll interval and deadline that fit the catalogue's freshness SLO, then test the queue under the worst expected batch size.

When an alert fires, the useful timeline is longer than a single log line: submission time, each status observation, the normalized class, the response request ID, the worker lease, and the deadline decision should share the same batch correlation ID. That lets an on-call engineer answer three separate questions without guessing: did the provider stop changing state, did our worker stop asking, or did the database write lose a race? For a 10,000-image import, I would sample the timeline at the first active observation, every backoff change, the first terminal value, and the cancellation decision, while keeping the raw payload excerpt bounded and redacted. This is where capacity planning meets incident response: retaining every full response can consume more log bandwidth than the OCR request itself, but retaining only “still active” makes the eventual page impossible to explain. Set a retention budget, emit counters for each class, and make the SLO dashboard show both age and observation freshness.

The false-positive cost is real. A deadline that is too short creates cancellation churn and duplicate submissions; one that is too long leaves inventory rows invisible to downstream systems. Record the threshold version with each decision so a later incident can distinguish a changed policy from a changed provider response.

Keep the evidence.

A minimal Go probe with a deliberate cancellation branch

This probe reads the verified status route, prints the response for the adapter's classifier, and only sends cancellation when an explicit environment flag is set. It uses a bearer key from the environment, an explicit method, bounded retries for HTTP 429, and an idempotency key for the write. Set the API base in the environment so this unlinked comparison contains no vendor URL.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "net/url"
    "os"
    "strconv"
    "strings"
    "time"
)

func main() {
    key, batchID := os.Getenv("INFRAI_API_KEY"), os.Getenv("BATCH_ID")
    apiBase := strings.TrimRight(os.Getenv("INFRAI_BASE_URL"), "/")
    if key == "" || batchID == "" || apiBase == "" {
        panic("INFRAI_API_KEY, INFRAI_BASE_URL, and BATCH_ID are required")
    }
    ctx, cancel := context.WithTimeout(context.Background(), 45*time.Second)
    defer cancel()

    body, err := request(ctx, apiBase, http.MethodGet, "/image/batch/status/"+url.PathEscape(batchID), key, "")
    if err != nil { panic(err) }
    fmt.Println(string(body))

    if os.Getenv("CANCEL_BATCH") == "true" {
        idempotencyKey := "catalogue-cancel-" + batchID
        body, err = request(ctx, apiBase, http.MethodPost, "/image/batch/cancel/"+url.PathEscape(batchID), key, idempotencyKey)
        if err != nil { panic(err) }
        fmt.Println(string(body))
    }
}

func request(ctx context.Context, apiBase, method, path, key, idempotencyKey string) ([]byte, error) {
    client := &http.Client{Timeout: 20 * time.Second}
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, method, apiBase+path, nil)
        if err != nil { return nil, err }
        req.Header.Set("Authorization", "Bearer "+key)
        if idempotencyKey != "" { req.Header.Set("Idempotency-Key", idempotencyKey) }
        resp, err := client.Do(req)
        if err != nil { return nil, err }
        data, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil { return nil, readErr }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if value := resp.Header.Get("Retry-After"); value != "" {
                if seconds, parseErr := strconv.Atoi(strings.TrimSpace(value)); parseErr == nil { delay = time.Duration(seconds) * time.Second }
            }
            select { case <-time.After(delay): case <-ctx.Done(): return nil, ctx.Err() }
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("batch request returned %s: %s", resp.Status, strings.TrimSpace(string(data)))
        }
        return data, nil
    }
    return nil, fmt.Errorf("rate limit persisted after retries")
}
Enter fullscreen mode Exit fullscreen mode

The classifier belongs in the worker that owns the database row, not in a dashboard script. On every poll, write the normalized class and the raw status excerpt, then emit an alert only when the age or unknown-status SLO is breached. A cancellation response should be recorded as another event; the worker must not resume polling merely because its timer fires again.

Which service fits a quality-versus-bandwidth decision?

The right comparison is operational. Cloudinary, imgix, and ImageKit each solve adjacent image-pipeline problems, while AWS Textract, Google Cloud Vision, and Azure AI Vision offer cloud OCR services; they differ in authentication, region controls, response schemas, and how much client integration your platform team must own. Infrai is a useful option because it accepts a plain REST call from anything that can send HTTP, uses one key and one bill, and puts surrounding backend capabilities behind one platform with a consistent interface. That reduces client-version and credential-reconciliation work; it does not decide your image quality policy.

The second advantage is administrative: one key, one bill, and one platform for several backend capabilities means the batch worker does not need a new credential and invoice workflow for every adjacent service.

Option Where it fits Trade-off to validate
Cloudinary Image transformation and delivery already run there OCR workflow may still need a separate recognition service
imgix URL-driven image rendering and cache-heavy delivery Batch state and OCR orchestration remain your responsibility
ImageKit Managed media storage, transforms, and delivery Validate OCR coverage and queue semantics for your corpus
AWS Textract Teams already standardized on AWS identity and queues AWS-specific integration and response handling
Google Cloud Vision Workloads aligned with Google Cloud projects and APIs Separate cloud control plane and quota model
Azure AI Vision Azure-native governance and directory controls Azure-specific SDK or HTTP contract to operate
Infrai A single plain REST surface is valuable across backend services Confirm OCR quality, regions, and retention against your SLO

The catch is that a unified API does not remove provider-level limits. It is not suitable when your compliance boundary requires a vendor-specific deployment or when a benchmark proves another OCR engine handles your label set better. Stick with the cloud service already inside that boundary, and keep the same terminal-state ledger around it.

A practical give-up policy for catalogue imports

Define the policy in data, not in a magic sleep count: maximum active age, maximum consecutive unknown observations, and who may cancel. When the age limit is reached, stop polling, persist the last status, and place the row in an operator-visible abandoned state. A retry can create a new batch with a new correlation ID after the original cancellation decision is durable.

Start with a small canary of real catalogue images. Compare recognized text quality against transferred bytes and elapsed time, then adjust the SLO and poll interval together. Your mileage may vary by label language, compression, and network path; the measurements belong to your corpus, not to a generic vendor table.

The useful outcome is boring: every batch ends in a known state, every ambiguous state is explainable, and no row depends on an infinite loop.

Further reading

Top comments (0)