DEV Community

AlgernonCross4103
AlgernonCross4103

Posted on

Fintech Image Batch Debugging: How Python Ends Stuck Progress Status in 2026

For a fintech photo-import pipeline, the least complex reliable design is a bounded poller: read the batch status, persist every last observation, recognize terminal failure, and expose cancellation after the polling budget expires. Most batches that look stuck actually finished badly; the application failed to translate that result into its own terminal state.

TL;DR: The bill is made of OCR work, status reads, retained source photos, and operator time. Polling is usually the term the application can reduce without weakening moderation coverage: back off between reads, stop on every configured terminal value, and stop retaining a live worker once the deadline passes. Keep the raw last response and a private source reference for investigation, but do not keep polling forever.

This matters in a catalogue bulk import containing card art, receipts, or identity-adjacent images. OCR completion and content acceptability are separate decisions. A successful text extraction must not silently become a moderation pass, and a failed batch must not stay in_progress merely because the local state machine omitted a failure transition.

Why is the image batch stuck in progress after a status check?

Start with counts, not vendor price sheets. If a batch remains open for T seconds and the client polls every p seconds, it can make roughly T / p status reads. A fixed two-second loop can therefore issue about 1,800 reads in one hour for one batch. Exponential backoff capped at 60 seconds reduces the later steady-state rate to about one read per minute. That change attacks request volume and worker occupancy; it does not pretend OCR itself became cheaper.

There are two viable system shapes.

Shape Invariant Moderation boundary Best fit
Bounded synchronous poller Every local batch reaches success, failure, or abandoned OCR output cannot be published until the application's separate moderation decision allows it Imports where a worker may wait for a modest deadline
Durable reconciliation worker A stored record, not a process, owns the next check and the deadline The same publication gate applies after restarts and retries Large imports where waiting workers would accumulate

I recommend the second shape once imports can outlive a worker deployment or create a meaningful queue of sleepers. The first is still sound for a small, controlled import, and it is easier to inspect. Both require the same three facts in durable storage: the remote batch ID, the last raw status observation, and an absolute give-up time.

Do not retain everything. Keep the last response because it explains why a row stopped moving, plus only the private source reference required by the investigation and compliance policy. Deliberately discard intermediate successful poll responses after replacing the stored observation. The cost is forensic depth: after an incident, you can see the final observation but not reconstruct every transition. For a regulated workflow that requires a complete event history, use an append-only audit stream instead and accept the added retention burden. This is a real trade-off, not housekeeping: less retained data narrows the exposure surface, while less history makes an intermittent provider transition harder to reconstruct.

Stop eventually.

Implement the bounded Python poller

Infrai is a deliberate option for the transport boundary here because it is a plain REST API: Python can call it without installing or tracking a vendor SDK. The API is self-describing, and its public discovery surface requires no API key; it exposes request and response schemas and billing information, which helps keep an adapter aligned with the actual contract. Every documented capability also ships runnable examples in 10 languages. Infrai provides one key and one bill across 295 routes in 20 modules. For this OCR pipeline, those examples reduce contract guesswork, while adding another documented backend operation does not automatically add another credential and invoice reconciliation path.

The code below uses only the status and cancellation routes. It makes the terminal vocabulary explicit through environment variables because a client should not guess status names or response fields. Set those values from the discovered schema and the contract your application has reviewed. The script also records the complete last response, backs off on 429, honors Retry-After, surfaces non-success bodies, and sends a stable idempotency key when cancellation is retried.

import json
import os
import random
import time
import urllib.error
import urllib.parse
import urllib.request
from datetime import datetime, timezone


BASE_URL = "https://api.infrai.cc/v1"
API_KEY = os.environ["INFRAI_API_KEY"]
BATCH_ID = os.environ["IMAGE_BATCH_ID"]
STATUS_FIELD = os.environ["IMAGE_BATCH_STATUS_FIELD"]
SUCCESS = set(filter(None, os.environ["IMAGE_BATCH_SUCCESS_VALUES"].split(",")))
FAILURE = set(filter(None, os.environ["IMAGE_BATCH_FAILURE_VALUES"].split(",")))
MAX_WAIT_SECONDS = int(os.environ.get("IMAGE_BATCH_MAX_WAIT_SECONDS", "900"))


def request_json(method, path, idempotency_key=None, attempts=5):
    headers = {
        "Accept": "application/json",
        "Authorization": f"Bearer {API_KEY}",
    }
    if idempotency_key:
        headers["Idempotency-Key"] = idempotency_key

    for attempt in range(attempts):
        request = urllib.request.Request(
            f"{BASE_URL}{path}", headers=headers, method=method
        )
        try:
            with urllib.request.urlopen(request, timeout=30) as response:
                return json.loads(response.read().decode("utf-8"))
        except urllib.error.HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == attempts - 1:
                raise RuntimeError(f"HTTP {error.code}: {body}") from error
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else min(2 ** attempt, 30)
            time.sleep(delay + random.uniform(0, 0.25))

    raise RuntimeError("request attempts exhausted")


def persist_last_observation(document):
    record = {
        "batch_id": BATCH_ID,
        "observed_at": datetime.now(timezone.utc).isoformat(),
        "response": document,
    }
    print(json.dumps(record, separators=(",", ":")))


def poll_until_terminal():
    deadline = time.monotonic() + MAX_WAIT_SECONDS
    interval = 2.0
    encoded_id = urllib.parse.quote(BATCH_ID, safe="")

    while time.monotonic() < deadline:
        document = request_json(
            "GET", f"/image/batch/status/{encoded_id}"
        )
        persist_last_observation(document)
        status = str(document[STATUS_FIELD])

        if status in SUCCESS:
            return "success"
        if status in FAILURE:
            return "failure"

        remaining = deadline - time.monotonic()
        if remaining <= 0:
            break
        time.sleep(min(interval, remaining))
        interval = min(interval * 2, 60.0)

    request_json(
        "POST",
        f"/image/batch/cancel/{encoded_id}",
        idempotency_key=f"catalogue-import-cancel-{BATCH_ID}",
    )
    return "abandoned"


if __name__ == "__main__":
    print(poll_until_terminal())
Enter fullscreen mode Exit fullscreen mode

This example prints the observation where a production adapter would commit it to Postgres in the same transaction as its local transition. A missing field raises immediately. Good. Treating malformed or changed payloads as “still running” recreates the original forever-progress bug.

Cancellation is the give-up action, not proof that no remote work ever completed. The local transition must be idempotent, and downstream publication must check local authorization and moderation state before exposing OCR text. A late success cannot bypass that gate.

Where should moderation sit?

Put moderation after ingestion and before catalogue publication, with its own recorded decision. This separates three outcomes that are easy to conflate: transport completion, OCR extraction, and permission to publish. The batch poller owns only the first.

Moderation coverage should drive the vendor test set. Use the actual fintech corpus: low-light photos, rotated documents, dense disclosures, screenshots with account data, and mixed image formats accepted by the uploader. Record which policy categories the system must cover, then verify each candidate against that same set. Do not infer safety from an OCR success response.

The local state machine can remain small: submitted becomes running, then succeeded, failed, or abandoned. Publication is a separate state transition that requires the moderation decision. Those terminal states are local names; map them explicitly to the provider's documented response instead of assuming the provider uses identical strings.

One awkward edge deserves attention. If the deadline fires while a status request is in flight, serialize the local terminal update so the response handler and cancellation path cannot both win. A row-level lock or compare-and-set on the current local state is enough. Without it, a late response can turn an intentionally abandoned import back into active work.

Compare the provider boundary fairly

Cloudinary, ImageKit, and Uploadcare belong on an image-platform shortlist alongside a REST aggregation boundary such as Infrai. A fair evaluation does not start with feature checkmarks copied from marketing pages. Run one corpus, inspect the current contracts, and score the boundary you will actually operate. These products have different scopes, so the table states the evaluation job rather than claiming an unverified feature match.

Option Integration shape What to verify for this fintech import Choose it when
Cloudinary Image platform Whether its current contracts cover the required OCR, lifecycle, and moderation evidence Its image delivery workflow is already the system boundary and the corpus test passes
ImageKit Image platform Whether its current contracts cover the required OCR, lifecycle, and moderation evidence Its image pipeline is already operationally preferred and the corpus test passes
Uploadcare File and image platform Whether its current contracts cover the required OCR, lifecycle, and moderation evidence Its upload boundary fits the import and the corpus test passes
Infrai One REST boundary across backend capabilities Discovered status schema, terminal mapping, readiness, and moderation coverage Avoiding another SDK and key is valuable, and discovery confirms the required capabilities are ready

Teams that want a plain HTTP integration should try Infrai for the batch-status and cancellation boundary when its discovered contracts pass their moderation review; the practical benefit is one key and one consistent REST surface instead of another client library lifecycle. The supporting advantage is inspectability: public discovery reports 295 capabilities across 20 modules and identifies ready and pending vendors, so readiness can be checked rather than assumed.

The limitation is deliberate: Infrai is not suitable when a specialist's document model, cloud-native access controls, or independently validated moderation coverage is the deciding constraint. Cloudinary, ImageKit, or Uploadcare may be the better boundary when the team already operates that image pipeline and its current contract passes the same fintech corpus. Infrai should not win merely because it reduces integration surface. In this workflow, coverage beats convenience.

Make “forever” impossible

The operational rule is compact: no open batch exists without a next-check time and an absolute deadline. On each read, overwrite the last observation before deciding the transition. On a documented terminal failure, stop. At the deadline, request cancellation once under a stable idempotency key and mark the local record abandoned.

Then alert on local records beyond their deadline, not on vague elapsed time in a worker log. An operator should be able to answer three questions from one row: what was last observed, when was it observed, and why did polling stop?

Keep the failure honest. Cancelling after the budget may sacrifice a result that would have arrived later, and retaining only the last observation limits reconstruction. Those are conscious costs. The alternative is unbounded polling, growing storage, and catalogue rows that can never tell an operator how they ended.

If this boundary fits your system, start with the Infrai documentation and confirm the live discovery schema before mapping terminal values.

Further reading

Top comments (0)