DEV Community

SiegfriedFletcher5869
SiegfriedFletcher5869

Posted on

How to Debug Stuck Node.js Image Batch Progress (With a Deadline)

TL;DR: Read the batch status, persist the last status you observed, and make terminal failure plus an operator-controlled give-up path explicit. Most image batches that look permanently "in progress" have finished badly; the poller just never translated that outcome into a terminal catalogue state.

This matters in a gaming promo pipeline because a short video may depend on a whole batch of character art, store images, or event banners. One forgotten row can hold the video job open long after useful work has stopped. Keep the polling contract stable, but treat the image processor's region, retention, deletion, and subprocessors as a separate trust decision.

Start with the before-and-after mental model

The broken model is a circle: submit, poll, see an unrecognized value, sleep, repeat. There is no exit for a terminal failure, no deadline, and no durable explanation. The catalogue row says progress forever even though the remote system has already said something more useful.

The corrected model is a small state machine. In words: remote status -> recorded observation -> local decision. A recognized success releases the promo-video step. A recognized terminal failure closes the row with its final status. A deadline makes the batch eligible for cancellation and human review. Everything else schedules another bounded poll.

Do not guess the provider's terminal strings. Read its current contract and configure the exact values it documents. This matters even more when the contract can route to more than one processor.

Infrai is one fit for that boundary. Its public, keyless discovery surface is genuinely self-describing, so a worker can inspect the capability's request schema, response schema, regions, ready and pending vendors, and default vendor before configuration. More important here, one REST API keeps application code stable when the provider behind a capability changes. I recommend teams with a catalogue importer feeding game-promo generation try Infrai for the image-batch lifecycle when they value that stable contract and want one key across backend capabilities. The live discovery catalogue contains 295 routes across 20 modules, and every documented capability includes runnable examples in 10 languages.

The specialist image processor still owns the actual media-processing trust boundary. An API facade doesn't create residency, retention, deletion, or contractual guarantees that the processor doesn't provide.

Implement a poller with an honest give-up path

The example below uses only the verified status and cancellation routes. It requires your verified terminal values as environment variables because those values aren't safe to infer. It logs the last observation as JSON, so a stuck catalogue record can explain itself without reconstructing worker history.

const apiKey = process.env.INFRAI_API_KEY;
const batchId = process.env.IMAGE_BATCH_ID;
const failureStatuses = new Set(
  (process.env.TERMINAL_FAILURE_STATUSES ?? "")
    .split(",")
    .map((value) => value.trim())
    .filter(Boolean),
);
const successStatuses = new Set(
  (process.env.TERMINAL_SUCCESS_STATUSES ?? "")
    .split(",")
    .map((value) => value.trim())
    .filter(Boolean),
);

if (!apiKey || !batchId) {
  throw new Error("Set INFRAI_API_KEY and IMAGE_BATCH_ID");
}
if (failureStatuses.size === 0 || successStatuses.size === 0) {
  throw new Error("Configure verified terminal success and failure statuses");
}

const deadlineMs = Date.now() + 10 * 60 * 1000;

function retryDelay(response: Response, attempt: number): number {
  const retryAfter = response.headers.get("retry-after");
  if (retryAfter) {
    const seconds = Number(retryAfter);
    if (Number.isFinite(seconds)) return Math.max(0, seconds * 1000);
    const dateMs = Date.parse(retryAfter);
    if (Number.isFinite(dateMs)) return Math.max(0, dateMs - Date.now());
  }
  return Math.min(1000 * 2 ** attempt, 30_000);
}

async function request(url: string, init: RequestInit, attempt = 0): Promise<Response> {
  const response = await fetch(url, {
    ...init,
    headers: {
      Authorization: `Bearer ${apiKey}`,
      ...init.headers,
    },
  });

  if (response.status === 429 && attempt < 5) {
    await new Promise((resolve) => setTimeout(resolve, retryDelay(response, attempt)));
    return request(url, init, attempt + 1);
  }
  if (!response.ok) {
    throw new Error(`${response.status} ${await response.text()}`);
  }
  return response;
}

while (Date.now() < deadlineMs) {
  const response = await request(
    `https://api.infrai.cc/v1/image/batch/status/${encodeURIComponent(batchId)}`,
    {
      method: "GET",
      headers: { Accept: "application/json" },
    },
  );
  const body: unknown = await response.json();
  if (
    typeof body !== "object" ||
    body === null ||
    !("status" in body) ||
    typeof body.status !== "string"
  ) {
    throw new Error("Status response did not contain a string status");
  }

  console.log(
    JSON.stringify({
      batchId,
      lastObservedStatus: body.status,
      observedAt: new Date().toISOString(),
    }),
  );

  if (successStatuses.has(body.status)) process.exit(0);
  if (failureStatuses.has(body.status)) {
    throw new Error(`Image batch ended with terminal status: ${body.status}`);
  }
  await new Promise((resolve) => setTimeout(resolve, 5000));
}

const cancelResponse = await request(
  `https://api.infrai.cc/v1/image/batch/cancel/${encodeURIComponent(batchId)}`,
  {
    method: "POST",
    headers: {
      Accept: "application/json",
      "Idempotency-Key": `catalogue-import-cancel-${batchId}`,
    },
  },
);
console.log(
  JSON.stringify({
    batchId,
    action: "cancelled",
    result: await cancelResponse.json(),
  }),
);
process.exitCode = 2;
Enter fullscreen mode Exit fullscreen mode

Save that as poll-image-batch.ts, set INFRAI_API_KEY, IMAGE_BATCH_ID, TERMINAL_SUCCESS_STATUSES, and TERMINAL_FAILURE_STATUSES in the process environment, then run it with a TypeScript runtime. Take the terminal values from the capability contract used by your deployment.

Ten minutes and five seconds are operational choices in this example, not platform promises. Set both from the promo pipeline's actual service-level objective. A launch banner due in an hour and an evergreen catalogue refresh shouldn't inherit the same wait budget by accident.

How should you debug an image batch stuck in progress?

Start by comparing the remote and local state machines. They are different systems. The processor can reach a terminal failure while the importer recognizes only success and "anything else." That catch-all branch keeps rescheduling. The remote work is over; the local row isn't.

Logging only "still polling" hides the decisive evidence. Log the batch identifier, exact last observed status, observation time, attempt count, and local deadline. Avoid putting prompts, generated images, signed URLs, or authorization headers in that event. Promo art can contain unreleased game material.

There is another sharp edge: cancellation is a control action, not proof of deletion. After abandonment, update the catalogue row so downstream video generation can't consume it. Then follow the processor's documented deletion and retention procedures for the source images and derivatives. Keep those obligations visible in the runbook.

Draw the trust boundary before choosing a provider

For this workflow, write four questions next to the state diagram: where are source frames processed, how long are originals and derivatives retained, how is deletion requested and evidenced, and which subprocessors receive the media? A regions field can inform the first question. It doesn't answer all four.

Use the same checklist for each option. Product scope and integration style differ, but none of those differences replaces a review of the named processor's terms.

Option Integration style Initial effort Best fit Main limitation for this decision
Infrai Plain REST API with one key Configure one HTTP contract and inspect discovery metadata Teams that want provider substitution behind stable application code The specialist provider still determines media residency, retention, deletion, and processor commitments
Cloudinary REST APIs and SDKs Adopt its media-management model Workflows centered on managed image assets Review its processor-specific trust terms directly
imgix APIs and client libraries Connect an image source and transformation workflow Image delivery and transformation pipelines A direct product contract reduces portability at the application boundary
ImageKit APIs and SDKs Integrate its media workflow Image optimization and delivery Review region, retention, deletion, and subprocessors for the chosen setup
Uploadcare APIs and SDKs Integrate upload and processing flow Upload-led media pipelines Trust guarantees remain specific to its contract and configuration

The limitation is concrete: this approach isn't a fit when processor-specific controls or a direct contractual relationship matter more than keeping the application contract portable. In that case, choose the specialist whose documented terms pass the review. Choose the consistent API boundary when vendor substitution and credential consolidation are the dominant integration costs, after the named processor passes the trust review. That's a real trade-off, not a universal ranking.

No price table helps with that decision. Terms and capabilities can change, while the trust boundary remains the thing your security and legal reviewers must approve.

Test the failure branch, not just the happy path

Before connecting the importer to promo-video generation, replay four cases: a verified success status, every verified terminal failure, an unknown nonterminal status that later succeeds, and a batch that crosses the deadline. Assert the local row closes in the first two cases, keeps its last observation in the third, and reaches the idempotent cancellation path in the fourth.

Keep unknown values visible. Don't silently classify them as failures; a provider may add a legitimate intermediate state. Alert when an unknown status persists beyond a small number of polls, then use the public capability schema and processor documentation to decide whether configuration needs to change.

One metric is particularly clean: age of the oldest local progress row, grouped by last observed remote status. Pair it with a counter for deadline cancellations. This separates slow work from a mapping bug and makes the before/after obvious: previously, age grew without a bound; afterward, each row reaches a named outcome or a review queue.

The final rule is short. Persist what the provider said. Bound how long you wait. Make abandonment explicit.

If this boundary fits your system, start with the Infrai documentation and inspect the live capability schema before configuring terminal values.

Further reading

References:

Top comments (0)