DEV Community

Snap AI
Snap AI

Posted on Fully Autonomous

How we put 41 AI image and video models behind one credit ledger

Calling an image or video model is the easy part of a multi-model AI app. Every gateway has a decent SDK and the happy path takes an afternoon. The parts that took real thought in Snap AI Studio, a web studio that puts 41 image, video and audio models on one credit balance, were money and state:

  • take the credits before a job runs, without letting two tabs spend the same balance
  • give them back exactly once when a provider fails, no matter how many code paths notice the failure
  • finish jobs that nobody is polling, because the user closed the tab
  • make Stripe webhooks safe to replay

This is how each of those works in a Next.js 16 app with Prisma 7 and Postgres. Nothing here is specific to AI; any app that sells metered work to a third party has the same four problems.

One interface, several gateways

Each model in the registry declares which gateways can serve it. OpenRouter handles most image and video models; fal covers what OpenRouter does not expose (lip sync, text to speech, music, upscaling, background removal, face swap). A mock provider implements the same interface and returns placeholder images and sample clips, so the whole app runs with zero API keys. That last part matters more than it sounds: every refund path below can be tested locally by putting a magic word in the prompt that makes the mock fail on purpose.

For a job, the gateway named by MODEL_PROVIDER is tried first, then any other gateway that has a key and a verified id for that model.

Debit first, in one transaction

A job row and the debit are created in the same transaction, before anything is sent to a provider. The debit is a conditional update, so the balance check and the decrement are one statement:

const updated = await tx.user.updateMany({
    where: { id: userId, credits: { gte: amount } },
    data: { credits: { decrement: amount } },
});
const user = await tx.user.findUniqueOrThrow({ where: { id: userId }, select: { credits: true } });
if (updated.count === 0) throw new InsufficientCreditsError(amount, user.credits);
await tx.creditLedger.create({
    data: { userId, delta: -amount, balanceAfter: user.credits, reason, ...extra },
});
Enter fullscreen mode Exit fullscreen mode

Two requests racing for the last credits cannot both pass, because Postgres evaluates credits >= amount and applies the decrement atomically for each row update. Every movement also writes a ledger row with the balance after it, so the balance can always be explained line by line.

Only after the transaction commits does the app call provider.submit. If submit throws, the job is failed and refunded through the same path as any other failure.

Refund exactly once

Failures arrive from several places: the submit call, the request that polls the job, the background sweeper, a timeout. Any two of them can notice the same failure at the same moment. The fix is to make "fail this job" a claim that only one caller can win:

const claimed = await tx.job.updateMany({
    where: { id: jobId, refunded: false, status: { in: ["queued", "running"] } },
    data: { status: "failed", error: message, refunded: refund, finishedAt: new Date() },
});
if (claimed.count === 0) return null;
if (refund && job.cost > 0)
    await grantCredits(tx, job.userId, job.cost, "job_refund", { jobId, note: "Refund for failed generation" });
Enter fullscreen mode Exit fullscreen mode

The where clause is the whole trick. The first caller moves the job out of queued/running and the refund happens inside the same transaction. Every later caller matches zero rows and returns. No locks, no flags in memory, and it survives a restart halfway through.

Jobs nobody is watching

The polling endpoint advances a job when the browser asks for it: it calls the provider, copies outputs to storage on success, and refunds on failure. But users close tabs. A background sweeper started from instrumentation.ts runs every 20 seconds and advances any job that is not terminal. Jobs older than 30 minutes are failed and refunded through the same claim as above, so a provider that never answers cannot hold credits forever.

Sync images, async video

OpenRouter's video API is asynchronous: create the job, poll its status, then download the content with the API key. Images are different: the image endpoint is synchronous, and one image can take over a minute. Holding an HTTP request open that long breaks the "submit returns an id" contract the rest of the pipeline relies on, so submit starts the image request in the background and returns an id at once; the poller reads the result from an in-process map.

That map is the one place the design assumes a single app instance. If the app ever scales out, image tasks move to a queue. Writing the assumption down next to the code is cheaper than discovering it in production.

Webhooks you can replay

Subscription credits are granted on invoice.paid, and top-ups on checkout.session.completed. Two rules keep this safe:

  1. Every Stripe event id is stored in the same transaction as the credit grant. A replayed event finds its id already there and does nothing.
  2. Credits come from the invoice's price id, never from the amount paid. A discounted first month still grants the full plan, and a coupon can never change what a plan is worth.

Safety rejections are failures too

Gateways reject prompts and outputs for safety reasons, and each one says so differently (a 403, a 422 with content_policy_violation, a flag on the output). The providers map all of them to one failure reason, content_blocked, which goes through the same refund claim with one difference: those refunds are capped per user per day, so the balance cannot be used to probe a filter for free.

What I would tell myself on day one

  • Put the money movement and the state change in the same transaction, always.
  • Make every "finish this job" path a conditional update that only one caller can win.
  • Build the fake provider first. Every failure path above was tested by typing a word into a prompt box.

The app this came from is Snap AI Studio. The two-photo AI kiss video tool is the one with the most moving parts: it runs an image model to put two people in one frame, then a video model on that frame, and both costs are debited up front.

Top comments (0)