DEV Community

Cover image for The AI API Was the Easy Part: Building a Production AI SaaS with Queues, Provider Abstraction, and Billing
Demon SDA
Demon SDA

Posted on

The AI API Was the Easy Part: Building a Production AI SaaS with Queues, Provider Abstraction, and Billing

When I started building AI Music, the core workflow looked almost trivial:

User prompt
   ↓
Backend
   ↓
AI API
   ↓
Generated song
Enter fullscreen mode Exit fullscreen mode

For a prototype, that is enough.
The interesting problems started later.
Generation could take tens of seconds or several minutes. Then I added payments, personal voice workflows, video generation, multiple external AI services, retries, worker crashes, and the requirement that a user should never lose either their result or the credits they paid for an operation.
At some point I realized that the AI API itself had become one of the simplest parts of the system.
The difficult part was everything around it.
Moving long-running AI jobs out of HTTP requests
The first implementation you naturally want to write looks something like this:

app.post("/generate", async (req, reply) => {
  const result = await aiProvider.generate(req.body)

  return result
})
Enter fullscreen mode Exit fullscreen mode

This works when the operation takes a few seconds.
It becomes much less attractive when generation takes a minute.
During that time:

  • the browser can disconnect;
  • a reverse proxy can terminate the request;
  • the API instance can restart;
  • the external AI service can continue processing even after our connection disappears. So I moved long-running operations into background jobs. The flow became:
POST /generation
       ↓
create database record
       ↓
validate billing
       ↓
enqueue BullMQ job
       ↓
return generationId
Enter fullscreen mode Exit fullscreen mode

The frontend now tracks an operation instead of waiting for one long HTTP request:

QUEUED
PROCESSING
COMPLETED
FAILED
Enter fullscreen mode Exit fullscreen mode

The actual AI work is performed by a worker.
That small architectural change solved several problems at once.
The HTTP request became short-lived.
The generation state no longer depended on one Node.js process.
And most importantly, unfinished jobs could be recovered after a crash or redeployment.
The architecture evolved into Web / API / Worker
The system eventually settled into a structure like this:

Web / Client
      ↓
Fastify API
      ↓
PostgreSQL
      ↓
Redis / BullMQ
      ↓
Background Worker
      ↓
AI Providers
      ↓
Object Storage
Enter fullscreen mode Exit fullscreen mode

The API is responsible for things such as:

  • authentication;
  • validation;
  • domain state;
  • authorization;
  • credits and billing;
  • enqueueing long-running jobs.
    The worker owns the expensive and slow work:

  • calling external AI services;

  • polling job status;

  • processing provider responses;

  • downloading results;

  • saving media;

  • retries and recovery.

The browser never calls AI vendors directly.
It also does not need to know where media is physically stored or which payment or AI service is currently being used.
That separation became important later when the provider stack started changing.
One AI provider is easy. The second one changes the design.
With a single provider, it is very tempting to use its SDK directly throughout the codebase:

vendor.generate()
vendor.getStatus()
vendor.extend()
Enter fullscreen mode Exit fullscreen mode

That works until another service is introduced.
Maybe one service is better for music generation.
Another provides personal voice functionality.
Another handles video.
If vendor-specific calls are spread across API routes, worker code, and business services, changing providers becomes expensive very quickly.
So I introduced a provider abstraction layer.
Conceptually:

Product Logic
     ↓
Provider Layer
     ↓
Music / Voice / Video implementations
Enter fullscreen mode Exit fullscreen mode

The product asks for a capability:

provider.generateMusic()
instead of referring to a specific vendor:
someVendorSdk.createTrack()
Enter fullscreen mode Exit fullscreen mode

The rule became:

Product logic should depend on capabilities, not vendor APIs.

This sounds obvious, but it becomes extremely valuable once provider pricing, API contracts, reliability, or supported features start changing.
A simple provider factory was still not enough
Initially, a function like this seems sufficient:
getCurrentMusicProvider()
But there is a subtle problem.
Imagine a user creates an asset using Provider A.
A week later, the default provider changes to Provider B.
The user returns and wants to continue working with the old asset.
If the application simply uses the new default provider, Provider B may have no idea what the old external ID means.
This led me to separate two routing problems.
Creation routing
For a new asset, the question is:
Which provider should create it?

The answer can depend on:

  • capabilities;
  • feature flags;
  • product rules;
  • availability;
  • the operation being requested.

Existing-asset routing
For an existing asset, the question is different:
Which provider owns this asset?

The provider relationship is persisted with the domain entity.
A simplified version looks like this:
const provider =
track.provider ??
generation.provider ??
song.provider
An existing asset therefore keeps affinity with the provider that created it.
This is especially important for voice-related workflows.
Moving an existing voice asset to a different vendor is not necessarily a harmless fallback. It can produce a technically valid but completely different result.
Why I prefer fail-closed behavior
Automatic fallback looks attractive:

try {
  return await providerA.generate()
} catch {
  return await providerB.generate()
}
Enter fullscreen mode Exit fullscreen mode

For stateless requests, this can sometimes be reasonable.
For stateful AI assets, it can be dangerous.

A different provider may:

  • not understand the existing ID;
  • not support the requested capability;
  • return an incompatible result;
  • create a new asset that the application mistakenly associates with the old one. For these operations I prefer fail-closed behavior.

For example:

provider mismatch
      ↓
0 credits spent
      ↓
0 external HTTP calls
      ↓
explicit domain error
Enter fullscreen mode Exit fullscreen mode

It is less magical, but much easier to reason about.
Billing changes the reliability requirements
Many distributed-system problems become more serious once money is involved.

Suppose this happens:

  1. the user starts a generation;
  2. credits are deducted;
  3. the provider accepts the task;
  4. the worker crashes;
  5. BullMQ retries the job. Without idempotency, the retry can create another paid generation or deduct credits twice. The order of operations therefore matters. A simplified paid operation looks more like this:
load asset
   ↓
verify provider affinity
   ↓
verify capability
   ↓
validate external IDs
   ↓
spend credits
   ↓
call provider
Enter fullscreen mode Exit fullscreen mode

The system should establish that an operation is valid before charging for it.
If failure happens after credits have been spent, the failure path needs compensation:

provider failure
↓
refund original spend

This is also why credits are more useful as a ledger than as a mutable number.
Instead of only knowing:
balance = 452
the system can explain why:
+500 purchase
-24 generation
-120 voice operation
+120 refund
-24 generation
That makes both debugging and support much easier.
Idempotency is not optional for paid AI jobs

Idempotency shows up everywhere in this kind of product:

  • payments;
  • queue retries;
  • callbacks;
  • background workers;
  • external AI requests.

A payment webhook can arrive twice.
A worker can execute the same job again after a transient error.
Neither event should create a second business side effect.
For long-running AI operations, I want enough durable state to determine whether the operation already exists:

  1. operationId
  2. idempotencyKey
  3. externalTaskId
  4. billingEntryId
  5. status

Then a retry can resume an existing operation instead of blindly creating another one.
External task IDs are part of domain state
Consider this failure:

Worker
  ↓
POST /generate
  ↓
Provider creates task #abc123
  ↓
Worker crashes
Enter fullscreen mode Exit fullscreen mode

If the application restarts without knowing about abc123, it may send another POST.
Now there are two provider tasks and potentially twice the cost.
That means the external task ID is not just an implementation detail.
It is recovery state.
Once the external side effect has happened, the system needs enough durable information to continue working with it after its own process restarts.
The general rule I ended up using is:
If an external side effect has already happened, our system must be able to resume from it.

A retry is not always the same operation
It is easy to think of retry logic as:
await job.retry()
But safe retry behavior depends on when the failure happened.
For example:

before external request
→ safe to retry

external task already created
→ resume polling

external request accepted but task ID not persisted
→ reconciliation required

billing succeeded but provider permanently failed
→ refund

Those are very different failure modes.
This is why retry policy gradually became part of the business workflow rather than just a queue configuration option.
Reconciliation became necessary
Queues are useful, but they do not eliminate every inconsistent state.
A worker can crash between two durable writes.
A redeployment can happen while an external task is still processing.
So I added reconciliation logic for jobs that stay in an intermediate state too long.
Conceptually:

find PROCESSING jobs older than threshold
↓
inspect externalTaskId
↓
query provider
↓
complete / fail / refund / retry

For free operations this is mostly a reliability concern.
For paid operations it becomes a financial correctness concern.
Provider errors should not leak into the UI
Another useful boundary is error normalization.
The frontend should not receive provider-specific errors such as:

{
  "provider_status": 402,
  "vendor_code": "SOME_VENDOR_ERROR_123"
}
Enter fullscreen mode Exit fullscreen mode

The provider layer maps external responses into domain errors:
PROVIDER_UNAVAILABLE
PROVIDER_AFFINITY_MISMATCH
OPERATION_UNAVAILABLE
The API then maps those into the application's public HTTP contract.
That means I can replace an external integration without forcing the frontend to understand a completely new error model.
The browser should know domain concepts, not vendors
This became a broader rule in the project.
The frontend works with entities such as:

  1. song
  2. generation
  3. voiceProfile
  4. shareVideo
  5. payment

It does not work with vendor-specific objects.
The same principle applies to object storage.
The browser should not need direct knowledge of storage implementation details.
Media is served through application-controlled access, typically using short-lived signed URLs.
This keeps infrastructure decisions behind the backend boundary.
I also split music, voice, and video providers
At first, one large interface can look convenient:

interface Provider {
  generateMusic()
  cloneVoice()
  generateVideo()
  animateImage()
}
Enter fullscreen mode Exit fullscreen mode

But that interface quickly becomes a collection of unrelated capabilities.
I eventually split them conceptually:

ai-providers/
music/
voice/
video/

Each capability group has its own contract.
This matters because their operational characteristics are different.
Music generation, voice processing, and video generation have different:

  • request models;
  • polling behavior;
  • timeouts;
  • concurrency constraints;
  • costs;
  • post-processing requirements.

Changing the video stack should not force a redesign of music generation.
That turned out to be a useful test for whether the abstraction was in the right place.
The worker became a real runtime component
The worker started as a place to execute background jobs.
Over time it became a distinct runtime with responsibilities of its own.
The API handles:

  1. auth
  2. validation
  3. domain state
  4. billing
  5. enqueue
  6. The worker handles:
  7. provider calls
  8. polling
  9. downloads
  10. storage
  11. retries
  12. reconciliation

This also affects scaling.
The API can scale according to HTTP traffic.
Workers need to scale according to job volume and provider concurrency limits.
Those are not the same thing.
Adding ten worker replicas does not help if an external provider only allows five concurrent jobs.
Object storage is part of the workflow
AI services often return temporary media URLs.
Using those URLs as permanent application assets creates unnecessary dependency on the provider.
The URL may expire.
Its format may change.
It may expose implementation details.
So successful results are copied into application-controlled object storage.
The application then owns the media lifecycle.

This gives more control over:

  • retention;
  • signed access;
  • previews;
  • thumbnails;
  • post-processing;
  • migration between providers.

Infrastructure costs appear earlier than users
One lesson surprised me because it was not AI-specific at all.
A SaaS can have very few users and still run infrastructure 24/7.
Database compute, Redis, API services, workers, monitoring, polling jobs — each component looks inexpensive in isolation.
Together they become noticeable.
I eventually had to investigate why database compute was staying active even when the product was idle.
Background monitoring and periodic jobs can easily prevent autosuspend.
The lesson was:
Before product-market fit, infrastructure optimization is less about saving every cent and more about making sure idle infrastructure is actually idle.

What I would design earlier next time
If I started the product again, I would introduce several things earlier:

  • provider abstraction;
  • idempotency for every paid operation;
  • persisted external task IDs;
  • failure-path design alongside the happy path;
  • separate routing for new assets and existing assets. Not because an MVP needs enterprise architecture from day one. But because these are the areas that become painful first when an AI demo turns into a paid product. Final architecture A very simplified view now looks like this:

Web / Client
↓
API
↓
Domain validation
↓
Billing
↓
Queue
↓
Worker
↓
Provider abstraction
↓
External AI services
↓
Object Storage

Around that flow are the mechanisms that make it production-safe:
idempotency
asset affinity
retries
reconciliation
refunds
observability
signed media access
The AI model is important.
But most of the production engineering ended up happening around it.
The main lesson
AI makes it incredibly cheap to build a convincing prototype:
prompt → API → result
Production brings you back to classic software engineering:
queues, durable state, retries, idempotency, recovery, billing, observability, and security.
For me, that became the most interesting part of building the product.
Not calling the model.
Building a system that still behaves correctly when a provider is slow, a worker crashes, a callback arrives twice, or a user has already paid for an operation.

Links
Live product:
[https://voice-to-song.com/en]

Public engineering showcase:
[https://github.com/Demon5611/AI-Music-Showcase]

linkedin:
[https://www.linkedin.com/in/dmitriy-sedov]

The showcase intentionally excludes production credentials, merchant configuration, internal infrastructure settings, and some commercial logic. It exists as a public engineering view of the architecture and design decisions.

Top comments (0)