<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Demon SDA</title>
    <description>The latest articles on DEV Community by Demon SDA (@demon_sda_f1c9cdf9faa5b76).</description>
    <link>https://dev.to/demon_sda_f1c9cdf9faa5b76</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F2861976%2F0a0859f5-f6f9-4dbe-8f52-6349a62729d6.png</url>
      <title>DEV Community: Demon SDA</title>
      <link>https://dev.to/demon_sda_f1c9cdf9faa5b76</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/demon_sda_f1c9cdf9faa5b76"/>
    <language>en</language>
    <item>
      <title>The AI API Was the Easy Part: Building a Production AI SaaS with Queues, Provider Abstraction, and Billing</title>
      <dc:creator>Demon SDA</dc:creator>
      <pubDate>Tue, 06 Oct 2026 13:18:15 +0000</pubDate>
      <link>https://dev.to/demon_sda_f1c9cdf9faa5b76/the-ai-api-was-the-easy-part-building-a-production-ai-saas-with-queues-provider-abstraction-and-op8</link>
      <guid>https://dev.to/demon_sda_f1c9cdf9faa5b76/the-ai-api-was-the-easy-part-building-a-production-ai-saas-with-queues-provider-abstraction-and-op8</guid>
      <description>&lt;p&gt;When I started building AI Music, the core workflow looked almost trivial:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User prompt
   ↓
Backend
   ↓
AI API
   ↓
Generated song
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For a prototype, that is enough.&lt;br&gt;
The interesting problems started later.&lt;br&gt;
Generation could take tens of seconds or several minutes. Then I added payments, personal voice workflows, video generation, multiple external AI services, retries, worker crashes, and the requirement that a user should never lose either their result or the credits they paid for an operation.&lt;br&gt;
At some point I realized that the AI API itself had become one of the simplest parts of the system.&lt;br&gt;
The difficult part was everything around it.&lt;br&gt;
Moving long-running AI jobs out of HTTP requests&lt;br&gt;
The first implementation you naturally want to write looks something like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/generate&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;reply&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;aiProvider&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt;
&lt;span class="p"&gt;})&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This works when the operation takes a few seconds.&lt;br&gt;
It becomes much less attractive when generation takes a minute.&lt;br&gt;
During that time:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the browser can disconnect;&lt;/li&gt;
&lt;li&gt;a reverse proxy can terminate the request;&lt;/li&gt;
&lt;li&gt;the API instance can restart;&lt;/li&gt;
&lt;li&gt;the external AI service can continue processing even after our connection disappears.
So I moved long-running operations into background jobs.
The flow became:
&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;POST /generation
       ↓
create database record
       ↓
validate billing
       ↓
enqueue BullMQ job
       ↓
return generationId
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The frontend now tracks an operation instead of waiting for one long HTTP request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;QUEUED
PROCESSING
COMPLETED
FAILED
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The actual AI work is performed by a worker.&lt;br&gt;
That small architectural change solved several problems at once.&lt;br&gt;
The HTTP request became short-lived.&lt;br&gt;
The generation state no longer depended on one Node.js process.&lt;br&gt;
And most importantly, unfinished jobs could be recovered after a crash or redeployment.&lt;br&gt;
The architecture evolved into Web / API / Worker&lt;br&gt;
The system eventually settled into a structure like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Web / Client
      ↓
Fastify API
      ↓
PostgreSQL
      ↓
Redis / BullMQ
      ↓
Background Worker
      ↓
AI Providers
      ↓
Object Storage
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The API is responsible for things such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;authentication;&lt;/li&gt;
&lt;li&gt;validation;&lt;/li&gt;
&lt;li&gt;domain state;&lt;/li&gt;
&lt;li&gt;authorization;&lt;/li&gt;
&lt;li&gt;credits and billing;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;enqueueing long-running jobs.&lt;br&gt;
The worker owns the expensive and slow work:&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;calling external AI services;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;polling job status;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;processing provider responses;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;downloading results;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;saving media;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;retries and recovery.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The browser never calls AI vendors directly.&lt;br&gt;
It also does not need to know where media is physically stored or which payment or AI service is currently being used.&lt;br&gt;
That separation became important later when the provider stack started changing.&lt;br&gt;
One AI provider is easy. The second one changes the design.&lt;br&gt;
With a single provider, it is very tempting to use its SDK directly throughout the codebase:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nx"&gt;vendor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="nx"&gt;vendor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getStatus&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="nx"&gt;vendor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;extend&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That works until another service is introduced.&lt;br&gt;
Maybe one service is better for music generation.&lt;br&gt;
Another provides personal voice functionality.&lt;br&gt;
Another handles video.&lt;br&gt;
If vendor-specific calls are spread across API routes, worker code, and business services, changing providers becomes expensive very quickly.&lt;br&gt;
So I introduced a provider abstraction layer.&lt;br&gt;
Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Product Logic
     ↓
Provider Layer
     ↓
Music / Voice / Video implementations
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The product asks for a capability:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nx"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generateMusic&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="nx"&gt;instead&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;referring&lt;/span&gt; &lt;span class="nx"&gt;to&lt;/span&gt; &lt;span class="nx"&gt;a&lt;/span&gt; &lt;span class="nx"&gt;specific&lt;/span&gt; &lt;span class="nx"&gt;vendor&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;span class="nx"&gt;someVendorSdk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;createTrack&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The rule became:&lt;/p&gt;

&lt;p&gt;Product logic should depend on capabilities, not vendor APIs.&lt;/p&gt;

&lt;p&gt;This sounds obvious, but it becomes extremely valuable once provider pricing, API contracts, reliability, or supported features start changing.&lt;br&gt;
A simple provider factory was still not enough&lt;br&gt;
Initially, a function like this seems sufficient:&lt;br&gt;
getCurrentMusicProvider()&lt;br&gt;
But there is a subtle problem.&lt;br&gt;
Imagine a user creates an asset using Provider A.&lt;br&gt;
A week later, the default provider changes to Provider B.&lt;br&gt;
The user returns and wants to continue working with the old asset.&lt;br&gt;
If the application simply uses the new default provider, Provider B may have no idea what the old external ID means.&lt;br&gt;
This led me to separate two routing problems.&lt;br&gt;
Creation routing&lt;br&gt;
For a new asset, the question is:&lt;br&gt;
Which provider should create it?&lt;/p&gt;

&lt;p&gt;The answer can depend on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;capabilities;&lt;/li&gt;
&lt;li&gt;feature flags;&lt;/li&gt;
&lt;li&gt;product rules;&lt;/li&gt;
&lt;li&gt;availability;&lt;/li&gt;
&lt;li&gt;the operation being requested.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Existing-asset routing&lt;br&gt;
For an existing asset, the question is different:&lt;br&gt;
Which provider owns this asset?&lt;/p&gt;

&lt;p&gt;The provider relationship is persisted with the domain entity.&lt;br&gt;
A simplified version looks like this:&lt;br&gt;
const provider =&lt;br&gt;
  track.provider ??&lt;br&gt;
  generation.provider ??&lt;br&gt;
  song.provider&lt;br&gt;
An existing asset therefore keeps affinity with the provider that created it.&lt;br&gt;
This is especially important for voice-related workflows.&lt;br&gt;
Moving an existing voice asset to a different vendor is not necessarily a harmless fallback. It can produce a technically valid but completely different result.&lt;br&gt;
Why I prefer fail-closed behavior&lt;br&gt;
Automatic fallback looks attractive:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;providerA&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;providerB&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For stateless requests, this can sometimes be reasonable.&lt;br&gt;
For stateful AI assets, it can be dangerous.&lt;/p&gt;

&lt;p&gt;A different provider may:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;not understand the existing ID;&lt;/li&gt;
&lt;li&gt;not support the requested capability;&lt;/li&gt;
&lt;li&gt;return an incompatible result;&lt;/li&gt;
&lt;li&gt;create a new asset that the application mistakenly associates with the old one.
For these operations I prefer fail-closed behavior.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;provider mismatch
      ↓
0 credits spent
      ↓
0 external HTTP calls
      ↓
explicit domain error
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It is less magical, but much easier to reason about.&lt;br&gt;
Billing changes the reliability requirements&lt;br&gt;
Many distributed-system problems become more serious once money is involved.&lt;/p&gt;

&lt;p&gt;Suppose this happens:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;the user starts a generation;&lt;/li&gt;
&lt;li&gt;credits are deducted;&lt;/li&gt;
&lt;li&gt;the provider accepts the task;&lt;/li&gt;
&lt;li&gt;the worker crashes;&lt;/li&gt;
&lt;li&gt;BullMQ retries the job.
Without idempotency, the retry can create another paid generation or deduct credits twice.
The order of operations therefore matters.
A simplified paid operation looks more like this:
&lt;/li&gt;
&lt;/ol&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;load asset
   ↓
verify provider affinity
   ↓
verify capability
   ↓
validate external IDs
   ↓
spend credits
   ↓
call provider
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The system should establish that an operation is valid before charging for it.&lt;br&gt;
If failure happens after credits have been spent, the failure path needs compensation:&lt;/p&gt;

&lt;p&gt;provider failure&lt;br&gt;
      ↓&lt;br&gt;
refund original spend&lt;/p&gt;

&lt;p&gt;This is also why credits are more useful as a ledger than as a mutable number.&lt;br&gt;
Instead of only knowing:&lt;br&gt;
balance = 452&lt;br&gt;
the system can explain why:&lt;br&gt;
+500 purchase&lt;br&gt;
-24  generation&lt;br&gt;
-120 voice operation&lt;br&gt;
+120 refund&lt;br&gt;
-24  generation&lt;br&gt;
That makes both debugging and support much easier.&lt;br&gt;
Idempotency is not optional for paid AI jobs&lt;/p&gt;

&lt;p&gt;Idempotency shows up everywhere in this kind of product:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;payments;&lt;/li&gt;
&lt;li&gt;queue retries;&lt;/li&gt;
&lt;li&gt;callbacks;&lt;/li&gt;
&lt;li&gt;background workers;&lt;/li&gt;
&lt;li&gt;external AI requests.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A payment webhook can arrive twice.&lt;br&gt;
A worker can execute the same job again after a transient error.&lt;br&gt;
Neither event should create a second business side effect.&lt;br&gt;
For long-running AI operations, I want enough durable state to determine whether the operation already exists:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;operationId&lt;/li&gt;
&lt;li&gt;idempotencyKey&lt;/li&gt;
&lt;li&gt;externalTaskId&lt;/li&gt;
&lt;li&gt;billingEntryId&lt;/li&gt;
&lt;li&gt;status&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Then a retry can resume an existing operation instead of blindly creating another one.&lt;br&gt;
External task IDs are part of domain state&lt;br&gt;
Consider this failure:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Worker
  ↓
POST /generate
  ↓
Provider creates task #abc123
  ↓
Worker crashes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the application restarts without knowing about abc123, it may send another POST.&lt;br&gt;
Now there are two provider tasks and potentially twice the cost.&lt;br&gt;
That means the external task ID is not just an implementation detail.&lt;br&gt;
It is recovery state.&lt;br&gt;
Once the external side effect has happened, the system needs enough durable information to continue working with it after its own process restarts.&lt;br&gt;
The general rule I ended up using is:&lt;br&gt;
If an external side effect has already happened, our system must be able to resume from it.&lt;/p&gt;

&lt;p&gt;A retry is not always the same operation&lt;br&gt;
It is easy to think of retry logic as:&lt;br&gt;
await job.retry()&lt;br&gt;
But safe retry behavior depends on when the failure happened.&lt;br&gt;
For example:&lt;/p&gt;

&lt;p&gt;before external request&lt;br&gt;
→ safe to retry&lt;/p&gt;

&lt;p&gt;external task already created&lt;br&gt;
→ resume polling&lt;/p&gt;

&lt;p&gt;external request accepted but task ID not persisted&lt;br&gt;
→ reconciliation required&lt;/p&gt;

&lt;p&gt;billing succeeded but provider permanently failed&lt;br&gt;
→ refund&lt;/p&gt;

&lt;p&gt;Those are very different failure modes.&lt;br&gt;
This is why retry policy gradually became part of the business workflow rather than just a queue configuration option.&lt;br&gt;
Reconciliation became necessary&lt;br&gt;
Queues are useful, but they do not eliminate every inconsistent state.&lt;br&gt;
A worker can crash between two durable writes.&lt;br&gt;
A redeployment can happen while an external task is still processing.&lt;br&gt;
So I added reconciliation logic for jobs that stay in an intermediate state too long.&lt;br&gt;
Conceptually:&lt;/p&gt;

&lt;p&gt;find PROCESSING jobs older than threshold&lt;br&gt;
        ↓&lt;br&gt;
inspect externalTaskId&lt;br&gt;
        ↓&lt;br&gt;
query provider&lt;br&gt;
        ↓&lt;br&gt;
complete / fail / refund / retry&lt;/p&gt;

&lt;p&gt;For free operations this is mostly a reliability concern.&lt;br&gt;
For paid operations it becomes a financial correctness concern.&lt;br&gt;
Provider errors should not leak into the UI&lt;br&gt;
Another useful boundary is error normalization.&lt;br&gt;
The frontend should not receive provider-specific errors such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"provider_status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;402&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"vendor_code"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"SOME_VENDOR_ERROR_123"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The provider layer maps external responses into domain errors:&lt;br&gt;
PROVIDER_UNAVAILABLE&lt;br&gt;
PROVIDER_AFFINITY_MISMATCH&lt;br&gt;
OPERATION_UNAVAILABLE&lt;br&gt;
The API then maps those into the application's public HTTP contract.&lt;br&gt;
That means I can replace an external integration without forcing the frontend to understand a completely new error model.&lt;br&gt;
The browser should know domain concepts, not vendors&lt;br&gt;
This became a broader rule in the project.&lt;br&gt;
The frontend works with entities such as:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;song&lt;/li&gt;
&lt;li&gt;generation&lt;/li&gt;
&lt;li&gt;voiceProfile&lt;/li&gt;
&lt;li&gt;shareVideo&lt;/li&gt;
&lt;li&gt;payment&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;It does not work with vendor-specific objects.&lt;br&gt;
The same principle applies to object storage.&lt;br&gt;
The browser should not need direct knowledge of storage implementation details.&lt;br&gt;
Media is served through application-controlled access, typically using short-lived signed URLs.&lt;br&gt;
This keeps infrastructure decisions behind the backend boundary.&lt;br&gt;
I also split music, voice, and video providers&lt;br&gt;
At first, one large interface can look convenient:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kr"&gt;interface&lt;/span&gt; &lt;span class="nx"&gt;Provider&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;generateMusic&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
  &lt;span class="nf"&gt;cloneVoice&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
  &lt;span class="nf"&gt;generateVideo&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
  &lt;span class="nf"&gt;animateImage&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But that interface quickly becomes a collection of unrelated capabilities.&lt;br&gt;
I eventually split them conceptually:&lt;/p&gt;

&lt;p&gt;ai-providers/&lt;br&gt;
  music/&lt;br&gt;
  voice/&lt;br&gt;
  video/&lt;/p&gt;

&lt;p&gt;Each capability group has its own contract.&lt;br&gt;
This matters because their operational characteristics are different.&lt;br&gt;
Music generation, voice processing, and video generation have different:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;request models;&lt;/li&gt;
&lt;li&gt;polling behavior;&lt;/li&gt;
&lt;li&gt;timeouts;&lt;/li&gt;
&lt;li&gt;concurrency constraints;&lt;/li&gt;
&lt;li&gt;costs;&lt;/li&gt;
&lt;li&gt;post-processing requirements.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Changing the video stack should not force a redesign of music generation.&lt;br&gt;
That turned out to be a useful test for whether the abstraction was in the right place.&lt;br&gt;
The worker became a real runtime component&lt;br&gt;
The worker started as a place to execute background jobs.&lt;br&gt;
Over time it became a distinct runtime with responsibilities of its own.&lt;br&gt;
The API handles:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;auth&lt;/li&gt;
&lt;li&gt;validation&lt;/li&gt;
&lt;li&gt;domain state&lt;/li&gt;
&lt;li&gt;billing&lt;/li&gt;
&lt;li&gt;enqueue&lt;/li&gt;
&lt;li&gt;The worker handles:&lt;/li&gt;
&lt;li&gt;provider calls&lt;/li&gt;
&lt;li&gt;polling&lt;/li&gt;
&lt;li&gt;downloads&lt;/li&gt;
&lt;li&gt;storage&lt;/li&gt;
&lt;li&gt;retries&lt;/li&gt;
&lt;li&gt;reconciliation&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This also affects scaling.&lt;br&gt;
The API can scale according to HTTP traffic.&lt;br&gt;
Workers need to scale according to job volume and provider concurrency limits.&lt;br&gt;
Those are not the same thing.&lt;br&gt;
Adding ten worker replicas does not help if an external provider only allows five concurrent jobs.&lt;br&gt;
Object storage is part of the workflow&lt;br&gt;
AI services often return temporary media URLs.&lt;br&gt;
Using those URLs as permanent application assets creates unnecessary dependency on the provider.&lt;br&gt;
The URL may expire.&lt;br&gt;
Its format may change.&lt;br&gt;
It may expose implementation details.&lt;br&gt;
So successful results are copied into application-controlled object storage.&lt;br&gt;
The application then owns the media lifecycle.&lt;/p&gt;

&lt;p&gt;This gives more control over:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;retention;&lt;/li&gt;
&lt;li&gt;signed access;&lt;/li&gt;
&lt;li&gt;previews;&lt;/li&gt;
&lt;li&gt;thumbnails;&lt;/li&gt;
&lt;li&gt;post-processing;&lt;/li&gt;
&lt;li&gt;migration between providers.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Infrastructure costs appear earlier than users&lt;br&gt;
One lesson surprised me because it was not AI-specific at all.&lt;br&gt;
A SaaS can have very few users and still run infrastructure 24/7.&lt;br&gt;
Database compute, Redis, API services, workers, monitoring, polling jobs — each component looks inexpensive in isolation.&lt;br&gt;
Together they become noticeable.&lt;br&gt;
I eventually had to investigate why database compute was staying active even when the product was idle.&lt;br&gt;
Background monitoring and periodic jobs can easily prevent autosuspend.&lt;br&gt;
The lesson was:&lt;br&gt;
Before product-market fit, infrastructure optimization is less about saving every cent and more about making sure idle infrastructure is actually idle.&lt;/p&gt;

&lt;p&gt;What I would design earlier next time&lt;br&gt;
If I started the product again, I would introduce several things earlier:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;provider abstraction;&lt;/li&gt;
&lt;li&gt;idempotency for every paid operation;&lt;/li&gt;
&lt;li&gt;persisted external task IDs;&lt;/li&gt;
&lt;li&gt;failure-path design alongside the happy path;&lt;/li&gt;
&lt;li&gt;separate routing for new assets and existing assets.
Not because an MVP needs enterprise architecture from day one.
But because these are the areas that become painful first when an AI demo turns into a paid product.
Final architecture
A very simplified view now looks like this:&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Web / Client&lt;br&gt;
      ↓&lt;br&gt;
API&lt;br&gt;
      ↓&lt;br&gt;
Domain validation&lt;br&gt;
      ↓&lt;br&gt;
Billing&lt;br&gt;
      ↓&lt;br&gt;
Queue&lt;br&gt;
      ↓&lt;br&gt;
Worker&lt;br&gt;
      ↓&lt;br&gt;
Provider abstraction&lt;br&gt;
      ↓&lt;br&gt;
External AI services&lt;br&gt;
      ↓&lt;br&gt;
Object Storage&lt;/p&gt;

&lt;p&gt;Around that flow are the mechanisms that make it production-safe:&lt;br&gt;
idempotency&lt;br&gt;
asset affinity&lt;br&gt;
retries&lt;br&gt;
reconciliation&lt;br&gt;
refunds&lt;br&gt;
observability&lt;br&gt;
signed media access&lt;br&gt;
The AI model is important.&lt;br&gt;
But most of the production engineering ended up happening around it.&lt;br&gt;
The main lesson&lt;br&gt;
AI makes it incredibly cheap to build a convincing prototype:&lt;br&gt;
prompt → API → result&lt;br&gt;
Production brings you back to classic software engineering:&lt;br&gt;
queues, durable state, retries, idempotency, recovery, billing, observability, and security.&lt;br&gt;
For me, that became the most interesting part of building the product.&lt;br&gt;
Not calling the model.&lt;br&gt;
Building a system that still behaves correctly when a provider is slow, a worker crashes, a callback arrives twice, or a user has already paid for an operation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Links&lt;/strong&gt;&lt;br&gt;
Live product:&lt;br&gt;
[&lt;a href="https://voice-to-song.com/en" rel="noopener noreferrer"&gt;https://voice-to-song.com/en&lt;/a&gt;]&lt;/p&gt;

&lt;p&gt;Public engineering showcase:&lt;br&gt;
[&lt;a href="https://github.com/Demon5611/AI-Music-Showcase" rel="noopener noreferrer"&gt;https://github.com/Demon5611/AI-Music-Showcase&lt;/a&gt;]&lt;/p&gt;

&lt;p&gt;linkedin:&lt;br&gt;
[&lt;a href="https://www.linkedin.com/in/dmitriy-sedov" rel="noopener noreferrer"&gt;https://www.linkedin.com/in/dmitriy-sedov&lt;/a&gt;]&lt;/p&gt;

&lt;p&gt;The showcase intentionally excludes production credentials, merchant configuration, internal infrastructure settings, and some commercial logic. It exists as a public engineering view of the architecture and design decisions.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>node</category>
      <category>sass</category>
    </item>
  </channel>
</rss>
