DEV Community

Howard Shaw
Howard Shaw

Posted on

I Integrated 20+ AI Models. The Hard Part Wasn't the AI.

When I integrated my first AI model, the whole thing felt almost ridiculously simple.

There was an API, a prompt, a request, and eventually a result. I wrapped the API call in my application, handled the response, stored the output, and moved on to the next feature. It was one of those tasks where you could spend more time reading the documentation than actually writing the code.

That impression did not last very long.

Over the following months, I ended up integrating more than 20 image and video models into a single product. At first, I thought this would mostly be a matter of writing a few provider adapters and keeping a common interface. In reality, the adapters were probably the easy part. The difficult work was figuring out what "the same feature" even meant when every model had slightly different assumptions, capabilities, input requirements, behavior, pricing, and failure modes.

This was one of those projects where the deeper I got into it, the less I believed the simple version of the problem.

My original plan was much cleaner

I started with a fairly conventional idea.

The application would have a common generation interface, and each provider would sit behind an adapter. Something roughly like this:

type GenerateInput = {
  prompt?: string
  imageUrl?: string
  videoUrl?: string
  duration?: number
  aspectRatio?: string
}

type GenerateResult = {
  jobId: string
}

interface ModelAdapter {
  generate(input: GenerateInput): Promise<GenerateResult>
  getStatus(jobId: string): Promise<JobStatus>
}
Enter fullscreen mode Exit fullscreen mode

The idea was to keep the rest of the application ignorant of provider-specific details. If I wanted to add another model later, I would just implement another adapter.

This worked pretty well with the first couple of models. By the third one, I started adding exceptions. Then more exceptions appeared.

One provider wanted a publicly accessible image URL. Another worked better with uploaded assets. Another had a separate asset creation step before a generation request could even be made. Some models supported reference images, but the meaning of a reference image was not necessarily the same from one model to another. One provider offered a webhook while another expected the application to poll for status. Some exposed durations as arbitrary values, while others only allowed a short list of predefined options.

None of these differences were particularly difficult on their own. The problem was that they accumulated.

Eventually, the common interface started filling up with optional fields and special cases. At that point, I had to admit that I had built an abstraction that was trying too hard to make different things look identical.

A universal API sounds better than it actually is

This is probably the part of the project where I learned the most.

When you're designing an abstraction, there is a natural temptation to define the ideal interface first and then force every implementation to fit it. That can work very well when the implementations are variations of the same underlying system. AI models are often not.

Take something that sounds simple, like "image to video."

You might assume that the application sends an image and a prompt, and the model creates a video. Technically, that's often true. In practice, the differences are substantial. Some models treat the image primarily as the starting frame. Others use it more as visual guidance. Some preserve the subject and composition extremely well but are less reliable when the requested motion becomes complicated. Others are much better at camera movement. Some handle faces and people reasonably well, while another may behave very differently with the exact same source image.

From the outside, all of these can be presented as "image to video."

From an engineering and product perspective, they are not the same capability.

That distinction matters once users start selecting models based on what they are actually trying to create. If the application hides all of those differences behind one generic form, you can end up with a UI that looks beautifully consistent but gives users very little information about what will actually happen.

So I gradually stopped trying to make every model behave identically.

The goal became much more modest: provide a common foundation for the things that really are common, while allowing each provider to expose the capabilities that make it different.

That turned out to be a much healthier abstraction.

The model itself is only one part of the job

This was the other thing I seriously underestimated.

When I first started working with image and video generation, I naturally focused on the model call. That's the interesting part, after all. You send the input, something impressive comes back, and that's the part people usually see in demos.

In a real product, the model call is only a small section of the flow.

A user uploads an image. The application has to validate the file, store it somewhere, and make sure the provider can access it. The request has to be converted into the format that particular provider expects. A generation job needs to be created and associated with the user. The application needs to know whether the job is queued, processing, completed, or failed. If the provider reports an error, the application needs to decide whether that error is recoverable. When the generation finishes, the output needs to be stored and made available to the user.

And then there are all the things that don't happen on the happy path.

The user closes the browser while the generation is still running. A provider takes much longer than usual to respond. The request succeeds, but the output never becomes available. A retry creates a second generation when you didn't intend to create one. Two jobs finish at different times and arrive out of order.

None of this is particularly glamorous. It also has almost nothing to do with the intelligence of the model.

But these are the things that determine whether a user feels that the product works.

Eventually I started treating generation as a state machine

I spent quite a bit of time trying to make the generation lifecycle elegant. In the end, the simplest mental model turned out to be the most useful.

A generation is a job with a state.

queued
  ↓
processing
  ↓
completed
Enter fullscreen mode Exit fullscreen mode

Or:

queued
  ↓
processing
  ↓
failed
Enter fullscreen mode Exit fullscreen mode

The internal details vary by provider, but the application usually doesn't need to know every provider-specific status. The adapter can deal with polling, webhooks, provider response formats, retries, and all the other unpleasant details. The rest of the application just needs a reliable representation of the job.

For example:

type JobStatus =
  | "queued"
  | "processing"
  | "completed"
  | "failed"
Enter fullscreen mode Exit fullscreen mode

That sounds almost too basic, but it made a big difference.

Once I treated the generation itself as a first-class job rather than just an API request, several other design decisions became easier. It became much clearer where retries belonged, what the frontend should display, and how to handle a job that finishes after the user has left the page.

The nice thing about this approach is that it isn't really an "AI architecture." It's just good asynchronous application design. The AI part happens to be what creates the jobs.

Costs turned out to be part of the architecture

There is another complication you don't notice when testing models one at a time.

Different models don't just produce different results. They can have very different costs and generation times.

When you offer several models in one product, the user isn't really choosing only between Model A and Model B. They are also choosing between a faster and slower generation, a cheaper and more expensive generation, and sometimes a more predictable and less predictable one.

That has consequences for product design.

For example, the model with the best benchmark or the most impressive demo is not necessarily the model I want as the default for every task. If it costs several times as much and takes substantially longer, asking every user to start there may not actually produce a better experience.

This is one reason why I now think about model selection as a product problem, not just an engineering problem.

The cheapest model is not always right. The most expensive model is not always right either. What matters is whether the model is a good fit for the task the user is trying to accomplish.

That sounds obvious when written down. It was much less obvious when I was in the middle of integrating them.

Most of the bugs live in the edge cases

The happy path is still easy.

You upload something, send a request, wait, and get a result.

The real application starts when that sequence breaks.

I have had plenty of cases where a model worked perfectly with the test inputs I used during integration, and then behaved differently once I started trying a wider variety of real-world inputs. Resolution, aspect ratio, image composition, faces, source quality, motion requirements, and other details can all change the outcome.

The tricky part is that a failure isn't always an explicit failure.

Sometimes the API returns a normal success response and the result is simply not very good.

That makes AI integration different from many traditional API integrations. With a payment API, you usually know whether the operation succeeded. With a generation model, "success" can mean that a video was technically produced even though it doesn't satisfy what the user actually asked for.

So there are really two types of reliability to think about.

There is infrastructure reliability, where the job completes and the system correctly handles the result. Then there is model reliability, where the output is actually useful for the requested task.

You need both.

Adding the 20th model was easier than adding the second

This was one of the more surprising things I learned.

The first model is difficult because everything is new. The second is often harder than expected because it exposes the differences you didn't account for. By the time you've integrated several providers and rewritten the abstraction a couple of times, you start to understand where the real boundaries should be.

Once that foundation is reasonably stable, adding another model is often much more mechanical.

That became a useful test for me. If integrating one more model requires changes all over the application, something is probably wrong with the architecture. The same is true if the adapter has become so complicated that nobody really understands what it is doing.

A good abstraction does not eliminate differences between systems. It gives those differences a sensible place to live.

That may be one of the most useful lessons I took from this entire project.

I stopped chasing feature parity

At some point, I realized I was asking the wrong question.

I had been thinking, "How can I make every model support the same set of options?"

The better question was, "What are the capabilities that actually make sense to expose across these models?"

Those are not the same thing.

Some models have capabilities that others simply don't have. Trying to hide that fact usually makes the product more confusing, not less.

So instead of forcing everything into one giant shared set of controls, I started thinking in terms of common capabilities and model-specific capabilities. The common parts should feel consistent. The differences should be visible when they matter.

That also changed how I think about model selection.

I don't care as much about saying that a product supports 20 or 30 models. The number sounds good in a headline, but it doesn't tell the user very much.

What is much more useful is knowing why you might choose one model over another.

One may be better for preserving a character. Another may be better for motion. Another may be faster. Another may be more affordable. Another may support a type of input that the others don't.

The differences are the useful information.

Documentation is only the beginning

I've also become much less willing to consider a model "integrated" just because the API call works according to the documentation.

Documentation tells you how a provider expects the API to behave. It does not always tell you what it feels like to use that API as part of a product.

You discover the rest through testing.

A particular aspect ratio may produce unexpectedly poor results. A reference image may work very well in one scenario and poorly in another. A capability may exist on paper but have practical limitations that aren't obvious until you put it into a real workflow.

So my integration process has changed.

I still read the documentation first. But I don't consider the job finished when the request returns a successful response. I want to see the model behave with the kinds of inputs and workflows that the actual product supports.

That takes more time up front, but it saves a lot of confusion later.

This changed how I look at new models

When another impressive model gets announced now, my first reaction is no longer just, "How good is it?"

I'm more interested in what it does differently.

Does it solve a problem that the existing models don't solve well? Is it noticeably better for a particular type of input? Is it fast enough to be practical? Is the API reliable enough to build around? Is the output quality consistent enough that users won't feel like they are rolling the dice?

Those questions are less exciting than benchmark screenshots, but they are much more useful when you're actually running a product.

At some point, adding another model just because it exists stops being meaningful.

Twenty models that behave roughly the same way are not necessarily better than five models with clearly different strengths.

That was not how I thought about multi-model products when I started.

I'm still figuring it out

I don't want to turn this into one of those "here are the lessons I learned and now everything makes sense" posts.

It doesn't.

I'm still integrating models. I'm still finding strange edge cases. I'm still changing parts of the architecture that seemed perfectly reasonable a few months earlier. That's probably not going to stop anytime soon, because the models themselves keep changing.

But after working through enough integrations, I do have a much clearer idea of where the real engineering work is.

It isn't the API call.

The interesting part is everything around it: normalizing inputs without pretending providers are identical, managing asynchronous jobs, handling failures, dealing with cost and latency, understanding what each model is actually good at, and turning all of that into something a normal user can interact with without needing to know any of it.

That is the part I didn't really understand when I integrated my first model.

And that's probably the biggest difference between adding an AI API to a web application and actually building a product around AI.

A small side project that turned into a much bigger lesson

The project that pushed me into all of this is VioEvo AI, an independent AI image and video creation platform where I've been experimenting with multiple generation models in the same product.

I originally thought the interesting part would be the models themselves.

After 20+ integrations, I'm much more interested in the engineering problems that appear when you try to make all of those very different systems feel like one product.

Top comments (0)