DEV Community

mandapi
mandapi

Posted on

How I Cut AI API Costs Without Rewriting My OpenAI Integration

If you're building with LLMs, the model itself is often the easiest part.

The annoying part comes later.

You start with one provider. Then you want to test another model. Soon you have multiple API keys, different pricing structures, different endpoints, separate billing accounts, and provider-specific code scattered across your project.

I ran into exactly this problem while building AI applications.

The problem with using multiple AI providers

Suppose an application needs access to several model families:

  • GPT for general reasoning and coding
  • Claude for long-form tasks
  • Gemini for another price/performance option
  • DeepSeek for inexpensive workloads

Using each provider directly can mean maintaining several integrations.

Even when APIs look similar, authentication, model names, endpoints, billing and availability can differ.

There is also a second problem: cost.

For experiments, agents and applications processing large numbers of tokens, API costs can become significant surprisingly quickly.

So I wanted two things:

  1. One interface for multiple model providers.
  2. The ability to choose cheaper models or routes without rewriting the application.

OpenAI compatibility makes this much easier

A useful approach is to standardize around the OpenAI API format.

Instead of changing application logic whenever you change providers, you keep essentially the same request structure and change the base URL and model.

For example, an application using the OpenAI Python SDK can look roughly like this:

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="YOUR_OPENAI_COMPATIBLE_ENDPOINT"
)

response = client.chat.completions.create(
    model="YOUR_MODEL",
    messages=[
        {
            "role": "user",
            "content": "Explain why API compatibility matters."
        }
    ]
)

print(response.choices[0].message.content)
Enter fullscreen mode Exit fullscreen mode

The important part isn't the few lines of code.

It's that the application no longer needs to be tightly coupled to a single model provider.

I ended up building this into MandAPI

While working on this problem, I built MandAPI, an OpenAI-compatible multi-model API gateway.

The idea is deliberately simple:

one API format, one API key, multiple AI model families.

Instead of maintaining completely separate integrations, developers can access models from families such as GPT, Claude, Gemini and DeepSeek through an OpenAI-compatible interface.

That also makes price experimentation much easier.

If a workload doesn't require the most expensive model, you can move it to a cheaper model without redesigning the application.

Why this matters for AI API costs

A common mistake is using the most capable model for every request.

In a real application, workloads are usually mixed.

Some requests need strong reasoning.

Others are simple extraction, classification, rewriting, summarization or conversational tasks.

A better architecture can look like:

Complex reasoning
        ↓
High-capability model

Normal generation
        ↓
Mid-cost model

Simple/high-volume tasks
        ↓
Low-cost model
Enter fullscreen mode Exit fullscreen mode

Once multiple models share a compatible interface, routing workloads this way becomes much easier.

For high-token applications, the difference can become substantial.

Don't compare models only by headline price

Token price is important, but it isn't the only variable.

When comparing AI APIs, I now look at:

  • input token price
  • output token price
  • context limits
  • model quality
  • latency
  • streaming support
  • compatibility
  • availability
  • payment friction
  • actual cost for my workload

The cheapest model on paper isn't necessarily the cheapest model for the application.

A model that needs twice as many attempts to produce an acceptable answer may actually cost more.

Multi-model APIs are also useful as an abstraction layer

There is another benefit that is easy to underestimate: avoiding provider lock-in.

Your application talks to an interface rather than being designed around one specific provider.

That makes it easier to:

  • benchmark new models
  • change models when pricing changes
  • add fallback routes
  • test cheaper alternatives
  • migrate workloads
  • build model-routing systems

This becomes increasingly useful because AI model pricing and capabilities change very quickly.

What I'm building next

I'm continuing to work on MandAPI as both an API gateway and a model/pricing research project.

One area I'm particularly interested in is transparent LLM price comparison.

I've started publishing some of the underlying developer resources and pricing research openly on GitHub:

MandAPI Developer Kit

The goal is to make it easier for developers to compare models based on actual API economics rather than marketing pages.

I'm also especially interested in making AI APIs easier to purchase and use in markets where international billing can be inconvenient, including Brazil.

Final thought

You don't necessarily need to rewrite your AI application to experiment with different providers.

An OpenAI-compatible abstraction layer can make model switching surprisingly simple.

And once switching becomes simple, price becomes something you can optimize continuously instead of something you're locked into.

If you're building an AI product with significant API usage, it's worth designing for model portability from the beginning.

I'd be interested to hear how other developers are handling multi-model routing and API costs.

Top comments (0)