<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: mandapi</title>
    <description>The latest articles on DEV Community by mandapi (@puchi_fan_d7a71680ed6a559).</description>
    <link>https://dev.to/puchi_fan_d7a71680ed6a559</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4158923%2Fb3bf01d8-d149-43b0-84db-7931047662d5.png</url>
      <title>DEV Community: mandapi</title>
      <link>https://dev.to/puchi_fan_d7a71680ed6a559</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/puchi_fan_d7a71680ed6a559"/>
    <language>en</language>
    <item>
      <title>How I Cut AI API Costs Without Rewriting My OpenAI Integration</title>
      <dc:creator>mandapi</dc:creator>
      <pubDate>Mon, 05 Oct 2026 16:16:35 +0000</pubDate>
      <link>https://dev.to/puchi_fan_d7a71680ed6a559/how-i-cut-ai-api-costs-without-rewriting-my-openai-integration-5b1a</link>
      <guid>https://dev.to/puchi_fan_d7a71680ed6a559/how-i-cut-ai-api-costs-without-rewriting-my-openai-integration-5b1a</guid>
      <description>&lt;p&gt;If you're building with LLMs, the model itself is often the easiest part.&lt;/p&gt;

&lt;p&gt;The annoying part comes later.&lt;/p&gt;

&lt;p&gt;You start with one provider. Then you want to test another model. Soon you have multiple API keys, different pricing structures, different endpoints, separate billing accounts, and provider-specific code scattered across your project.&lt;/p&gt;

&lt;p&gt;I ran into exactly this problem while building AI applications.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem with using multiple AI providers
&lt;/h2&gt;

&lt;p&gt;Suppose an application needs access to several model families:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;GPT for general reasoning and coding&lt;/li&gt;
&lt;li&gt;Claude for long-form tasks&lt;/li&gt;
&lt;li&gt;Gemini for another price/performance option&lt;/li&gt;
&lt;li&gt;DeepSeek for inexpensive workloads&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Using each provider directly can mean maintaining several integrations.&lt;/p&gt;

&lt;p&gt;Even when APIs look similar, authentication, model names, endpoints, billing and availability can differ.&lt;/p&gt;

&lt;p&gt;There is also a second problem: &lt;strong&gt;cost&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;For experiments, agents and applications processing large numbers of tokens, API costs can become significant surprisingly quickly.&lt;/p&gt;

&lt;p&gt;So I wanted two things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;One interface for multiple model providers.&lt;/li&gt;
&lt;li&gt;The ability to choose cheaper models or routes without rewriting the application.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  OpenAI compatibility makes this much easier
&lt;/h2&gt;

&lt;p&gt;A useful approach is to standardize around the OpenAI API format.&lt;/p&gt;

&lt;p&gt;Instead of changing application logic whenever you change providers, you keep essentially the same request structure and change the base URL and model.&lt;/p&gt;

&lt;p&gt;For example, an application using the OpenAI Python SDK can look roughly like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;YOUR_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;YOUR_OPENAI_COMPATIBLE_ENDPOINT&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;YOUR_MODEL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Explain why API compatibility matters.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important part isn't the few lines of code.&lt;/p&gt;

&lt;p&gt;It's that the application no longer needs to be tightly coupled to a single model provider.&lt;/p&gt;

&lt;h2&gt;
  
  
  I ended up building this into MandAPI
&lt;/h2&gt;

&lt;p&gt;While working on this problem, I built &lt;a href="https://mandapi.com" rel="noopener noreferrer"&gt;MandAPI&lt;/a&gt;, an OpenAI-compatible multi-model API gateway.&lt;/p&gt;

&lt;p&gt;The idea is deliberately simple:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;one API format, one API key, multiple AI model families.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Instead of maintaining completely separate integrations, developers can access models from families such as GPT, Claude, Gemini and DeepSeek through an OpenAI-compatible interface.&lt;/p&gt;

&lt;p&gt;That also makes price experimentation much easier.&lt;/p&gt;

&lt;p&gt;If a workload doesn't require the most expensive model, you can move it to a cheaper model without redesigning the application.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters for AI API costs
&lt;/h2&gt;

&lt;p&gt;A common mistake is using the most capable model for every request.&lt;/p&gt;

&lt;p&gt;In a real application, workloads are usually mixed.&lt;/p&gt;

&lt;p&gt;Some requests need strong reasoning.&lt;/p&gt;

&lt;p&gt;Others are simple extraction, classification, rewriting, summarization or conversational tasks.&lt;/p&gt;

&lt;p&gt;A better architecture can look like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Complex reasoning
        ↓
High-capability model

Normal generation
        ↓
Mid-cost model

Simple/high-volume tasks
        ↓
Low-cost model
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once multiple models share a compatible interface, routing workloads this way becomes much easier.&lt;/p&gt;

&lt;p&gt;For high-token applications, the difference can become substantial.&lt;/p&gt;

&lt;h2&gt;
  
  
  Don't compare models only by headline price
&lt;/h2&gt;

&lt;p&gt;Token price is important, but it isn't the only variable.&lt;/p&gt;

&lt;p&gt;When comparing AI APIs, I now look at:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;input token price&lt;/li&gt;
&lt;li&gt;output token price&lt;/li&gt;
&lt;li&gt;context limits&lt;/li&gt;
&lt;li&gt;model quality&lt;/li&gt;
&lt;li&gt;latency&lt;/li&gt;
&lt;li&gt;streaming support&lt;/li&gt;
&lt;li&gt;compatibility&lt;/li&gt;
&lt;li&gt;availability&lt;/li&gt;
&lt;li&gt;payment friction&lt;/li&gt;
&lt;li&gt;actual cost for my workload&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The cheapest model on paper isn't necessarily the cheapest model for the application.&lt;/p&gt;

&lt;p&gt;A model that needs twice as many attempts to produce an acceptable answer may actually cost more.&lt;/p&gt;

&lt;h2&gt;
  
  
  Multi-model APIs are also useful as an abstraction layer
&lt;/h2&gt;

&lt;p&gt;There is another benefit that is easy to underestimate: avoiding provider lock-in.&lt;/p&gt;

&lt;p&gt;Your application talks to an interface rather than being designed around one specific provider.&lt;/p&gt;

&lt;p&gt;That makes it easier to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;benchmark new models&lt;/li&gt;
&lt;li&gt;change models when pricing changes&lt;/li&gt;
&lt;li&gt;add fallback routes&lt;/li&gt;
&lt;li&gt;test cheaper alternatives&lt;/li&gt;
&lt;li&gt;migrate workloads&lt;/li&gt;
&lt;li&gt;build model-routing systems&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This becomes increasingly useful because AI model pricing and capabilities change very quickly.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'm building next
&lt;/h2&gt;

&lt;p&gt;I'm continuing to work on MandAPI as both an API gateway and a model/pricing research project.&lt;/p&gt;

&lt;p&gt;One area I'm particularly interested in is transparent LLM price comparison.&lt;/p&gt;

&lt;p&gt;I've started publishing some of the underlying developer resources and pricing research openly on GitHub:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/mandapi-ai/developer-kit" rel="noopener noreferrer"&gt;MandAPI Developer Kit&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The goal is to make it easier for developers to compare models based on actual API economics rather than marketing pages.&lt;/p&gt;

&lt;p&gt;I'm also especially interested in making AI APIs easier to purchase and use in markets where international billing can be inconvenient, including Brazil.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final thought
&lt;/h2&gt;

&lt;p&gt;You don't necessarily need to rewrite your AI application to experiment with different providers.&lt;/p&gt;

&lt;p&gt;An OpenAI-compatible abstraction layer can make model switching surprisingly simple.&lt;/p&gt;

&lt;p&gt;And once switching becomes simple, &lt;strong&gt;price becomes something you can optimize continuously instead of something you're locked into.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If you're building an AI product with significant API usage, it's worth designing for model portability from the beginning.&lt;/p&gt;

&lt;p&gt;I'd be interested to hear how other developers are handling multi-model routing and API costs.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
    </item>
  </channel>
</rss>
