DEV Community

SarangAI
SarangAI

Posted on

How I Built a Unified API Gateway for 200+ AI Models (Architecture Deep Dive)

Every developer who has worked with more than one AI provider knows the pain:

  • OpenAI has its own API key
  • Anthropic has its own billing dashboard
  • Google has its own SDK
  • DeepSeek has its own rate limits

If you want to use 4 different models, you have to manage 4 accounts, 4 invoices, and 4 sets of documentation. This isn't just annoying — it's a bottleneck that stops developers from experimenting with new models.

I built SarangAI to solve this. It's a unified AI gateway that routes all requests to 200+ models through a single OpenAI-compatible endpoint.

In this article, I'll walk through the architecture behind it, the design decisions I made, and the technical challenges that came up.

High-Level Architecture

At a high level, SarangAI consists of 4 main components:

  1. API Gateway — receives requests from clients (CLI, IDE, or any app)
  2. Router — decides which model handles the request
  3. Provider Adapters — translate OpenAI-compatible format to each provider's format
  4. Billing & Rate Limiter — manages prepaid balance and per-user rate limits

The flow is simple:

Client → API Gateway → Router → Provider Adapter → AI Provider
                ↓
          Billing & Rate Limiter
Enter fullscreen mode Exit fullscreen mode

Why OpenAI-Compatible?

This is the most important design decision I made.

When I started building SarangAI, I had two options:

Option 1: Build my own API format, my own docs, my own SDK.
Option 2: Use the OpenAI format, which has become the de-facto standard.

I chose option 2. Here's why:

  • Zero migration — if your code already uses the OpenAI SDK, just change the base URL and API key. Done.
  • Massive ecosystem — thousands of libraries and tools already support the OpenAI format.
  • Familiar — developers don't need to learn a new API.

This is what makes SarangAI usable in minutes, not hours.

Technical Challenge #1: Normalizing Responses

Every provider has a different response format. OpenAI has choices[0].message.content, Anthropic has content[0].text, Google has yet another structure.

The solution: adapter pattern. Each provider has an adapter that:

  1. Translates the request from OpenAI format to the provider's format
  2. Translates the response from the provider's format back to OpenAI format
  3. Handles errors and retry logic

This keeps client-side code clean — they don't need to know which provider is being used.

Technical Challenge #2: Streaming

Streaming responses are tricky. Every provider sends chunks differently:

  • OpenAI uses Server-Sent Events (SSE) with data: {...}
  • Anthropic has its own event types (content_block_delta, etc.)
  • Google has yet another streaming format

In SarangAI, I normalize all streaming to the same SSE format as OpenAI. So clients only need to handle one streaming format.

Technical Challenge #3: Rate Limiting & Billing

Since SarangAI uses a prepaid IDR top-up model, I need to:

  1. Track every request and calculate cost based on token usage
  2. Deduct from the user's balance in real-time
  3. Handle race conditions with concurrent requests

For this, I use a combination of Redis (for fast balance checks) and a database (for audit trails).

Technical Challenge #4: Model Routing

One of SarangAI's main features is instant model switching. Users can change models without restarting their app.

This means the router has to:

  • Read the model config from the request
  • Validate that the model is available
  • Route to the correct adapter
  • Handle fallback if a provider is down

I made this router stateless, so it can scale horizontally without issues.

The CLI: sarangai-cli

Besides the API gateway, I also built a CLI tool that works directly from the terminal:

npm install -g sarangai-cli
sarang
Enter fullscreen mode Exit fullscreen mode

The CLI connects to the SarangAI endpoint and gives you an interactive workspace. You can:

  • Switch models with a single command
  • See token usage
  • Check your balance

This is especially useful for developers who live in the terminal.

Lessons Learned

1. Standards matter.
Choosing the OpenAI-compatible format was the best decision I made. It's what makes adoption fast.

2. Adapter patterns save lives.
Without clean adapters, adding a new provider would be a nightmare.

3. Prepaid > Subscription for developer tools.
Developers hate monthly subscriptions. Prepaid gives them a sense of full control.

4. Documentation is a feature.
No matter how good your architecture is, if the docs are bad, nobody will use it.

Try It Yourself

If you work with multiple AI models regularly, give SarangAI a try:

🔗 https://sarangai.id

Install the CLI:

npm install -g sarangai-cli
Enter fullscreen mode Exit fullscreen mode

I'm curious: what's your current setup for handling multiple AI providers? Do you use a library? Or manage them one by one?

Share in the comments - I'd love to hear how other developers handle this.

Top comments (0)