On my first AI feature I called one provider straight from the code that needed it. It worked, and for a while that was the right call.
Then I wanted a cheaper model for the boring tasks, like tagging and short summaries. That's when I noticed the provider was everywhere: keys in three places, retry logic copied around, and no clear idea what each feature actually cost.
Chalk's post this week cited an a16z survey where 81% of CIOs at Global 2000 companies now use three or more model families. I'm nowhere near that scale, but even with two providers the mess shows up fast.
What the layer does
It's small. Roughly:
- One function per task type, like
summarizeorclassify, not per provider. - The model for each task comes from config, so switching is a one line change.
- A timeout and one fallback model, so a slow provider doesn't hang a user request.
- A daily spend limit per key. I got this one after a retry loop ran all night once.
- A log line per request with task, model, tokens and latency.
What it cost me
Another thing to maintain. Tool calling and structured output formats differ between providers, so the abstraction leaks. I ended up with small adapter code per provider anyway.
It also tempts you to over build. I'd skip it until a second provider is actually on the table.
What I'd do again
The per task log. Seeing that one feature used most of the tokens changed what I worked on next more than any routing trick did.
If you're running more than one model, did you build your own layer or use a gateway?
Top comments (2)
the daily spend limit is the hardest lesson to learn. a runaway retry loop on a 429 rate limit can drain a month's api budget while you sleep.
i built a similar routing layer using a cloudflare edge worker to dynamically switch between three models based on task complexity (a 120b general coder, a 20b speed fallback, and an r1 reasoning specialist). and your point about the abstraction leaking is spot on. trying to normalize tool calling and structured json outputs across different providers (like groq vs deepseek) is where the adapter code always gets messy.
but the per-task logging is the real roi. seeing exactly which routing path burned the most tokens completely changed how i optimized my prompt budgets.
curious about your adapter code: did you end up using a strict schema validator (like zod) in your layer to catch the structured output leaks before they hit your app, or do you just let the app handle the parsing fallback? 🐯
I started with a home-grown wrapper in Python and hit the exact same schema leak you did. Basic text generation and token counts normalize easily, but tool definitions and error codes differ just enough between APIs that a single generic schema wrapper breaks down.
What ended up working best was keeping the layer strictly about routing and telemetry, while leaving serialization to the caller. The wrapper handles the timeout, the fallback trigger, and the per-task SQLite log. If provider A throws a schema error or 429, the fallback catches it, but the wrapper doesn't try to normalize the tool payload itself. That kept the helper under two hundred lines and avoided building a third-party gateway clone from scratch.