DEV Community

Cover image for Using the right AI should be the easy part
BekahHW
BekahHW

Posted on Originally published at papercompute.com

Using the right AI should be the easy part

#ai

Picture an office building with no front desk. Everyone who works there has their own key to every door. Visitors walk in off the street and find a room. Contractors let themselves in at night. Nobody knows who is in the building, which rooms are busy, or what any of it costs until the utility bill arrives.

That's most companies' AI setup right now. Every tool that talks to a model has its own key and its own route out the door. Harness asked 700 engineering leaders about it this year. Most companies know they need to do something about AI spend, but they don't have the visibility to make educated decisions.

harness report data as a pie chart

The fix is a front desk. In AI infrastructure it's called an AI gateway: one door every request goes through on the way to a model. Everyone still gets to their meeting. Someone just wrote down who came in, where they went, and when they left.

The desk is step one. Two more layers make it smarter, and each one is built on the one below it. Start with the simplest version of the question: what does a front desk know by the end of its first week that nobody in your building can tell you right now?

An AI gateway gives you one place to make decisions

Without a gateway, every AI tool talks directly to a model provider. Right now, every tool is like its own island. It has its own provider connection, spend, rules, decisions about what model to use. A gateway gives you a place to control all of that.

Every AI request passes through the same point (the gateway), so you can see which tools and teams are making requests, track spending, apply policies across teams, and decide where each request should go.

Think of it like this: all your AI tools → one control point → the right model.

That gives you one place every request passes through. Now you can see which tools and teams are making requests, track what they're spending, apply the same policies across all of them, and decide where each request should go. The tools don't change. The base URL does.

An AI gateway directs where the request should go. Maybe that's OpenAI. Maybe it's Anthropic. Maybe one provider is unavailable this afternoon and you need another. The gateway makes that call once, for everyone, instead of every tool carrying its own version of the logic.

teams going through the gateway to the right model/provider

Your teams don’t all need the same model, provider, budget, or policy. A gateway lets you make those decisions once instead of asking every engineer, team member, and tool to make them independently. The goal is to take that mental load away from the people doing the work.

Policies might look like this:

  • Product gets a fast, capable default model.
  • Platform is allowed to use the strongest models for harder infrastructure work.
  • Support gets a cheaper model for high-volume summarization.
  • A sensitive workload can be restricted to a particular provider.
  • Everyone can have different budgets, rate limits, or fallback behavior.

The person setting up the gateway defines the guardrails once. After that, they mostly disappear. An engineer opens their agent and gets to work. They don’t have to remember which model they’re supposed to use, which provider is approved, or whether a particular task is worth the more expensive option. The system already knows the rules. The right way becomes the easy way.

An AI gateway can route based on who is asking. A semantic gateway can route based on what they’re asking it to do.

A semantic gateway asks what the request needs

A semantic gateway looks at the request itself and asks: "what kind of work is this?"

semantic gateway separating the types of work to different powered models

A simple question goes to a smaller, cheaper model. A hard coding problem goes to a stronger one. The person using the agent doesn't have to stop and choose.

That matters more than it sounds, because today the choosing falls on the person using the AI tool. The problem isn’t necessarily that one person uses up the team’s AI budget (though it certainly can be). But if every engineer is being asked to spend their own allowance wisely that means deciding, task by task, which model is worth the cost.Developers have been posting about exactly this since the weekly caps arrived.

Cutting model spend is the smaller win. The bigger one is that you can give your best engineers the most capable models when that capability actually makes their work better, without paying that price for every request the company makes. Using the cheapest model everywhere should never be the goal. Using the right amount of intelligence for each piece of work is.


"The employee makes fewer decisions. The organization makes better decisions."

That’s a much better system than asking every engineer to make that calculation themselves.
A semantic router can recognize that a task looks simple or complex. It can choose a model based on that classification. What it can’t know from the request alone is whether the cheaper model struggled, whether the stronger model was overkill, or whether the engineer had to step in and fix the result.

Choosing the right model is only half the problem. You also need to know whether it was the right choice.

The bigger opportunity is an intelligence platform you own

What happens when a system can learn from the work your team is already doing? Every agent session contains evidence: what someone tried, what worked, what failed, what they corrected, which model handled it, which tools were called, and what finally got the job done.

The record matters after the request is over. Most of a gateway's value is synchronous: what should happen to this request, right now? There's an asynchronous value too: what can we learn from everything that happened this month? Once the desk is keeping the record, the traces, the sessions, the corrections people typed, the skills that got written, the evaluations, and the dreams and inceptions an agent runs on its own history stop being separate features. They become one loop: the next decision uses what happened last time. Continue, retry, hand the work to another agent, or stop and ask a person before spending another $16 on a fix that already failed once.

A reflection identifies what mattered in one session. A dream looks across those reflections for patterns: a failure your team keeps hitting, an approach that consistently works, work that doesn't actually need your strongest model. An inception takes one of those lessons and tests it against the original work to see whether applying it actually changes the outcome.

In other words, the goal isn't to have an AI system invent advice about how your team should work. It's to learn from what your team has already proven. And then you can start thinking about where that knowledge belongs. Maybe a repeated workflow becomes a skill. Maybe a pattern about successful work becomes context an agent can retrieve. Maybe evidence that a smaller model performs just as well changes a routing rule. Maybe a known failure causes an agent to try a different approach, or stops it before it burns another $16 doing the same thing again.

We aren't at a point where all of that happens automatically, and I don't think it should happen without evidence. But the pieces start to connect. First, remove friction by putting policy and routing behind the scenes. Then you learn from the work. Then you use what you've learned to remove the next bit of friction.

Where to start if your AI requests still skip the front desk

The first layer is one base-URL change. Three moves, in order:

  • 01. Put the front desk in. Even with one provider. Our gateway docs cover the setup, and the maturity model shows what each level buys. The bill-by-team answer arrives the first month.
  • 02. Keep the full session record. A request log tells you what was sent. The record of what came back, and what the person said next, is what the third layer runs on.
  • 03. Read a week of follow-ups. If more than one in six is someone saying "still failing," you have the same gap we do, and now you have the number to show whoever owns the budget.

The goal isn’t cheaper AI. It’s spending intelligence where it matters.

Top comments (0)