DEV Community

T. Alam
T. Alam

Posted on

We Stopped Hardcoding a Model Provider Into Every Feature. Here's the Router That Fixed It

A couple of months ago we had four different features in production, each one calling a different model provider directly, each one with its own little pile of error handling, its own retry logic, its own idea of what a "timeout" should mean. OpenAI for the chat widget. A cheaper Hugging Face model for the classification step nobody thought about until the AWS bill showed up. Claude for anything that touched a document. It worked, technically. It also meant that every time we wanted to try a new model, or a cheaper one, or route around an outage, someone had to go change code in four places and redeploy four things.

The fix wasn't "pick one provider and stick with it forever." It was building a router — one small agent whose entire job is deciding where a request should go, sitting in front of everything else. This post is that router, built with DNotifier's actual defineAgent and Workflow primitives, not a diagram with boxes and arrows that doesn't compile.

The problem with "just pick the best model"

Here's the thing nobody tells you until you've shipped a few of these systems: most of your traffic doesn't need your best model. We looked at three weeks of real requests hitting one of our internal tools and something like 80% of them were dead simple — short factual lookups, basic rewrites, one-line classifications. The other 20% actually needed real reasoning. Running everything through the expensive model because some requests need it is just burning money on the easy majority.

So the router's job isn't "find the best model." It's "figure out what this specific request actually needs, then send it to the cheapest thing that can handle it."

What we're building

One classifier agent that looks at an incoming request and decides how hard it is. Two (or more) answer agents behind it, each pointed at a different model. A workflow that ties them together and logs every routing decision so you can actually see what happened later, instead of guessing.

request → [classifier agent] → "simple" or "complex"
                                      │
                    ┌─────────────────┴─────────────────┐
                    ▼                                     ▼
           [small/cheap model]                   [larger/capable model]
                    │                                     │
                    └─────────────────┬─────────────────┘
                                       ▼
                                   response
                                (logged + observable)
Enter fullscreen mode Exit fullscreen mode

The routing decision is itself a small, cheap model call — not a guess encoded in application logic.

Step 1: the classifier

This is the part people over-engineer. You don't need a big model to decide if a question is simple — you need a small, fast one that's good at exactly one narrow job.

import { DNotifier } from "@dnotifier-realtime/dnotifier";

const routerAgent = DNotifier.defineAgent({
  name: "complexity-router",
  model: "huggingface/meta-llama/Llama-3.1-8B-Instruct:fastest",
  async run(ctx) {
    const classification = await ctx.sendAI({
      message: {
        text: `Classify this request as "simple" or "complex".
        Simple: factual lookups, short rewrites, basic classification.
        Complex: multi-step reasoning, long-document analysis, nuanced judgment calls.
        Respond with exactly one word.

        Request: "${ctx.input.text}"`,
      },
    });
    ctx.state.complexity = classification.text.trim().toLowerCase();
    return ctx.state.complexity;
  },
});
Enter fullscreen mode Exit fullscreen mode

Notice the :fastest hint on the model string. That's us telling DNotifier we care more about latency than shaving fractions of a cent on this particular call — it's an 8B model running one classification, it should come back almost instantly. We'll flip that priority for the expensive path in a second.

One thing worth being honest about: this is a soft classifier, not a guarantee. Occasionally it'll call something "simple" that really wasn't. That's fine — we'll get to why that's a recoverable problem later, not a fatal one.

Step 2: the two answer paths

const simpleAnswerAgent = DNotifier.defineAgent({
  name: "simple-answer-agent",
  model: "huggingface/meta-llama/Llama-3.1-8B-Instruct:fastest",
  async run(ctx) {
    const answer = await ctx.sendAI({ message: { text: ctx.input.text } });
    ctx.state.answer = answer.text;
    return answer.text;
  },
});

const complexAnswerAgent = DNotifier.defineAgent({
  name: "complex-answer-agent",
  model: "huggingface/openai/gpt-oss-120b:cheapest",
  async run(ctx) {
    const answer = await ctx.sendAI({ message: { text: ctx.input.text } });
    ctx.state.answer = answer.text;
    return answer.text;
  },
});
Enter fullscreen mode Exit fullscreen mode

See the difference in routing hints — :fastest on the small model, :cheapest on the big one. That's not a copy-paste mistake. On the small model, latency is what matters, and the cost difference between backing providers is basically noise. On the bigger, pricier call, the cost spread between backing providers is actually worth optimizing for, so we tell DNotifier to go find the cheapest route to that model instead of the fastest one. Small detail, but it adds up once you're running this at any real volume.

Step 3: wire it into a workflow

const routingWorkflow = new DNotifier.Workflow({
  name: "complexity-based-model-router",
  description: "Classifies request complexity and routes to an appropriately sized model",
  observability: true,
  async entry(ctx) {
    await ctx.agents.run(routerAgent, { text: ctx.input.text });

    if (ctx.state.complexity === "complex") {
      await ctx.agents.run(complexAnswerAgent, { text: ctx.input.text });
    } else {
      await ctx.agents.run(simpleAnswerAgent, { text: ctx.input.text });
    }

    return { complexity: ctx.state.complexity, answer: ctx.state.answer };
  },
});

routingWorkflow.registerAgents([routerAgent, simpleAnswerAgent, complexAnswerAgent]);
Enter fullscreen mode Exit fullscreen mode

observability: true is doing more work here than its one line suggests. Once this is live, every routing decision — what came in, what it got classified as, which model actually answered — shows up in the dashboard as it happens. The first time a teammate asked "why did the bot give a weak answer to this," we didn't have to reproduce anything. We opened the run and looked.

Step 4: actually run it

const notifier = new DNotifier({
  appId: process.env.DNOTIFIER_APP_ID,
  secret: process.env.DNOTIFIER_SECRET,
  userId: "routing-system",
  transport: "ws",
  WebSocketImpl: WebSocket,
});

await notifier.connect();

const outcome = await notifier.runWorkflow(routingWorkflow, {
  text: "What year did the Berlin Wall fall?",
});

console.log(outcome.complexity, "→", outcome.answer);
// simple → 1989
Enter fullscreen mode Exit fullscreen mode

Same workflow, harder question:

const outcome2 = await notifier.runWorkflow(routingWorkflow, {
  text: "Given these three quarterly reports, what's driving the margin decline and is it structural or seasonal?",
});

console.log(outcome2.complexity, "→", outcome2.answer);
// complex → [a real, reasoned answer from the larger model]
Enter fullscreen mode Exit fullscreen mode

Nobody flipped a switch between those two calls. The router looked at the second question, decided it actually needed reasoning depth, and sent it somewhere else — automatically.

Where this stops being a toy example

The version above routes between two Hugging Face model sizes because that's a clean example, but nothing about defineAgent ties you to one provider. This is the part that actually changed how we think about this stuff: because every agent is just a name, a model string, and a run function, the exact same pattern routes across any combination of providers you've got connected.

We eventually extended this same shape to include a local model running through Ollama, for one specific category of request where the data genuinely couldn't leave our own infrastructure:

const localOnlyAgent = DNotifier.defineAgent({
  name: "local-only-agent",
  provider: "ollama",
  model: "llama3.2",
  async run(ctx) {
    const answer = await ctx.sendAI({ message: { text: ctx.input.text } });
    ctx.state.answer = answer.text;
    return answer.text;
  },
});
Enter fullscreen mode Exit fullscreen mode

Same defineAgent shape. Same workflow wiring. The only thing that changed is which agent gets picked, and now the router isn't just choosing between "cheap" and "expensive" — it's choosing between "cloud" and "stays on our own hardware," based on whatever classification rule you want to write into the router's prompt.

The same pattern, generalized past two model sizes into any mix of providers you've got connected — cloud or local.

We'll get into the specifics of running Ollama in production (and the very dumb networking mistake we made the first time) in the next post.

Why we bothered with the extra agent instead of an if/else in application code

We got asked this internally more than once — why not just check input.length > 200 or some heuristic in plain code and skip the model call entirely? Two honest reasons. First, a length check is a terrible proxy for complexity — "what's 2+2" and a three-sentence question requiring real synthesis can be the same length. Second, the classifier call is cheap enough (one small, fast model call) that the cost of running it on every single request, even the ones that end up routed to the expensive model anyway, is trivial next to what it saves on the 80% that don't need the expensive model at all.

And when the router does misclassify something — it happens, rarely — it's a soft failure, not a crash. A complex question that gets routed to the simple model comes back with a weaker answer, not an error. We watch the classification and the answer quality together over time through the observability dashboard, and when the pattern shows the router being too aggressive about sending things downmarket, we adjust the prompt. It's tuning, not firefighting.

If you're juggling more than one provider right now and every new one means another SDK, another auth flow, another set of error codes to learn — this pattern is worth the hour it takes to set up. It paid for itself for us within the first week.

Top comments (0)