DEV Community

Antonio Berben
Antonio Berben

Posted on

Route each prompt to the right model, with a decision model in your cluster and one agentgateway policy

A hands-on lab, every file included. A 170-line ExtProc asks a decision model which tier should answer, and agentgateway routes on the reply. Runs on kind, on a CPU, and the prompts stay home.

Classify prompts with Laya, the open Jev; route them with agentgateway

Every platform team I talk to lands in the same meeting sooner or later. The agents send everything to the biggest model. Finance wants the bill down. Somebody says "route by prompt", somebody else opens the vLLM Semantic Router docs and finds a full platform: classifiers, a semantic cache, tooling. It is a good project, and if you need what it does, it is the right answer. This team needed one decision, and that is a smaller problem.

The decision is small. Does this prompt need the strong model, or will the fast one do? Written down like that, it is a classification with two labels. And this September a new kind of model showed up that does exactly that, and nothing else.

So this lab is the small version. One ExtProc, 172 lines of Go, that reads the prompt, asks a decision model which tier should answer, and hands the answer to agentgateway as a header. agentgateway does the routing, the way it already routes everything else. Everything you need is in this article: the Go file, five manifests, and the commands in order. Nothing to clone. You will run it on kind, on your laptop, with the classifier on a CPU inside the cluster.

One thing to be clear about before the first command. The classifier is real: Laya runs in the cluster and every probability in this article came out of it. The two models it routes to are not. They are two nginx pods that answer any chat completion with a fixed JSON naming their tier. That is not to save money. It is so you can see, in the response body, which model the gateway actually sent the request to. A real model answers the question, which tells you nothing about the route it took; a stand-in that says [strong-model] answered does. Swap them for vLLM or a hosted provider at the end and nothing in the routing changes.

A model that does not write

On 15 September 2026 TypeSafe AI released Jev, which they call a System One model. You send it a block of state and one or more typed questions, and it answers all of them in one pass with a choice, a score or a yes-or-no probability. It never generates a sentence. Output tokens are free because there are none. The hosted API is a single endpoint, POST /v1/systemone, and the response carries probabilities and a confidence for every question you asked.

Within a week the open weights community had built the same shape on top of open models. The one I use here is Laya: a ModernBERT-large classifier of 421 million parameters, Apache 2.0, with a built-in server that exposes the same /v1/systemone contract. It runs on a CPU. On my MacBook, inside kind, a decision plus the whole gateway round trip takes about 200 ms.

Why does this matter for routing? Three properties, and they are the three things a router wants:

  • The answer is one of the labels you sent. There is no text to parse and no way to get a label you did not define.
  • You get a probability for every label, so "not sure" is a number you can compare against a threshold instead of a feeling.
  • It is small. The question is short, the answer is a few floats, and there is no reason for it to cost more than a lookup.

A decision model answers a typed question with a probability per label; it never writes text

The request looks like this. The state is the last user message. The question is a choice with two criteria, and the criteria are plain English descriptions of what each tier is for:

{
  "model": "jev-latest",
  "state": "Design a multi-region rate limiter that stays consistent under network partition, compare token bucket with sliding window, then implement it in Go with unit tests.",
  "questions": {
    "tier": {
      "type": "choice",
      "instructions": "Which model tier should answer this user request?",
      "criteria": {
        "fast": "Short factual questions, greetings, translations, one-line definitions, rewording a sentence. Anything a small model answers well in a few lines.",
        "strong": "Multi-step reasoning, writing or debugging code, proofs, legal, financial or security analysis, long structured documents. Anything where a wrong answer is costly."
      }
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

And this is the shape of what comes back, from Laya, inside the cluster:

{
  "model": "laya-rl-agent",
  "answers": {
    "tier": {
      "type": "choice",
      "choice": "strong",
      "probabilities": {"fast": 0.259, "strong": 0.741},
      "confidence": 0.16
    }
  },
  "usage": {"input_tokens": 127, "output_tokens": 0}
}
Enter fullscreen mode Exit fullscreen mode

The router reads choice and the probability of that choice. Jev also returns a calibrated confidence in the 0 to 1 range; Laya's confidence lives on its own scale, so the probability is the signal that works for both, and it is the one the router compares.

Fast or strong is one question, not the only one

I picked cost tiers for this lab because they are the easiest to see, and for a while that was all I had in mind. It was a conversation with my colleague Rinor Maloku that widened it. The router only knows labels and their descriptions; it has no idea what "fast" or "strong" mean. So the labels can be anything: effort, domain, task type. Write the descriptions, and the same router classifies on that instead. That is what makes "model": "auto" worth having. It no longer means "cheap or expensive". It means "the right model for this request". Some axes I have seen teams need:

  • Effort. trivial, moderate, deep: how much reasoning the answer needs, which is what decides whether a reasoning model earns its latency.
  • Task type. code, prose, data, conversation: send code to a coding model and a customer reply to a model that writes well, instead of asking one model to do both.
  • Domain. legal, finance, support, general: the tier is a fine-tuned or RAG-backed model per domain, and the router is the only place that needs to know the mapping.
  • Sensitivity. contains-personal-data, contains-credentials, clean: not a tier at all, but a label the route can use to keep some prompts on the model that never leaves the cluster.
  • Language. es, en, de: route to the model that is strong in the user's language, or add the label as a header your telemetry can group by.

One request to the decision model can carry several of these questions at once; they are evaluated in parallel and isolated from each other. This lab asks one, so the plumbing stays visible. Adding a second is another entry in the questions map and another header.

Where the decision has to happen

Gateway API routes on paths and headers. It has no body matcher, by design, because a route is supposed to be cheap. So a decision made from the prompt has to become a header before the route is chosen, or it changes nothing.

agentgateway, one of the open source projects under the Agentic AI Foundation, has a phase for exactly this. A traffic policy attached to the Gateway with phase: PreRouting runs its filters before route selection. Put an extProc there, let it read the buffered body and write a header, and the HTTPRoute that comes next sees the header as if the client had sent it.

The ExtProc runs in PreRouting and writes x-selected-model before the HTTPRoute matches

Six steps, and the agent sees none of them:

  1. The agent sends {"model": "auto", "messages": [...]} to /v1/chat/completions.
  2. agentgateway buffers the body and streams it to the router over gRPC, the Envoy ext_proc protocol.
  3. The router takes the last user message and asks the decision model the choice question above.
  4. The router writes x-selected-model: strong-model on the request, plus three x-decision-* headers that explain itself.
  5. The HTTPRoute matches on x-selected-model and picks the AgentgatewayBackend for that tier.
  6. The backend pins the model name and rewrites it into the body, so the client's auto never reaches a model server.

Two fields in the policy carry all of this. phase: PreRouting, because at the default PostRouting the backend is already chosen when the header appears. And requestBodyMode: Buffered, because the router needs the whole prompt before it can answer.

One router, and the URL decides where your prompts go

The router speaks one contract, so where the decision model lives is a deployment choice, not a code change. DECISION_URL points at api.typesafe.ai and you are using Jev, hosted in the United States. It points at a Service in your namespace and you are using Laya, and nothing leaves the cluster.

That choice is worth a paragraph, because it is easy to make by accident. A router reads every prompt by construction. Not a sample, not the ones that fail: all of them, before any policy has looked at them. Whatever the router sends to its classifier, it sends for one hundred percent of your traffic. With Jev that is the last user message, without the system prompt and without the history, under TypeSafe's data processing addendum with EU standard contractual clauses. With Laya it is an HTTP call to the pod next door.

For me, in Europe, that is the whole reason the lab runs Laya. Not because the hosted option is worse at the job. Because the classifier is the one component that sees everything, and I want to decide where it sits with my eyes open. The same binary, the same policy and the same route serve both, and switching is one variable and one key.

One router, two destinations: Laya in the cluster or Jev hosted, chosen by DECISION_URL

What you need

  • Docker, kind, kubectl, helm, jq and curl. No Go toolchain: the router compiles inside its pod.
  • About 6 GB of RAM free for kind, most of it for Laya.
  • Twenty minutes. The router pod fetches two Go modules and compiles (about two minutes); the Laya pod installs its package and downloads the checkpoint from Hugging Face (about five).
  • No API key of any kind.

Everything below goes into the agentgateway-system namespace. Save each YAML block to the file name in its heading, or paste it into kubectl apply -f -.

Step 0: a cluster and a gateway

kind create cluster --name jev-router

kubectl apply --server-side --force-conflicts \
  -f https://github.com/kubernetes-sigs/gateway-api/releases/download/v1.6.1/standard-install.yaml

helm upgrade --install agentgateway-crds oci://cr.agentgateway.dev/charts/agentgateway-crds \
  --version 1.5.0 --create-namespace --namespace agentgateway-system
helm upgrade --install agentgateway oci://cr.agentgateway.dev/charts/agentgateway \
  --version 1.5.0 --namespace agentgateway-system --wait
Enter fullscreen mode Exit fullscreen mode

One Gateway, one HTTP listener. Everything in the lab is reached through a port-forward, nothing is exposed. 00-gateway.yaml:

apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
  name: agentgateway-proxy
  namespace: agentgateway-system
spec:
  gatewayClassName: agentgateway
  listeners:
  - name: http
    port: 80
    protocol: HTTP
    allowedRoutes:
      namespaces:
        from: Same
Enter fullscreen mode Exit fullscreen mode
kubectl apply -f 00-gateway.yaml
Enter fullscreen mode Exit fullscreen mode

Step 1: two tiers

The two tiers are stand-ins on purpose: an nginx that answers every chat completion with a fixed OpenAI-shaped JSON naming its tier. llm-fast answers [fast-model] answered, llm-strong answers [strong-model] answered. The point is not to avoid a provider bill. It is that the answer itself becomes the evidence of the route: when the body says strong-model, the request went to the strong tier, and there is no other way it could have got there. A real model would answer the question, and the route would stay invisible. Replace them with vLLM, Ollama or a hosted provider once you trust the routing; the policy, the route and the router do not change.

The fast tier is a ConfigMap, a Deployment, a Service and one AgentgatewayBackend. 10-tier-fast.yaml:

# Two stand-in model servers. Each answers every chat completion with a fixed
# OpenAI-shaped JSON that names its tier, so you can see which one the gateway picked.
apiVersion: v1
kind: ConfigMap
metadata:
  name: llm-fast
  namespace: agentgateway-system
data:
  default.conf: |
    server {
      listen 8080;
      location / {
        default_type application/json;
        return 200 '{"id":"mock","object":"chat.completion","model":"fast-model","choices":[{"index":0,"message":{"role":"assistant","content":"[fast-model] answered"},"finish_reason":"stop"}],"usage":{"prompt_tokens":1,"completion_tokens":1,"total_tokens":2}}';
      }
    }
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: llm-fast
  namespace: agentgateway-system
spec:
  selector: {matchLabels: {app: llm-fast}}
  template:
    metadata: {labels: {app: llm-fast}}
    spec:
      containers:
      - name: nginx
        image: nginx:1.27-alpine
        ports: [{containerPort: 8080}]
        volumeMounts: [{name: conf, mountPath: /etc/nginx/conf.d}]
      volumes: [{name: conf, configMap: {name: llm-fast}}]
---
apiVersion: v1
kind: Service
metadata:
  name: llm-fast
  namespace: agentgateway-system
spec:
  selector: {app: llm-fast}
  ports: [{port: 8080}]
---
# One gateway backend per tier. `model` is both the value the router writes into
# x-selected-model and the value the gateway rewrites into the body before forwarding.
apiVersion: agentgateway.dev/v1alpha1
kind: AgentgatewayBackend
metadata:
  name: llm-fast
  namespace: agentgateway-system
spec:
  ai:
    provider:
      openai:
        model: fast-model
      host: llm-fast.agentgateway-system.svc.cluster.local
      port: 8080
Enter fullscreen mode Exit fullscreen mode

The strong tier is the same file with the names swapped, so let sed write it:

kubectl apply -f 10-tier-fast.yaml
sed 's/fast-model/strong-model/g; s/llm-fast/llm-strong/g' 10-tier-fast.yaml | kubectl apply -f -
Enter fullscreen mode Exit fullscreen mode

The backend's model field does two jobs. It is the value the router will write into x-selected-model, and it is the value agentgateway substitutes into the body's model field before forwarding. Same name in three places: the router's table, the backend, and the HTTPRoute.

Step 2: Laya, in the cluster

A python:3.11-slim pod that installs laya[serve] with a CPU torch wheel at start and runs laya-serve. 20-laya.yaml:

# Laya (ModernBERT, Apache 2.0) behind its Jev-compatible server, on CPU.
# The pod installs the package at start and downloads the checkpoint from
# Hugging Face; the startup probe gives it up to 25 minutes.
apiVersion: apps/v1
kind: Deployment
metadata:
  name: laya
  namespace: agentgateway-system
spec:
  progressDeadlineSeconds: 1800
  selector: {matchLabels: {app: laya}}
  template:
    metadata: {labels: {app: laya}}
    spec:
      # The Service is called laya, so Kubernetes would inject LAYA_PORT=tcp://...
      # and laya-serve reads that as its port. Links off, port explicit.
      enableServiceLinks: false
      containers:
      - name: laya
        image: python:3.11-slim
        command: ["sh", "-c"]
        args:
        - pip install --quiet torch --index-url https://download.pytorch.org/whl/cpu &&
          pip install --quiet "laya[serve]==0.3.21" && exec laya-serve
        env:
        - {name: LAYA_HOST, value: "0.0.0.0"}
        - {name: LAYA_PORT, value: "8000"}
        - {name: LAYA_DEVICE, value: cpu}
        - {name: LAYA_MODELS, value: english}
        - {name: LAYA_PRELOAD, value: "1"}
        - {name: LAYA_THREADS, value: "4"}
        - {name: USE_TF, value: "0"}
        - {name: HF_HOME, value: /cache}
        ports: [{containerPort: 8000}]
        startupProbe: {httpGet: {path: /health, port: 8000}, periodSeconds: 15, failureThreshold: 100}
        readinessProbe: {httpGet: {path: /health, port: 8000}, periodSeconds: 10}
        resources: {requests: {cpu: "1", memory: 2Gi}, limits: {memory: 4Gi}}
        volumeMounts: [{name: cache, mountPath: /cache}]
      volumes: [{name: cache, emptyDir: {sizeLimit: 6Gi}}]
---
apiVersion: v1
kind: Service
metadata:
  name: laya
  namespace: agentgateway-system
spec:
  selector: {app: laya}
  ports: [{port: 8000}]
Enter fullscreen mode Exit fullscreen mode
kubectl apply -f 20-laya.yaml
Enter fullscreen mode Exit fullscreen mode

Three settings earn their place. enableServiceLinks: false, because the Service is called laya and Kubernetes would otherwise inject LAYA_PORT=tcp://10.96.x.x:8000 into the pod, which laya-serve reads as its listen port; the port is set explicitly instead. LAYA_PRELOAD loads the checkpoint before the server opens its port, so /health doubles as a readiness signal. And the startupProbe gives the install and the first download room. Start it now and let it work while you do the next two steps.

Step 3: the router

This is the whole ExtProc. Save it as main.go:

// decision-router: an agentgateway ExtProc that asks a decision model which
// model tier should answer, and writes the answer as a request header.
package main

import (
    "bytes"
    "encoding/json"
    "fmt"
    "io"
    "log"
    "net"
    "net/http"
    "os"
    "strconv"
    "time"

    corev3 "github.com/envoyproxy/go-control-plane/envoy/config/core/v3"
    extprocv3 "github.com/envoyproxy/go-control-plane/envoy/service/ext_proc/v3"
    "google.golang.org/grpc"
)

// The routing table: label -> description the decision model reads, and the
// model name the gateway backend pins. Edit these, restart, routing changes.
var tiers = map[string]struct{ Description, Model string }{
    "fast":   {"Short factual questions, greetings, translations, one-line definitions, rewording a sentence. Anything a small model answers well in a few lines.", "fast-model"},
    "strong": {"Multi-step reasoning, writing or debugging code, proofs, legal, financial or security analysis, long structured documents. Anything where a wrong answer is costly.", "strong-model"},
}

var (
    decisionURL = os.Getenv("DECISION_URL")             // http://laya:8000 or https://api.typesafe.ai
    decisionKey = os.Getenv("DECISION_API_KEY")         // empty = no Authorization header
    defaultTier = env("DEFAULT_TIER", "fast")
    threshold, _ = strconv.ParseFloat(env("CONFIDENCE_THRESHOLD", "0.55"), 64)
    client      = &http.Client{Timeout: 2 * time.Second}
)

func env(k, d string) string {
    if v := os.Getenv(k); v != "" {
        return v
    }
    return d
}

// decide returns the headers for one request. It never fails: anything odd
// lands on the default tier with the reason in x-decision-reason.
func decide(body []byte) map[string]string {
    var req struct {
        Model    string `json:"model"`
        Messages []struct {
            Role    string `json:"role"`
            Content string `json:"content"`
        } `json:"messages"`
    }
    if err := json.Unmarshal(body, &req); err == nil && req.Model != "" && req.Model != "auto" {
        return map[string]string{"x-selected-model": req.Model, "x-decision-reason": "passthrough"}
    }
    prompt := ""
    for _, m := range req.Messages {
        if m.Role == "user" {
            prompt = m.Content
        }
    }
    tier, conf, reason := defaultTier, 0.0, "fallback"
    if prompt != "" {
        if label, p, err := ask(prompt); err != nil {
            log.Printf("decision api: %v", err)
        } else if _, ok := tiers[label]; ok && p >= threshold {
            tier, conf, reason = label, p, "classified"
        } else {
            conf, reason = p, "low-confidence"
        }
    }
    return map[string]string{
        "x-selected-model":      tiers[tier].Model,
        "x-decision-tier":       tier,
        "x-decision-reason":     reason,
        "x-decision-confidence": fmt.Sprintf("%.3f", conf),
    }
}

// ask sends one choice question to the /v1/systemone API and returns the label
// and the probability of that label.
func ask(prompt string) (string, float64, error) {
    criteria := map[string]string{}
    for name, t := range tiers {
        criteria[name] = t.Description
    }
    payload, _ := json.Marshal(map[string]any{
        "model": "jev-latest", "state": prompt,
        "questions": map[string]any{"tier": map[string]any{
            "type": "choice", "instructions": "Which model tier should answer this user request?", "criteria": criteria}},
    })
    req, _ := http.NewRequest("POST", decisionURL+"/v1/systemone", bytes.NewReader(payload))
    req.Header.Set("Content-Type", "application/json")
    if decisionKey != "" {
        req.Header.Set("Authorization", "Bearer "+decisionKey)
    }
    resp, err := client.Do(req)
    if err != nil {
        return "", 0, err
    }
    defer resp.Body.Close()
    var out struct {
        Answers struct {
            Tier struct {
                Choice        string             `json:"choice"`
                Probabilities map[string]float64 `json:"probabilities"`
            } `json:"tier"`
        } `json:"answers"`
    }
    raw, _ := io.ReadAll(resp.Body)
    if resp.StatusCode != 200 {
        return "", 0, fmt.Errorf("%s: %s", resp.Status, raw)
    }
    if err := json.Unmarshal(raw, &out); err != nil {
        return "", 0, err
    }
    return out.Answers.Tier.Choice, out.Answers.Tier.Probabilities[out.Answers.Tier.Choice], nil
}

func mutation(h map[string]string) *extprocv3.HeaderMutation {
    m := &extprocv3.HeaderMutation{RemoveHeaders: []string{"x-selected-model", "x-decision-tier", "x-decision-reason", "x-decision-confidence"}}
    for k, v := range h {
        m.SetHeaders = append(m.SetHeaders, &corev3.HeaderValueOption{Header: &corev3.HeaderValue{Key: k, RawValue: []byte(v)}})
    }
    return m
}

type server struct{ extprocv3.UnimplementedExternalProcessorServer }

// Process handles one HTTP exchange: request headers, request body, response headers.
func (server) Process(stream extprocv3.ExternalProcessor_ProcessServer) error {
    var decision map[string]string
    for {
        req, err := stream.Recv()
        if err != nil {
            return nil
        }
        var resp *extprocv3.ProcessingResponse
        switch v := req.Request.(type) {
        case *extprocv3.ProcessingRequest_RequestHeaders:
            // Strip the router's headers so a client cannot pick its own tier.
            resp = &extprocv3.ProcessingResponse{Response: &extprocv3.ProcessingResponse_RequestHeaders{
                RequestHeaders: &extprocv3.HeadersResponse{Response: &extprocv3.CommonResponse{HeaderMutation: mutation(nil)}}}}
        case *extprocv3.ProcessingRequest_RequestBody:
            decision = decide(v.RequestBody.Body)
            log.Printf("decision %v", decision)
            resp = &extprocv3.ProcessingResponse{Response: &extprocv3.ProcessingResponse_RequestBody{
                RequestBody: &extprocv3.BodyResponse{Response: &extprocv3.CommonResponse{HeaderMutation: mutation(decision), ClearRouteCache: true}}}}
        case *extprocv3.ProcessingRequest_ResponseHeaders:
            // Echo the decision on the response so the client can see it.
            resp = &extprocv3.ProcessingResponse{Response: &extprocv3.ProcessingResponse_ResponseHeaders{
                ResponseHeaders: &extprocv3.HeadersResponse{Response: &extprocv3.CommonResponse{HeaderMutation: mutation(decision)}}}}
        default:
            resp = &extprocv3.ProcessingResponse{}
        }
        if err := stream.Send(resp); err != nil {
            return err
        }
    }
}

func main() {
    lis, err := net.Listen("tcp", ":9090")
    if err != nil {
        log.Fatal(err)
    }
    srv := grpc.NewServer()
    extprocv3.RegisterExternalProcessorServer(srv, server{})
    log.Printf("decision-router on :9090, decision API %s, threshold %.2f", decisionURL, threshold)
    log.Fatal(srv.Serve(lis))
}
Enter fullscreen mode Exit fullscreen mode

Read it top to bottom once. The routing table is at the top: two labels, the description the decision model reads for each, and the model name the backend pins. The descriptions are the classifier's whole instruction set; the model was never trained on your tiers, so it needs to be told what they mean. Change them, restart the pod, and the routing changes. Swap them for code and prose, or legal and general, and nothing else in the file moves.

decide never fails the request. An explicit model other than auto is passed through untouched. A prompt goes to ask, which is the /v1/systemone call from the first section. A label at or above the threshold is classified; a label under it is low-confidence and goes to the default tier; a timeout or an error is fallback, also to the default tier. Every outcome is written to four headers so you can see it from the client.

Process is the ext_proc side. On the request headers phase the router strips every header it owns, so a client cannot pick its own tier by sending x-selected-model itself. On the request body phase it decides and returns a header mutation. On the response headers phase it echoes the same headers, which is how you read the decision without looking at logs.

The pod compiles this from a ConfigMap. 30-router.yaml:

# The router runs main.go straight from a ConfigMap: the pod fetches the two Go
# dependencies and compiles at start (about a minute). Build an image for production.
apiVersion: apps/v1
kind: Deployment
metadata:
  name: decision-router
  namespace: agentgateway-system
spec:
  selector: {matchLabels: {app: decision-router}}
  template:
    metadata: {labels: {app: decision-router}}
    spec:
      containers:
      - name: router
        image: golang:1.26-alpine
        workingDir: /src
        command: ["sh", "-c"]
        args:
        - cp /cm/main.go . && go mod init router >/dev/null 2>&1 &&
          go get github.com/envoyproxy/go-control-plane/envoy@v1.39.0 google.golang.org/grpc@v1.84.0 &&
          go mod tidy && exec go run .
        env:
        - {name: DECISION_URL, value: http://laya.agentgateway-system.svc.cluster.local:8000}
        - {name: DEFAULT_TIER, value: fast}
        - {name: CONFIDENCE_THRESHOLD, value: "0.55"}
        ports: [{name: grpc, containerPort: 9090}]
        startupProbe: {tcpSocket: {port: 9090}, periodSeconds: 10, failureThreshold: 30}
        readinessProbe: {tcpSocket: {port: 9090}, periodSeconds: 5}
        volumeMounts:
        - {name: src, mountPath: /cm}
        - {name: work, mountPath: /src}
      volumes:
      - {name: src, configMap: {name: decision-router-src}}
      - {name: work, emptyDir: {}}
---
# appProtocol h2c is what tells agentgateway to speak gRPC to this Service.
apiVersion: v1
kind: Service
metadata:
  name: decision-router
  namespace: agentgateway-system
spec:
  selector: {app: decision-router}
  ports: [{name: grpc, port: 9090, targetPort: grpc, appProtocol: kubernetes.io/h2c}]
Enter fullscreen mode Exit fullscreen mode
kubectl -n agentgateway-system create configmap decision-router-src --from-file=main.go
kubectl apply -f 30-router.yaml
Enter fullscreen mode Exit fullscreen mode

Two things to notice. appProtocol: kubernetes.io/h2c on the Service is what tells agentgateway to speak gRPC to it. And the Deployment does go get and go run at start, which suits a lab and not production; the last section says what to do instead.

Step 4: the policy and the route

This is the whole integration. 40-policy.yaml:

apiVersion: agentgateway.dev/v1alpha1
kind: AgentgatewayPolicy
metadata:
  name: decision-router
  namespace: agentgateway-system
spec:
  targetRefs:
  - group: gateway.networking.k8s.io
    kind: Gateway
    name: agentgateway-proxy
  traffic:
    phase: PreRouting
    extProc:
      backendRef:
        name: decision-router
        namespace: agentgateway-system
        port: 9090
      failureMode: FailOpen
      processingOptions:
        requestHeaderMode: Send
        requestBodyMode: Buffered
        responseHeaderMode: Send
        responseBodyMode: None
        requestTrailerMode: Skip
        responseTrailerMode: Skip
Enter fullscreen mode Exit fullscreen mode

Three fields to read twice. phase: PreRouting is what makes the header count for routing; the CRD documents this phase as the one for policies that need to influence the routing decision. requestBodyMode: Buffered gives the router the whole prompt in one message. failureMode: FailOpen covers the router pod itself being unreachable: the request continues to the default route instead of failing. The decision API being unreachable is the router's own job, and it handles it by falling back, so a classifier outage never becomes a gateway outage. responseBodyMode: None is deliberate: the router only needs response headers to echo its decision, so the answer body never takes the detour through gRPC.

Then one HTTPRoute that matches on the header the router wrote. The last rule is the safety net: no header, or a model name nobody knows, and the request goes to the default tier. 40-route.yaml:

apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: llm-tiers
  namespace: agentgateway-system
spec:
  parentRefs:
  - name: agentgateway-proxy
    namespace: agentgateway-system
    sectionName: http
  rules:
  - matches:
    - path: {type: PathPrefix, value: /v1/chat/completions}
      headers:
      - {type: Exact, name: x-selected-model, value: fast-model}
    backendRefs:
    - {name: llm-fast, group: agentgateway.dev, kind: AgentgatewayBackend}
  - matches:
    - path: {type: PathPrefix, value: /v1/chat/completions}
      headers:
      - {type: Exact, name: x-selected-model, value: strong-model}
    backendRefs:
    - {name: llm-strong, group: agentgateway.dev, kind: AgentgatewayBackend}
  - matches:
    - path: {type: PathPrefix, value: /}
    backendRefs:
    - {name: llm-fast, group: agentgateway.dev, kind: AgentgatewayBackend}
Enter fullscreen mode Exit fullscreen mode
kubectl apply -f 40-policy.yaml -f 40-route.yaml
kubectl -n agentgateway-system rollout status deploy/decision-router deploy/laya --timeout=1500s
kubectl -n agentgateway-system get agentgatewaypolicy decision-router
Enter fullscreen mode Exit fullscreen mode
NAME              ACCEPTED   ATTACHED   AGE
decision-router   True       True       2m
Enter fullscreen mode Exit fullscreen mode

Step 5: run it

Open a port-forward and ask the gateway three things.

kubectl -n agentgateway-system port-forward svc/agentgateway-proxy 8080:80 &
Enter fullscreen mode Exit fullscreen mode

An explicit model bypasses the classifier:

curl -s -D - -o /dev/null localhost:8080/v1/chat/completions -H 'content-type: application/json' \
  -d '{"model":"fast-model","messages":[{"role":"user","content":"Say hi."}]}' | grep -i x-
Enter fullscreen mode Exit fullscreen mode
x-decision-reason: passthrough
x-selected-model: fast-model
Enter fullscreen mode Exit fullscreen mode

auto with a simple prompt:

curl -s -D - localhost:8080/v1/chat/completions -H 'content-type: application/json' \
  -d '{"model":"auto","messages":[{"role":"user","content":"What is the capital of France?"}]}'
Enter fullscreen mode Exit fullscreen mode
x-decision-reason: classified
x-decision-confidence: 0.673
x-selected-model: fast-model
x-decision-tier: fast

{"id":"mock","object":"chat.completion","model":"fast-model","choices":[{"index":0,"message":{"role":"assistant","content":"[fast-model] answered"},"finish_reason":"stop"}],...}
Enter fullscreen mode Exit fullscreen mode

And auto with the rate limiter prompt from the top of the article:

x-decision-tier: strong
x-decision-reason: classified
x-decision-confidence: 0.741
x-selected-model: strong-model

{"id":"mock","object":"chat.completion","model":"strong-model","choices":[{"index":0,"message":{"role":"assistant","content":"[strong-model] answered"},...}
Enter fullscreen mode Exit fullscreen mode

The body is the point. The client sent auto and the model server received strong-model. That rewrite is the backend's model field at work.

One negative check, because it is the one I care about most. A client that tries to pick its own tier with a header gets the classifier's answer anyway, because the router strips its headers before deciding:

curl -s -D - -o /dev/null localhost:8080/v1/chat/completions -H 'content-type: application/json' \
  -H 'x-selected-model: strong-model' \
  -d '{"model":"auto","messages":[{"role":"user","content":"What is the capital of France?"}]}' | grep -i x-decision-tier
Enter fullscreen mode Exit fullscreen mode
x-decision-tier: fast
Enter fullscreen mode Exit fullscreen mode

Now sixteen labelled prompts through the classifier, with the tier it chose and the probability behind it:

while IFS='|' read -r expect prompt; do
  out=$(curl -s -D - -o /dev/null localhost:8080/v1/chat/completions -H 'content-type: application/json' \
    -d "$(jq -cn --arg p "$prompt" '{model:"auto",messages:[{role:"user",content:$p}]}')")
  got=$(grep -i '^x-decision-tier:' <<<"$out" | cut -d' ' -f2 | tr -d '\r')
  conf=$(grep -i '^x-decision-confidence:' <<<"$out" | cut -d' ' -f2 | tr -d '\r')
  [[ "$got" == "$expect" ]] && mark=' ' || mark='!'
  printf '%s %-7s %-7s %-6s %s\n' "$mark" "$expect" "$got" "$conf" "${prompt:0:60}"
done <<'PROMPTS'
fast|What is the capital of Portugal?
fast|Translate 'good morning' into German.
fast|Give me a one-line definition of a Kubernetes Service.
fast|Rewrite this sentence in a friendlier tone: Your ticket has been closed.
fast|How many minutes are there in three hours?
fast|Summarize in one sentence: Kind runs Kubernetes nodes as Docker containers.
fast|What does the acronym CRD stand for in Kubernetes?
fast|Say hello in Spanish.
strong|Design a multi-region rate limiter for an API that must stay consistent under partition, and explain the trade-offs between token bucket and sliding window, then write it in Go with unit tests.
strong|Here is a Go stack trace with a nil pointer dereference in a gRPC stream handler. Walk through the likely root causes step by step and propose a fix with a regression test.
strong|Compare the obligations of a provider versus a deployer of a high-risk AI system under the EU AI Act, with article references, and draft a compliance checklist for a bank.
strong|Write a 1500-word architecture document for migrating a monolith to event-driven microservices, covering data ownership, saga patterns, idempotency and observability.
strong|Prove that the sum of the first n odd numbers equals n squared, then generalize the argument to arithmetic progressions.
strong|Refactor this 400-line Python module into testable units, explain each SOLID violation you found, and produce the refactored code with type hints.
strong|Analyze the security implications of running an ext_proc filter before route selection, enumerate threat scenarios, and recommend mitigations with configuration examples.
strong|Given quarterly revenue, churn and CAC figures for three SaaS products, build a cohort model, identify which product to sunset and justify the decision quantitatively.
PROMPTS
Enter fullscreen mode Exit fullscreen mode

This is the run behind this article, on a MacBook with Laya on four CPU threads, each line 200 to 245 ms end to end through the gateway:

  fast    fast    0.678  What is the capital of Portugal?
  fast    fast    0.542  Translate 'good morning' into German.
  fast    fast    0.808  Give me a one-line definition of a Kubernetes Service.
  fast    fast    0.726  Rewrite this sentence in a friendlier tone: Your ticket has
  fast    fast    0.824  How many minutes are there in three hours?
  fast    fast    0.757  Summarize in one sentence: Kind runs Kubernetes nodes as Doc
  fast    fast    0.729  What does the acronym CRD stand for in Kubernetes?
  fast    fast    0.616  Say hello in Spanish.
  strong  strong  0.742  Design a multi-region rate limiter for an API that must stay
  strong  strong  0.714  Here is a Go stack trace with a nil pointer dereference in a
! strong  fast    0.534  Compare the obligations of a provider versus a deployer of a
  strong  strong  0.626  Write a 1500-word architecture document for migrating a mono
  strong  strong  0.604  Prove that the sum of the first n odd numbers equals n squar
  strong  strong  0.557  Refactor this 400-line Python module into testable units, ex
  strong  strong  0.646  Analyze the security implications of running an ext_proc fil
  strong  strong  0.590  Given quarterly revenue, churn and CAC figures for three SaaS
Enter fullscreen mode Exit fullscreen mode

Fifteen of sixteen. Read the marked line with the threshold in mind: Laya leaned strong at 0.534, under the 0.55 line, so the router sent it to the cheap tier and labelled the decision low-confidence. That is the threshold doing its job. When the classifier is not sure, the request goes to the tier you chose as the default, and the header tells you it happened. Lower the threshold and it moves; raise it and more do. The number is yours.

Every decision ends on a tier: classified, low confidence or fallback, all visible in headers

Switching to Jev

Same router, same policy, same route. Set DECISION_URL to https://api.typesafe.ai and give the pod DECISION_API_KEY from a Secret:

kubectl -n agentgateway-system create secret generic typesafe --from-literal=KEY="$TYPESAFE_API_KEY"
kubectl -n agentgateway-system set env deploy/decision-router \
  DECISION_URL=https://api.typesafe.ai --from=secret/typesafe --prefix=DECISION_API_
Enter fullscreen mode Exit fullscreen mode

That last command maps the Secret's KEY to DECISION_API_KEY. The request body from the first section is the one the router sends; model: jev-latest resolves to the current version. Jev's probabilities are the same field the router already reads, and Jev adds a calibrated confidence you could compare instead.

What changes is not the code. It is where the last user message of every auto request goes: to TypeSafe's service, hosted in the United States, under their data processing addendum, with zero data retention as an enterprise option. If that fits your data, the hosted model gives you TypeSafe's calibration and no pod to feed. If it does not, Laya is already running.

Before this goes near a customer

The lab is deliberately the smallest thing that shows the mechanism. What I would add, roughly in order:

Build an image for the router instead of go run in the pod, and move the tier table out of the source into a ConfigMap, so the platform team changes descriptions without touching Go.

Give Laya a PersistentVolumeClaim so a rescheduled pod is not a re-install and a re-download, and pin the version you tested. The emptyDir is fine on a laptop and nowhere else.

Tell a timeout from an error from an unknown label in the headers, and add the decision latency and the model id the API reported. The lab folds them into fallback because the tier is the same; a dashboard wants the difference.

Write your own eval set. Sixteen prompts show the mechanism; your traffic decides the descriptions and the threshold. The loop above is the whole tool.

Start in observe mode. The router already explains every decision in headers, so run it for a week with the HTTPRoute pointing everything at one tier, read x-decision-tier in your logs, and only then split the route.

Ask a second question in the same call. Questions to a decision model are evaluated in parallel and isolated from each other, so a noul asking "does this message contain personal data?" costs nothing extra and gives you a second header to route or block on.

Put the tiers behind failover. AgentgatewayBackend supports priority groups, so strong can be two providers in order, and the router does not need to know.

Next level

Three directions I would take it. More than two tiers: one more entry in the table, one more rule in the HTTPRoute, one more backend. A score question for urgency, routed to a header your rate limit policy can read. And the same router in front of vLLM on a GPU node, where the tiers are two real models and the stand-in answers become real ones.

If you build this and land on a different threshold, a better pair of descriptions, or a reason to put the classifier somewhere else, I want to hear it. That is usually the conversation where I learn something.


agentgateway is an open source project under the Agentic AI Foundation: docs at agentgateway.dev, source at github.com/agentgateway/agentgateway. Jev is documented at docs.typesafe.ai. Laya is at github.com/NandhaKishorM/laya.

Top comments (0)