DEV Community

Mika
Mika

Posted on

Why We Stopped Using LLMs to Route LLMs: The 90ms TypeSafe Jev Architecture

If you are orchestrating multi-step AI agents, your biggest bottleneck is not model intelligence. It is routing latency.

Every time your agent evaluates which tool to invoke or whether a task needs a heavyweight reasoning model vs a lightweight model, calling an autoregressive LLM adds 2 to 4 seconds of pure overhead.

In this guide, we will replace the slow LLM prompt-router with TypeSafe Jev (a dedicated System One decision model), cutting decision latency down to 95ms with zero output token fees.


The Core Concept: Decision Tensors vs Autoregressive Generation

Autoregressive models (like GPT-4o-mini or Claude Haiku) generate text token by token. For classification tasks, this means:

  • Massive prompt ingest (KV-cache compute)
  • 20-50 tokens of JSON syntax generation
  • Regex/JSON schema validation overhead
  • Inevitable schema syntax parsing failures

TypeSafe Jev takes a state string and structured questions, returning calibrated probability distributions directly via a single forward pass.


5-Minute Setup Guide

1. Requirements

  • Python 3.9+
  • OpenRouter API Key

2. Implementation (jev_router.py)

import urllib.request
import json
import os

OPENROUTER_API_KEY = os.environ.get("OPENROUTER_API_KEY")

def route_agent_action(prompt: str) -> dict:
    url = "https://openrouter.ai/api/v1/systemone"

    payload = {
        "model": "~typesafe/jev-latest",
        "state": f"Agent incoming task: {prompt}",
        "questions": {
            "is_code_task": {
                "type": "noul",
                "instructions": "Does this require code modification or terminal execution?"
            },
            "model_tier": {
                "type": "choice",
                "instructions": "Pick the most cost-effective model tier",
                "criteria": {
                    "light": "Simple Q&A, summarizing text, formatting markdown",
                    "medium": "Standard coding, multi-file searches, data processing",
                    "heavy": "Complex architecture, deep reasoning, critical refactoring"
                }
            }
        }
    }

    req = urllib.request.Request(
        url,
        data=json.dumps(payload).encode("utf-8"),
        headers={
            "Authorization": f"Bearer {OPENROUTER_API_KEY}",
            "Content-Type": "application/json",
            "HTTP-Referer": "https://dev.to/m1kulya",
            "X-Title": "Fast Agent Router"
        },
        method="POST"
    )

    with urllib.request.urlopen(req, timeout=5) as resp:
        res = json.loads(resp.read().decode("utf-8"))

    return {
        "needs_code": res["answers"]["is_code_task"]["noul"] > 0.6,
        "tier": res["answers"]["model_tier"]["choice"],
        "confidence": res["answers"]["model_tier"]["confidence"]
    }

if __name__ == "__main__":
    action = route_agent_action("Fix the memory leak in the C++ WebSocket pool")
    print(f"Target Tier: {action['tier']} | Needs Code: {action['needs_code']}")
Enter fullscreen mode Exit fullscreen mode

Production Benchmarks

  • p50 Latency: 95ms
  • p95 Latency: 210ms
  • Cost: \$0.04 per 1,000 decisions (output tokens are 100% free)
  • Failure Rate: 0.0% schema errors

Conclusion & Architecture Takeaway

Decouple your System 1 (deterministic routing, fast dispatch) from your System 2 (deep generation, reasoning). Your users get sub-second agent responsiveness, and your API bill drops by 80%+.

What are you currently using for intent classification in your agent stacks? Let me know in the comments.

Top comments (0)