<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: G POORNA PRUDHVI</title>
    <description>The latest articles on DEV Community by G POORNA PRUDHVI (@poornagurram).</description>
    <link>https://dev.to/poornagurram</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F49880%2Fbeb3d681-8136-47b4-acea-c80f6dd7f3e6.jpeg</url>
      <title>DEV Community: G POORNA PRUDHVI</title>
      <link>https://dev.to/poornagurram</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/poornagurram"/>
    <language>en</language>
    <item>
      <title>Vector Databases for RAG: How to Choose One and Wire It In</title>
      <dc:creator>G POORNA PRUDHVI</dc:creator>
      <pubDate>Fri, 09 Oct 2026 17:38:35 +0000</pubDate>
      <link>https://dev.to/poornagurram/vector-databases-for-rag-how-to-choose-one-and-wire-it-in-2g3n</link>
      <guid>https://dev.to/poornagurram/vector-databases-for-rag-how-to-choose-one-and-wire-it-in-2g3n</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://aiengineerinsights.com/blog/vector-database-for-rag/" rel="noopener noreferrer"&gt;aiengineerinsights.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;In short:&lt;/strong&gt; Does your RAG app need a dedicated vector database? pgvector vs Chroma vs Qdrant vs Weaviate vs Milvus vs Pinecone — open-source vs managed, and how to choose.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fya8tia0r407ivb467br4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fya8tia0r407ivb467br4.png" alt="Vector Databases for RAG: How to Choose One and Wire It In (2026)" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What a vector database does in a RAG pipeline
&lt;/h2&gt;

&lt;p&gt;In a &lt;a href="https://aiengineerinsights.com/blog/rag-vs-fine-tuning" rel="noopener noreferrer"&gt;retrieval-augmented generation pipeline&lt;/a&gt;, the vector database is the store that sits between your &lt;a href="https://aiengineerinsights.com/blog/rag-chunking-strategies" rel="noopener noreferrer"&gt;chunking step&lt;/a&gt; and the language model. It does two jobs.&lt;/p&gt;

&lt;p&gt;At &lt;strong&gt;ingest time&lt;/strong&gt;, each chunk is run through an embedding model to produce a vector, and that vector is written to the database together with an id, the original text, and whatever metadata you attach (source file, section, tenant, date).&lt;/p&gt;

&lt;p&gt;At &lt;strong&gt;query time&lt;/strong&gt;, the user's question is embedded with the same model, and the database returns the top-k stored vectors closest to it under a distance metric such as cosine, inner product, or L2. Those k chunks are the "context" that gets pasted into the prompt.&lt;/p&gt;

&lt;p&gt;The part that makes a vector database more than a table with a float array column is the &lt;strong&gt;approximate nearest-neighbour (ANN) index&lt;/strong&gt;. Exact search compares the query against every stored vector, which is fine for thousands of chunks and painful for millions.&lt;/p&gt;

&lt;p&gt;ANN indexes like HNSW (a graph-based index) and IVFFlat (a clustering-based index) trade a little recall for a large speed-up, and every option in this guide is built around one or more of them (&lt;a href="https://github.com/pgvector/pgvector" rel="noopener noreferrer"&gt;pgvector README — indexing&lt;/a&gt;; &lt;a href="https://qdrant.tech/documentation/concepts/indexing/" rel="noopener noreferrer"&gt;Qdrant docs — indexing&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;Two consequences follow. First, the database can only return what you indexed, so chunking quality caps retrieval quality before the database is ever involved. Second, the database's filtering and index behaviour directly shape what "top-k" means in practice — which is why the choice matters more than "it stores vectors" suggests.&lt;/p&gt;

&lt;h2&gt;
  
  
  Do you actually need a dedicated vector database?
&lt;/h2&gt;

&lt;p&gt;This is the decision most teams get wrong in one of two directions: either they stand up a new distributed system for a corpus of ten thousand chunks, or they bolt vectors onto an existing database and only discover its limits under production filtering load. The honest answer is that &lt;strong&gt;a Postgres extension is enough for a lot of RAG apps&lt;/strong&gt;, and a dedicated database earns its place under specific, nameable conditions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The case for pgvector.&lt;/strong&gt; pgvector adds a &lt;code&gt;vector&lt;/code&gt; column type to Postgres, performs exact nearest-neighbour search by default (perfect recall), and lets you add an HNSW or IVFFlat index for approximate search when the table grows.&lt;/p&gt;

&lt;p&gt;It supports L2, inner product, and cosine distance, and — because it is just Postgres — you keep ACID transactions, joins against your existing tables, point-in-time recovery, and every client library you already use (&lt;a href="https://github.com/pgvector/pgvector" rel="noopener noreferrer"&gt;pgvector README&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;If your chunks live next to the rows they came from, a single &lt;code&gt;SELECT ... ORDER BY embedding &amp;lt;=&amp;gt; $1 LIMIT 10&lt;/code&gt; with a &lt;code&gt;WHERE&lt;/code&gt; clause on tenant or document id is a complete retrieval layer with no second system to operate.&lt;/p&gt;

&lt;p&gt;A dedicated vector database earns its keep when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Scale.&lt;/strong&gt; Your vector count is heading past what a single Postgres node's memory and HNSW build time handle comfortably, and you need horizontal sharding or a distributed deployment designed around vectors (Milvus's distributed mode, for example — &lt;a href="https://milvus.io/docs" rel="noopener noreferrer"&gt;Milvus docs&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Filtering at scale.&lt;/strong&gt; Most real queries combine "nearest to this vector" with "and only from this tenant / date range / document type". Dedicated databases design their ANN index around payload filtering so the filter doesn't collapse recall; Qdrant's filterable HNSW is the canonical example (&lt;a href="https://qdrant.tech/documentation/concepts/filtering/" rel="noopener noreferrer"&gt;Qdrant docs — filtering&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hybrid search.&lt;/strong&gt; You want keyword (BM25) and vector scores fused in one query, which Weaviate ships natively (&lt;a href="https://weaviate.io/developers/weaviate/search/hybrid" rel="noopener noreferrer"&gt;Weaviate docs — hybrid search&lt;/a&gt;). In Postgres you can combine &lt;code&gt;tsvector&lt;/code&gt; full-text search with pgvector, but you are assembling the fusion yourself.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Index tuning.&lt;/strong&gt; You need to pick between HNSW, IVF, product quantisation, disk-based, or GPU indexes per collection and tune them independently of the rest of your database workload.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zero ops.&lt;/strong&gt; Nobody on the team wants to run or tune a database at all, and a managed service is worth the vendor dependency.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If none of those apply yet, start with pgvector (or Chroma for a local prototype) and keep the retrieval interface behind a thin function, so swapping the store later is a contained change rather than a rewrite.&lt;/p&gt;

&lt;h2&gt;
  
  
  The main options compared
&lt;/h2&gt;

&lt;p&gt;Below are the six options that come up in almost every "best vector database for RAG" conversation. The table sticks to documented characteristics — licence, hosting model, and what each is designed for — rather than benchmark rankings, which vary wildly with dataset, dimensionality, filter selectivity, and hardware.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;pgvector&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Type:&lt;/strong&gt; Postgres extension (open source, PostgreSQL licence)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hosting:&lt;/strong&gt; Wherever your Postgres runs — self-hosted or any managed Postgres that ships the extension&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Best for:&lt;/strong&gt; Teams already on Postgres who want vectors next to their relational data, with SQL joins and ACID transactions, without a second system&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Chroma&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Type:&lt;/strong&gt; Open-source vector database (Apache 2.0), embedded / local-first&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hosting:&lt;/strong&gt; In-process with a persistent client, self-hosted server, or Chroma Cloud&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Best for:&lt;/strong&gt; Prototyping and small-to-mid apps where a pip install and a local folder is the whole deployment story&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Qdrant&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Type:&lt;/strong&gt; Open-source vector database (Apache 2.0), written in Rust&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hosting:&lt;/strong&gt; Self-host (Docker / Kubernetes) or Qdrant Cloud&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Best for:&lt;/strong&gt; Production workloads that lean hard on metadata filtering alongside vector search, with an option to move to managed later&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Weaviate&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Type:&lt;/strong&gt; Open-source vector database (BSD-3-Clause) with a managed cloud&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hosting:&lt;/strong&gt; Self-host or Weaviate Cloud&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Best for:&lt;/strong&gt; Teams that want built-in hybrid (BM25 + vector) search and pluggable vectoriser modules out of the box&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Milvus&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Type:&lt;/strong&gt; Open-source vector database (Apache 2.0), LF AI &amp;amp; Data project&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hosting:&lt;/strong&gt; Milvus Lite (pip), Standalone (Docker), Distributed (Kubernetes), or Zilliz Cloud&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Best for:&lt;/strong&gt; Large-scale deployments that need a distributed architecture and a wide choice of index types, including disk- and GPU-based ones&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Pinecone&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Type:&lt;/strong&gt; Fully managed, closed-source vector database&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hosting:&lt;/strong&gt; Managed service only — no self-host option&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Best for:&lt;/strong&gt; Teams that want zero infrastructure to run and are comfortable with a vendor-hosted, proprietary store&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A few details worth knowing beyond the table. &lt;strong&gt;Chroma&lt;/strong&gt; can run entirely in-process (ephemeral or persisted to a directory) or as a client-server deployment, and its query API supports metadata filters via &lt;code&gt;where&lt;/code&gt; and document-text filters via &lt;code&gt;where_document&lt;/code&gt; (&lt;a href="https://docs.trychroma.com/" rel="noopener noreferrer"&gt;Chroma docs&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Qdrant&lt;/strong&gt; stores a JSON "payload" alongside each vector and supports sparse vectors and quantisation to shrink memory (&lt;a href="https://qdrant.tech/documentation/" rel="noopener noreferrer"&gt;Qdrant docs&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Weaviate&lt;/strong&gt; offers HNSW, flat, and dynamic index types plus vectoriser modules that call an embedding provider for you at import time (&lt;a href="https://weaviate.io/developers/weaviate" rel="noopener noreferrer"&gt;Weaviate docs&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Milvus&lt;/strong&gt; exposes the broadest index menu — FLAT, IVF variants, HNSW, DiskANN, and GPU indexes — and three deployment modes from a pip-installable Lite to a Kubernetes cluster (&lt;a href="https://milvus.io/docs" rel="noopener noreferrer"&gt;Milvus docs&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pinecone&lt;/strong&gt; is serverless and managed only, with namespaces for tenant isolation and metadata filtering on query (&lt;a href="https://docs.pinecone.io/" rel="noopener noreferrer"&gt;Pinecone docs&lt;/a&gt;).&lt;/p&gt;

&lt;h2&gt;
  
  
  Open source vs managed: the real trade-off
&lt;/h2&gt;

&lt;p&gt;Five of the six options above are open-source vector databases or extensions, and four of those (Chroma, Qdrant, Weaviate, Milvus via Zilliz) also sell a managed cloud tier. That makes the choice less "open source vs managed" than "who runs it, and what do you give up either way".&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Self-hosting an open source vector database&lt;/strong&gt; gives you control over index parameters, data locality (the vectors never leave your VPC), and predictable infrastructure cost at steady load.&lt;/p&gt;

&lt;p&gt;The price is operations: you own upgrades, backups, replication, memory sizing for HNSW graphs, and re-indexing when you change embedding models.&lt;/p&gt;

&lt;p&gt;For a team that already runs Postgres or Kubernetes this is often marginal work; for a two-person team shipping a product it can be the largest single chunk of non-feature effort.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A managed service&lt;/strong&gt; — Pinecone, or the cloud tier of an open-source option — removes that operational load entirely and usually scales without a re-architecture.&lt;/p&gt;

&lt;p&gt;What you give up is control and portability: index internals are the vendor's, cost scales with usage rather than with hardware, and migrating out means re-embedding or bulk-exporting your corpus.&lt;/p&gt;

&lt;p&gt;Choosing a managed tier of an &lt;em&gt;open-source&lt;/em&gt; database keeps an exit hatch open, since the same API is available self-hosted; a closed-source service does not.&lt;/p&gt;

&lt;p&gt;A reasonable rule: prototype on something you can run locally, and make the self-host vs managed call when you know your actual query volume, filter patterns, and who will be on-call.&lt;/p&gt;

&lt;h2&gt;
  
  
  A minimal vector database example for RAG (Chroma)
&lt;/h2&gt;

&lt;p&gt;Here is the entire chunk → embed → store → query loop using the Chroma vector database, because it needs nothing but &lt;code&gt;pip install chromadb&lt;/code&gt; and runs in-process. The same four calls — create a client, get a collection, add documents, query — map almost one-to-one onto every other option in this guide (&lt;a href="https://docs.trychroma.com/" rel="noopener noreferrer"&gt;Chroma docs — getting started&lt;/a&gt;).&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;chromadb&lt;/span&gt;

&lt;span class="c1"&gt;# 1. Persistent client: vectors are stored on disk at ./rag_db and survive restarts
&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;chromadb&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;PersistentClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;./rag_db&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# 2. A collection is a named index. Chroma embeds documents with its
#    default embedding function unless you pass embedding_function=...
&lt;/span&gt;&lt;span class="n"&gt;collection&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_or_create_collection&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;docs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# 3. Chunks from your chunking step (one string per chunk)
&lt;/span&gt;&lt;span class="n"&gt;chunks&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pgvector adds a vector type plus HNSW and IVFFlat indexes to Postgres.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Chroma can run in-process with a persistent client, no server needed.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Qdrant is an open-source vector database written in Rust.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="c1"&gt;# 4. Add: each chunk is embedded, indexed, and stored with an id and metadata
&lt;/span&gt;&lt;span class="n"&gt;collection&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;ids&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;chunk-&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;))],&lt;/span&gt;
    &lt;span class="n"&gt;documents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;metadatas&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;source&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;notes.md&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;chunk&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;))],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# 5. Query: the question is embedded the same way, then the top-k nearest
#    chunks come back. 'where' applies an optional metadata filter.
&lt;/span&gt;&lt;span class="n"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;collection&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;query_texts&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Which option runs inside Postgres?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;n_results&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;where&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;source&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;notes.md&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;dist&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;zip&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;documents&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;distances&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]):&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;dist&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;  &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three things to notice. &lt;strong&gt;You never called an embedding model directly&lt;/strong&gt; — Chroma applies a default embedding function to &lt;code&gt;documents&lt;/code&gt; on both &lt;code&gt;add&lt;/code&gt; and &lt;code&gt;query&lt;/code&gt;, and you can swap in OpenAI, Sentence Transformers, or your own function by passing &lt;code&gt;embedding_function=&lt;/code&gt; when you create the collection.&lt;/p&gt;

&lt;p&gt;Whatever you choose, ingest and query must use the same model or the distances are meaningless. &lt;strong&gt;Metadata is not optional in practice&lt;/strong&gt; — the &lt;code&gt;where&lt;/code&gt; filter is how you scope retrieval to a tenant, a document, or a date range, and it is the feature you will lean on most as the corpus grows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The distances come back with the documents&lt;/strong&gt;, which is what you will log when you start measuring retrieval quality in the next section.&lt;/p&gt;

&lt;p&gt;Swapping this for pgvector means a &lt;code&gt;CREATE EXTENSION vector&lt;/code&gt;, a table with a &lt;code&gt;vector(n)&lt;/code&gt; column, an &lt;code&gt;INSERT&lt;/code&gt; per chunk, and an &lt;code&gt;ORDER BY embedding &amp;lt;=&amp;gt; query_vector LIMIT k&lt;/code&gt; query — with the embedding call done in your own code (&lt;a href="https://github.com/pgvector/pgvector" rel="noopener noreferrer"&gt;pgvector README&lt;/a&gt;). For Qdrant, Weaviate, Milvus, and Pinecone the shape is the same: a client, a collection/index, an upsert of vectors plus payload, and a search with a filter.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to choose: a decision guide by stage
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Prototype / local development.&lt;/strong&gt; Chroma with a persistent client, or pgvector if you already have Postgres running. Optimise for iteration speed — you will change chunking and embedding models several times, and re-indexing a local store is free.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Already on Postgres, up to millions of vectors.&lt;/strong&gt; pgvector with an HNSW index. You get joins against your application data, one backup story, and no new system. Revisit when index build time, memory, or filtered-query recall becomes a measurable problem rather than a hypothetical one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Production with heavy metadata filtering or hybrid search.&lt;/strong&gt; Qdrant (filterable HNSW, payloads, sparse vectors), Weaviate (native hybrid search, vectoriser modules), or Milvus (distributed deployment, wide index choice). All three are open source, so you can self-host now and move to their managed cloud later without changing the client code.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zero-ops, willing to accept a vendor.&lt;/strong&gt; Pinecone, or the managed tier of one of the open-source databases. Pick the managed tier of an open-source option if you want the ability to self-host later; pick Pinecone if you never intend to.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Very large scale (hundreds of millions of vectors and up).&lt;/strong&gt; Milvus Distributed or a managed service designed for that scale. This is the only tier where a dedicated distributed architecture is a requirement rather than a preference.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Whatever you pick, put the retrieval call behind a single function that takes a query string and filter and returns chunks with scores. Every option here fits that interface, and it is the difference between a one-afternoon migration and a week of untangling.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to know your vector DB retrieval is actually good
&lt;/h2&gt;

&lt;p&gt;The database returns top-k vectors by distance — that is a mechanical guarantee, not a quality one. Whether those k chunks are the &lt;em&gt;right&lt;/em&gt; chunks depends on the embedding model, the chunking, the index's recall at your settings, and how well filters interact with the index.&lt;/p&gt;

&lt;p&gt;None of that is visible from the query response, so treat the store the way you'd treat any other pipeline component: hold everything else constant and measure retrieval.&lt;/p&gt;

&lt;p&gt;The two retrieval-side metrics to track are &lt;strong&gt;context precision&lt;/strong&gt; (of the chunks returned, how many were relevant) and &lt;strong&gt;context recall&lt;/strong&gt; (of the chunks that were relevant, how many were returned), computed against a labelled eval set of questions and their expected source chunks. Frameworks such as RAGAS implement both (&lt;a href="https://docs.ragas.io/" rel="noopener noreferrer"&gt;RAGAS docs&lt;/a&gt;), and our &lt;a href="https://aiengineerinsights.com/blog/rag-evaluation-metrics" rel="noopener noreferrer"&gt;guide to RAG evaluation metrics&lt;/a&gt; walks through how they're calculated and what thresholds teams commonly use.&lt;/p&gt;

&lt;p&gt;Two database-specific checks are worth adding.&lt;/p&gt;

&lt;p&gt;First, compare ANN recall against exact search on a sample: pgvector lets you drop the index (or raise &lt;code&gt;hnsw.ef_search&lt;/code&gt;) and re-run the same query to see how much recall the index is costing you (&lt;a href="https://github.com/pgvector/pgvector" rel="noopener noreferrer"&gt;pgvector README — query options&lt;/a&gt;); Qdrant and the others expose equivalent search-time parameters.&lt;/p&gt;

&lt;p&gt;Second, re-run your eval set with your most selective production filter applied — a store that scores well unfiltered and badly filtered is telling you something important about how its index handles that filter. A vector database is only as good as the retrieval quality you can measure through it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is a vector database in RAG?
&lt;/h3&gt;

&lt;p&gt;A vector database is the store that holds the embeddings of your document chunks and answers nearest-neighbour queries against them. In a RAG pipeline it sits between chunking and generation: at ingest time each chunk is embedded and written to the database with an id and metadata, and at query time the user's question is embedded with the same model and the database returns the top-k most similar chunks, which become the context the language model answers from.&lt;/p&gt;

&lt;h3&gt;
  
  
  Do I need a vector database for RAG?
&lt;/h3&gt;

&lt;p&gt;Not always. For small-to-mid scale — up to a few million vectors — pgvector inside Postgres is often enough, especially if you already run Postgres and want vectors next to your relational data.&lt;/p&gt;

&lt;p&gt;A dedicated vector database earns its place when you need to scale beyond what a single Postgres node handles comfortably, when you rely heavily on metadata filtering combined with vector search, or when you want hybrid (keyword plus vector) search and finer control over the ANN index.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is the best open-source vector database?
&lt;/h3&gt;

&lt;p&gt;There is no single winner — it depends on where you are. Chroma is the easiest to start with because it runs in-process from a pip install and persists to a local folder, which makes it ideal for local development and prototyping.&lt;/p&gt;

&lt;p&gt;Qdrant, Weaviate, and Milvus are all open-source options built for self-hosted production: Qdrant is known for filtering alongside vector search, Weaviate for built-in hybrid search and vectoriser modules, and Milvus for a distributed architecture and a wide choice of index types.&lt;/p&gt;

&lt;p&gt;Pick based on your deployment model and the features you will actually use, not on a leaderboard.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is pgvector good enough for RAG?
&lt;/h3&gt;

&lt;p&gt;Yes, for many production applications. pgvector adds a vector column type, exact nearest-neighbour search by default, and HNSW and IVFFlat indexes for approximate search, and it inherits Postgres transactions, joins, and backups.&lt;/p&gt;

&lt;p&gt;It is a strong choice up to millions of vectors, particularly if you are already on Postgres.&lt;/p&gt;

&lt;p&gt;The caveats are that ANN index build time and memory (especially for HNSW) grow with the table, and that scaling beyond one node is a Postgres scaling problem rather than a vector-database feature, so very large or very filter-heavy workloads may be better served by a dedicated store.&lt;/p&gt;

&lt;h3&gt;
  
  
  Chroma vs Pinecone: which should I use for RAG?
&lt;/h3&gt;

&lt;p&gt;They solve different problems. Chroma is open source and local-first: it runs in-process with a persistent client, so it is great for development, prototyping, and small self-hosted apps, and it also offers a server mode and a managed cloud.&lt;/p&gt;

&lt;p&gt;Pinecone is a fully managed, closed-source service: you never run infrastructure, it scales without you operating anything, and in return you accept a proprietary store and a vendor dependency.&lt;/p&gt;

&lt;p&gt;Use Chroma to build and iterate; move to Pinecone (or a managed tier of an open-source database) when you want someone else to run production.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do I know if my vector database is returning good results?
&lt;/h3&gt;

&lt;p&gt;Measure retrieval directly rather than judging the final answers by eye.&lt;/p&gt;

&lt;p&gt;Build a small eval set of questions with the chunks that should be retrieved, then compute context precision (how much of what was retrieved is relevant) and context recall (how much of what was relevant was retrieved) for each configuration — index type, top-k, filters, embedding model.&lt;/p&gt;

&lt;p&gt;Our guide to RAG evaluation metrics covers how those scores are computed. The database choice only matters to the extent that retrieval quality you can measure improves.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/pgvector/pgvector" rel="noopener noreferrer"&gt;pgvector — Open-source vector similarity search for Postgres (GitHub README)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.trychroma.com/" rel="noopener noreferrer"&gt;Chroma — Documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://qdrant.tech/documentation/" rel="noopener noreferrer"&gt;Qdrant — Documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://qdrant.tech/documentation/concepts/indexing/" rel="noopener noreferrer"&gt;Qdrant — Indexing concepts&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://qdrant.tech/documentation/concepts/filtering/" rel="noopener noreferrer"&gt;Qdrant — Filtering concepts&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://weaviate.io/developers/weaviate" rel="noopener noreferrer"&gt;Weaviate — Documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://weaviate.io/developers/weaviate/search/hybrid" rel="noopener noreferrer"&gt;Weaviate — Hybrid search&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://milvus.io/docs" rel="noopener noreferrer"&gt;Milvus — Documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.pinecone.io/" rel="noopener noreferrer"&gt;Pinecone — Documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.ragas.io/" rel="noopener noreferrer"&gt;RAGAS — Documentation (context precision and context recall)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>rag</category>
      <category>database</category>
      <category>llm</category>
    </item>
    <item>
      <title>MCP vs API: What the Model Context Protocol Actually Is</title>
      <dc:creator>G POORNA PRUDHVI</dc:creator>
      <pubDate>Fri, 09 Oct 2026 17:37:57 +0000</pubDate>
      <link>https://dev.to/poornagurram/mcp-vs-api-what-the-model-context-protocol-actually-is-ggl</link>
      <guid>https://dev.to/poornagurram/mcp-vs-api-what-the-model-context-protocol-actually-is-ggl</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://aiengineerinsights.com/blog/mcp-vs-api/" rel="noopener noreferrer"&gt;aiengineerinsights.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;In short:&lt;/strong&gt; MCP vs API: an open, model-facing standard that makes tools reusable across AI apps — turning M×N integrations into M+N. How it works and when to use each.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fattn20vkgzws8jtdzxty.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fattn20vkgzws8jtdzxty.png" alt="MCP vs API: What the Model Context Protocol Actually Is (and When to Use It)" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What is MCP (Model Context Protocol)?
&lt;/h2&gt;

&lt;p&gt;MCP is an &lt;strong&gt;open protocol for connecting AI applications to external tools and data sources.&lt;/strong&gt; It was introduced and open-sourced by &lt;a href="https://www.anthropic.com/news/model-context-protocol" rel="noopener noreferrer"&gt;Anthropic in late 2024&lt;/a&gt; and adopted across the industry through 2025.&lt;/p&gt;

&lt;p&gt;The problem it solves is boring but expensive: before MCP, every AI app needed custom integration code for every tool it wanted to use. Ten apps and ten tools meant a hundred bespoke integrations.&lt;/p&gt;

&lt;p&gt;MCP fixes that the way USB-C fixed chargers, or the way the Language Server Protocol fixed editor tooling: a single standard connector.&lt;/p&gt;

&lt;p&gt;You expose a capability once as an &lt;strong&gt;MCP server&lt;/strong&gt;, and any &lt;strong&gt;MCP host&lt;/strong&gt; — Claude, ChatGPT, an IDE like Cursor, or your own &lt;a href="https://aiengineerinsights.com/blog/what-are-ai-agents" rel="noopener noreferrer"&gt;AI agent&lt;/a&gt; — can discover and use it without custom glue.&lt;/p&gt;

&lt;p&gt;Anthropic's own phrase for it is a "USB-C port for AI applications." If you want the full plain-English explainer — what an MCP server is, the host/client/server roles, and the architecture — see &lt;a href="https://aiengineerinsights.com/blog/what-is-mcp" rel="noopener noreferrer"&gt;What Is MCP?&lt;/a&gt; This article focuses on how MCP compares to a plain API.&lt;/p&gt;

&lt;h2&gt;
  
  
  MCP vs API: the real difference
&lt;/h2&gt;

&lt;p&gt;The question "MCP vs API" is slightly the wrong frame, because MCP is built &lt;em&gt;on top of&lt;/em&gt; APIs — an MCP server almost always calls a REST or database API underneath. The useful contrast is what each is &lt;em&gt;for&lt;/em&gt;: an API is a general interface a developer codes against; MCP is a model-facing protocol that makes those interfaces discoverable and reusable by any AI app.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Traditional API&lt;/th&gt;
&lt;th&gt;MCP&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;What it is&lt;/td&gt;
&lt;td&gt;A general interface between two pieces of software&lt;/td&gt;
&lt;td&gt;A specific open protocol for connecting LLM apps to tools and data&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Who it's for&lt;/td&gt;
&lt;td&gt;Developers, who write integration code&lt;/td&gt;
&lt;td&gt;The model / agent, which calls tools at runtime&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Discovery&lt;/td&gt;
&lt;td&gt;Read the docs, hard-code each endpoint&lt;/td&gt;
&lt;td&gt;Self-describing — the host lists available tools automatically&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;What it exposes&lt;/td&gt;
&lt;td&gt;Endpoints (REST, GraphQL, RPC…)&lt;/td&gt;
&lt;td&gt;Standard Tools, Resources, and Prompts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Integration cost&lt;/td&gt;
&lt;td&gt;M×N — every app wired to every tool&lt;/td&gt;
&lt;td&gt;M+N — write one server, any host uses it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Transport&lt;/td&gt;
&lt;td&gt;Usually HTTP&lt;/td&gt;
&lt;td&gt;JSON-RPC 2.0 over stdio or streamable HTTP&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Relationship&lt;/td&gt;
&lt;td&gt;The thing being called&lt;/td&gt;
&lt;td&gt;Usually wraps an API so a model can use it&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The single biggest win is the last-but-one row. Custom integrations scale as M×N — every AI app times every tool. MCP collapses that to &lt;strong&gt;M+N&lt;/strong&gt;: write a server once, and every host that speaks MCP gets it for free.&lt;/p&gt;

&lt;h2&gt;
  
  
  How MCP works: architecture
&lt;/h2&gt;

&lt;p&gt;MCP is a client-server protocol built on &lt;a href="https://www.jsonrpc.org/specification" rel="noopener noreferrer"&gt;JSON-RPC 2.0&lt;/a&gt;. Three roles (shown in the diagram above):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Host.&lt;/strong&gt; The AI application the user interacts with — Claude Desktop, an IDE, an agent runtime.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Client.&lt;/strong&gt; Lives inside the host and holds one connection to each server.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Server.&lt;/strong&gt; Exposes capabilities over MCP, usually wrapping an API, database, or filesystem.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each server offers three standard primitives: &lt;strong&gt;Tools&lt;/strong&gt; (actions the model can call, like "create a GitHub issue"), &lt;strong&gt;Resources&lt;/strong&gt; (data/context the app can read, like a file or record), and &lt;strong&gt;Prompts&lt;/strong&gt; (reusable templates a user can invoke).&lt;/p&gt;

&lt;p&gt;Servers run over &lt;strong&gt;stdio&lt;/strong&gt; for local processes or &lt;strong&gt;streamable HTTP&lt;/strong&gt; for remote ones. Because the primitives are standard and self-describing, the host can list a server's tools at runtime — no hard-coding.&lt;/p&gt;

&lt;h2&gt;
  
  
  How MCP evolved — and what was hard at first
&lt;/h2&gt;

&lt;p&gt;MCP shipped fast and changed fast, so the version you read about in an early post is not quite the one you build on today. The first release (November 2024) supported two transports — &lt;strong&gt;stdio&lt;/strong&gt; for local servers and an &lt;strong&gt;HTTP+SSE&lt;/strong&gt; transport for remote ones — with a deliberately minimal spec. That minimalism is what got it adopted, but it also left real gaps the next revisions had to close.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The early pain points:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The HTTP+SSE transport was awkward.&lt;/strong&gt; It needed a long-lived server-sent-events stream plus a separate POST endpoint — stateful, hard to run on serverless, and with no way to resume a dropped connection.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Auth was underspecified.&lt;/strong&gt; Early remote servers had no standard way to authenticate, so everyone rolled their own.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security gaps.&lt;/strong&gt; Because tool descriptions are fed to the model, a malicious server can smuggle instructions ("tool poisoning" / prompt injection), and over-permissioned servers or token pass-through opened confused-deputy and token-theft risks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Breaking changes.&lt;/strong&gt; The spec moved so quickly that early adopters had to chase transport and SDK churn.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The revisions that fixed most of it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;2025-03-26.&lt;/strong&gt; The clunky HTTP+SSE transport was replaced by &lt;strong&gt;Streamable HTTP&lt;/strong&gt; — a single endpoint that works statelessly (serverless-friendly) and supports resumable streams — and a formal &lt;strong&gt;OAuth 2.1&lt;/strong&gt; authorization framework was added.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2025-06-18.&lt;/strong&gt; MCP servers were defined as OAuth &lt;strong&gt;Resource Servers&lt;/strong&gt; with Resource Indicators (RFC 8707) so tokens can't be reused where they shouldn't be, plus structured tool output, an &lt;em&gt;elicitation&lt;/em&gt; flow for servers to ask the user for input, removal of JSON-RPC batching, and a dedicated security best-practices document.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Governance opened up too, moving toward a public specification process rather than a single-vendor project. Net effect: the "USB-C for AI" idea survived, but the wiring behind it got a lot more production-ready.&lt;/p&gt;

&lt;h2&gt;
  
  
  RAG vs MCP: not the same thing
&lt;/h2&gt;

&lt;p&gt;These get compared a lot, but they solve different problems at different layers. &lt;strong&gt;RAG (retrieval-augmented generation) is a technique&lt;/strong&gt; for pulling relevant knowledge into a model's context before it answers. &lt;strong&gt;MCP is a protocol&lt;/strong&gt; for connecting tools and data sources to AI apps. They're complementary, not competing.&lt;/p&gt;

&lt;p&gt;In fact, the cleanest way to ship RAG to an agent is &lt;em&gt;through&lt;/em&gt; MCP: wrap your retrieval pipeline (embeddings, vector search, reranking) in an MCP server that exposes a &lt;code&gt;search_knowledge_base&lt;/code&gt; tool. Now any MCP host can do retrieval against your data with zero custom integration. RAG is the &lt;em&gt;what&lt;/em&gt;; MCP is one clean &lt;em&gt;how&lt;/em&gt; to deliver it.&lt;/p&gt;

&lt;h2&gt;
  
  
  When to use MCP vs a plain API
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Reach for MCP&lt;/strong&gt; when an LLM or agent needs to discover and call tools, especially if you want the same tool to work across multiple AI hosts, or you're building agents that plug into many capabilities.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reach for a plain API&lt;/strong&gt; when it's deterministic app-to-app integration with no model in the loop, when you need maximum control or performance, or when only one consumer will ever call it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Usually you use both:&lt;/strong&gt; the MCP server is a thin, model-friendly wrapper around your existing API, adding discovery, standard primitives, and reuse.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What developers actually say about MCP
&lt;/h2&gt;

&lt;p&gt;MCP's reception has been loud on both sides — worth knowing before you build on it. &lt;strong&gt;The praise:&lt;/strong&gt; it solved a real, expensive problem (the M×N integration mess), and adoption was extraordinary — thousands of community servers within months, plus cross-vendor support from OpenAI, Google, and Microsoft, made it a de facto standard faster than almost any recent protocol.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The criticism&lt;/strong&gt; clusters in three places:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Security is the loudest.&lt;/strong&gt; Researchers repeatedly warned that MCP makes it easy to wire an agent into what Simon Willison calls the "lethal trifecta" — access to private data, exposure to untrusted content, and a way to exfiltrate — so a poisoned tool description or malicious server can hijack an agent. The 2025 auth and security revisions were a direct response.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context and token burn.&lt;/strong&gt; Every connected server injects its tool definitions — names, descriptions, and schemas — into the model's context on &lt;em&gt;every&lt;/em&gt; call. Wire up a dozen servers and you can spend thousands of tokens before the user asks anything, driving up cost and latency; worse, too many tools measurably degrades the model's tool selection. The practical fix is to load only the servers a task needs (and lean on dynamic/lazy tool loading as hosts add it).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Transport churn.&lt;/strong&gt; Replacing HTTP+SSE so soon frustrated early adopters who'd already built on it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Do we even need it?"&lt;/strong&gt; Some argued function calling plus OpenAPI already covered most cases and that MCP re-invents the wheel; the counter is that runtime discovery, standard primitives, and cross-host reuse are exactly what ad-hoc function calling lacks.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The honest read in 2026: MCP won the standardization battle, and the sharp early edges (transport, auth, security) are mostly sanded down — but &lt;strong&gt;you still own the security of what you connect.&lt;/strong&gt; Treat every third-party server as untrusted, scope permissions tightly, and keep a human in the loop for risky actions.&lt;/p&gt;

&lt;h2&gt;
  
  
  MCP examples and ecosystem
&lt;/h2&gt;

&lt;p&gt;There's a large and growing library of ready-made MCP servers — filesystem, GitHub, Google Drive, Slack, Postgres, Puppeteer/browser, and web search among them — plus SDKs to build your own in TypeScript, Python, and other languages.&lt;/p&gt;

&lt;p&gt;On the host side, Claude Desktop, popular IDEs and coding agents (see our &lt;a href="https://aiengineerinsights.com/blog/best-ai-coding-agents" rel="noopener noreferrer"&gt;ranked AI coding agents&lt;/a&gt;), and custom agent frameworks all speak MCP. That two-sided adoption is exactly why writing one server is worth it: it lights up everywhere at once.&lt;/p&gt;

&lt;h2&gt;
  
  
  MCP vs A2A, ADK, and other agent protocols
&lt;/h2&gt;

&lt;p&gt;MCP isn't the only standard in this space, and the others mostly complement it rather than compete. The trick is to notice which &lt;em&gt;connection&lt;/em&gt; each one is about:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;MCP&lt;/strong&gt; — agent ↔ &lt;em&gt;tools and data&lt;/em&gt;. What this article is about.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A2A (Agent2Agent)&lt;/strong&gt; — agent ↔ &lt;em&gt;agent&lt;/em&gt;. Google's protocol for letting independent agents discover and delegate to each other, later donated to the Linux Foundation. MCP gives an agent its tools; A2A lets agents talk to one another. See our &lt;a href="https://aiengineerinsights.com/blog/google-a2a" rel="noopener noreferrer"&gt;deep dive on A2A&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ADK (Agent Development Kit)&lt;/strong&gt; — not a protocol but Google's open-source &lt;em&gt;framework&lt;/em&gt; for building agents; it speaks both MCP (for tools) and A2A (for agent-to-agent), which is a clean picture of how the layers stack.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ACP and others&lt;/strong&gt; — additional agent-communication efforts exist (e.g. IBM/BeeAI's ACP), but momentum in 2026 has consolidated around &lt;strong&gt;MCP for tools and A2A for agent-to-agent&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How to build an MCP server
&lt;/h2&gt;

&lt;p&gt;The minimum path:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Pick an SDK&lt;/strong&gt; — the official TypeScript or Python SDK is the usual starting point.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Define your primitives.&lt;/strong&gt; List the Tools, Resources, and Prompts you want to expose, with clear names and descriptions (the model reads these).
3.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Wrap the backend.&lt;/strong&gt; Each tool calls your existing API, database, or filesystem.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Choose a transport.&lt;/strong&gt; stdio for local, streamable HTTP for remote.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Connect a host&lt;/strong&gt; — point Claude Desktop, your IDE, or your agent at the server, and the tools appear automatically.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Building tool-using agents is core AI-engineering work in 2026. If you're leveling up toward it, our &lt;a href="https://aiengineerinsights.com/ai-engineering-roadmap" rel="noopener noreferrer"&gt;AI engineering roadmap&lt;/a&gt; and &lt;a href="https://aiengineerinsights.com/blog/ai-engineer-skills" rel="noopener noreferrer"&gt;skills checklist&lt;/a&gt; cover the fundamentals underneath MCP and agents.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is MCP in simple terms?
&lt;/h3&gt;

&lt;p&gt;MCP (Model Context Protocol) is an open standard for connecting AI apps to external tools and data. Instead of writing custom glue code for every tool in every AI app, you expose a tool once as an MCP server, and any MCP-compatible host — Claude, ChatGPT, an IDE, your own agent — can use it. Anthropic describes it as a 'USB-C port for AI applications.'&lt;/p&gt;

&lt;h3&gt;
  
  
  What is the difference between MCP and an API?
&lt;/h3&gt;

&lt;p&gt;An API is a general interface you write code against, one integration at a time. MCP is one standardized, model-facing protocol that makes tools self-describing and reusable across every AI app. They aren't competitors: an MCP server usually calls an API under the hood — MCP is the layer that lets an LLM discover and use that API in a uniform way.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is MCP better than a REST API?
&lt;/h3&gt;

&lt;p&gt;It's not better or worse — it operates at a different layer. Use a plain API for deterministic app-to-app integration with no model in the loop. Use MCP when you want an LLM or agent to discover and call tools, especially across multiple AI hosts. In practice MCP servers wrap REST APIs, so you often use both together.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is the difference between RAG and MCP?
&lt;/h3&gt;

&lt;p&gt;RAG (retrieval-augmented generation) is a technique for pulling relevant knowledge into a model's context. MCP is a protocol for connecting tools and data sources. They're at different layers and are often combined — you can expose a retrieval/RAG capability as an MCP server so any agent can search your knowledge base as a standard tool.&lt;/p&gt;

&lt;h3&gt;
  
  
  Who created MCP and is it open?
&lt;/h3&gt;

&lt;p&gt;MCP was introduced and open-sourced by Anthropic in late 2024. The specification is public, with SDKs in several languages, and it saw broad adoption across the industry through 2025 — including support from other model providers and many IDEs and agent frameworks.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is MCP secure?
&lt;/h3&gt;

&lt;p&gt;MCP is as secure as how you deploy it. Early versions had gaps — underspecified auth and prompt-injection / 'tool poisoning' risks — which the 2025 revisions addressed with an OAuth 2.1 framework, resource-server semantics, and a security best-practices spec.&lt;/p&gt;

&lt;p&gt;But you still own the risk of what you connect: treat third-party servers as untrusted, scope permissions tightly, and keep a human in the loop for sensitive actions.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is the difference between MCP and A2A?
&lt;/h3&gt;

&lt;p&gt;They cover different connections. MCP connects an agent to tools and data. A2A (Agent2Agent, from Google) connects agents to each other so they can discover and delegate work. They're complementary — an agent might use MCP for its tools and A2A to hand off to another agent. Google's ADK framework speaks both.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do I build an MCP server?
&lt;/h3&gt;

&lt;p&gt;Pick an official MCP SDK (TypeScript or Python are common), define the Tools, Resources, and Prompts you want to expose, wrap whatever API or data source they call, and run the server over stdio (for local) or HTTP (for remote). Point an MCP host — Claude Desktop, an IDE, or your agent — at it, and the tools become available automatically.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.anthropic.com/news/model-context-protocol" rel="noopener noreferrer"&gt;Anthropic — Introducing the Model Context Protocol&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://modelcontextprotocol.io/" rel="noopener noreferrer"&gt;Model Context Protocol — official documentation and specification&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://modelcontextprotocol.io/specification/versioning" rel="noopener noreferrer"&gt;MCP specification — versioning and revision history (2025-03-26, 2025-06-18)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/modelcontextprotocol/servers" rel="noopener noreferrer"&gt;MCP reference servers (GitHub)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://a2a-protocol.org/" rel="noopener noreferrer"&gt;A2A (Agent2Agent) protocol&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://google.github.io/adk-docs/" rel="noopener noreferrer"&gt;Google Agent Development Kit (ADK) documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://simonwillison.net/tags/lethal-trifecta/" rel="noopener noreferrer"&gt;Simon Willison — the "lethal trifecta" (agent security)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.jsonrpc.org/specification" rel="noopener noreferrer"&gt;JSON-RPC 2.0 specification&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>mcp</category>
      <category>api</category>
      <category>llm</category>
    </item>
    <item>
      <title>LLM Routing Explained: How an LLM Router Picks the Right Model</title>
      <dc:creator>G POORNA PRUDHVI</dc:creator>
      <pubDate>Fri, 09 Oct 2026 17:36:11 +0000</pubDate>
      <link>https://dev.to/poornagurram/llm-routing-explained-how-an-llm-router-picks-the-right-model-2mmp</link>
      <guid>https://dev.to/poornagurram/llm-routing-explained-how-an-llm-router-picks-the-right-model-2mmp</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://aiengineerinsights.com/blog/llm-routing/" rel="noopener noreferrer"&gt;aiengineerinsights.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; LLM routing sends each request to the cheapest model that can answer it acceptably, using a decision made before the expensive call — by rules, embedding similarity, a trained router (RouteLLM-style), or a calibrated decision model (Jev-style). Published routers report cost cuts from roughly 2× (RouteLLM) up to 98% (FrugalGPT cascades) on their own benchmarks; the number you get depends on how much of your traffic is genuinely easy. Default anything uncertain to the stronger model, log every decision, and measure router accuracy and quality delta on your own logs.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F94lgjdur958nx73vehsq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F94lgjdur958nx73vehsq.png" alt="LLM Routing Explained: How an LLM Router Picks the Right Model (RouteLLM, Semantic Router, Jev, AI Gateways) — 2026" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What is LLM routing, and what does an LLM router do?
&lt;/h2&gt;

&lt;p&gt;LLM routing (also called model routing) is the practice of choosing &lt;em&gt;which&lt;/em&gt; model answers a request at runtime rather than hard-wiring your application to one model. The router is whatever makes that choice.&lt;/p&gt;

&lt;p&gt;It sits between your application and your model providers, inspects each incoming request — the text, its length, the user, the tools in play — and dispatches it to a model tier: a small, fast model for the easy cases; a frontier model for the hard ones; sometimes a middle tier in between.&lt;/p&gt;

&lt;p&gt;The reason routing exists is an uncomfortable fact about production traffic: a large share of it is easy. Password resets, "what's your refund policy," classify-this-ticket, rewrite-this-sentence.&lt;/p&gt;

&lt;p&gt;Sending those to the same model you use for multi-step agent planning is paying frontier prices for work a model a tenth the size does fine.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://arxiv.org/abs/2305.05176" rel="noopener noreferrer"&gt;FrugalGPT paper&lt;/a&gt; framed the economics bluntly in 2023: API fees across popular models "can differ by two orders of magnitude," which is exactly the gap a router exploits.&lt;/p&gt;

&lt;p&gt;One distinction to fix before anything else, because the term "LLM router" is used for two different things:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Model routing&lt;/strong&gt; decides &lt;em&gt;which model&lt;/em&gt; should answer, based on the request. This is what RouteLLM, Semantic Router, NVIDIA's router, OpenRouter's auto router, and Jev do, each with a different signal. It is the subject of this post.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gateway routing&lt;/strong&gt; decides &lt;em&gt;which deployment&lt;/em&gt; of a model serves a request — which region, which provider key, which copy has rate-limit headroom — plus retries, fallbacks, and cooldowns. LiteLLM's Router is the canonical open-source example. It does not look at what the request is asking.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You usually want both, layered: a model router in front choosing the tier, a gateway behind it keeping each tier reliable. Conflating them is how teams buy a gateway, call it a router, and wonder why their bill didn't move.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does LLM routing actually save money? What the papers measured
&lt;/h2&gt;

&lt;p&gt;Yes, with a caveat that matters: every published number is on the authors' own benchmark and model pair. Three results anchor the field:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;FrugalGPT&lt;/strong&gt; (&lt;a href="https://arxiv.org/abs/2305.05176" rel="noopener noreferrer"&gt;Chen, Zaharia, Zou, 2023&lt;/a&gt;) reports that it "can match the performance of the best individual LLM (e.g. GPT-4) with up to 98% cost reduction" and can "improve the accuracy over GPT-4 by 4% with the same cost," using three strategies — prompt adaptation, LLM approximation, and the &lt;strong&gt;LLM cascade&lt;/strong&gt;: try a cheap model first, score its answer, and only escalate if the score is low.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hybrid LLM&lt;/strong&gt; (&lt;a href="https://arxiv.org/abs/2404.14618" rel="noopener noreferrer"&gt;Ding et al., 2024&lt;/a&gt;) trains a router on predicted query difficulty with a tunable quality level, and reports "up to 40% fewer calls to the large model, with no drop in response quality."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;RouteLLM&lt;/strong&gt; (&lt;a href="https://arxiv.org/abs/2406.18665" rel="noopener noreferrer"&gt;Ong et al., 2024&lt;/a&gt;) trains routers on human preference data to choose between a stronger and a weaker LLM, and reports it "significantly reduces costs — by over 2 times in certain cases — without compromising the quality of responses." Notably, the routers kept their performance "even when the strong and weak models are changed at test time."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The spread — 2× to 98% — is not the papers disagreeing. It is the savings being a function of three things you control: how much of your traffic is genuinely easy, how much cheaper your small tier is than your big one, and how accurately your router can tell the two apart.&lt;/p&gt;

&lt;p&gt;A router with perfect accuracy on traffic that is 90% easy saves close to 90% of frontier spend. The same router on agent-planning traffic that is 90% hard saves almost nothing and adds a hop.&lt;/p&gt;

&lt;p&gt;The routing research keeps pushing the accuracy term — a 2026 NVIDIA paper, &lt;a href="https://arxiv.org/abs/2603.20895" rel="noopener noreferrer"&gt;"LLM Router: Rethinking Routing with Prefill Activations"&lt;/a&gt;, routes on a model's internal prefill activations instead of surface features and reports closing 45.58% of the gap to an oracle router at 74.31% cost savings relative to the most expensive model — but the traffic-mix term is yours, and it is the one to measure first.&lt;/p&gt;

&lt;h2&gt;
  
  
  The four routing signals: rules, embeddings, learned routers, decision models
&lt;/h2&gt;

&lt;p&gt;Every router, open source or hosted, makes its decision from one of four kinds of signal — or a stack of them. The signal determines what the router can and cannot see, which is more important than which product wraps it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rules and metadata&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;What it decides on:&lt;/strong&gt; Token count, user tier, whether a tool is required, language, SLA — plain code&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost of the decision:&lt;/strong&gt; Free, microseconds, deterministic&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Strength:&lt;/strong&gt; Zero risk of being 'wrong' in a way you can't explain; perfect for hard constraints&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Weakness:&lt;/strong&gt; Cannot tell an easy question from a hard one by content; rules rot as traffic shifts&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Embedding similarity (semantic router)&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;What it decides on:&lt;/strong&gt; Which of your example-utterance groups the query is closest to in vector space&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost of the decision:&lt;/strong&gt; One embedding call per request (often local), tens of milliseconds&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Strength:&lt;/strong&gt; No training loop — add a route by adding example sentences; good for intent-shaped traffic&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Weakness:&lt;/strong&gt; Measures topic, not difficulty; a 'billing' question can be trivial or brutal&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Learned difficulty / preference router&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;What it decides on:&lt;/strong&gt; A trained model predicts which tier will answer well (or which model 'wins') for this query&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost of the decision:&lt;/strong&gt; A small model forward pass; needs labeled or preference data to train&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Strength:&lt;/strong&gt; Directly optimizes the thing you care about — quality vs cost; strongest published results&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Weakness:&lt;/strong&gt; Training data effort; drift when your models or traffic change; opaque&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Calibrated decision model (Jev-style)&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;What it decides on:&lt;/strong&gt; A typed Choice — simple / complex — plus a confidence trained to track accuracy&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost of the decision:&lt;/strong&gt; One non-autoregressive pass; TypeSafe quotes ~70–500 ms and $0.042 per million input tokens&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Strength:&lt;/strong&gt; No training data, labels defined at request time, and a confidence you can threshold on&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Weakness:&lt;/strong&gt; Calibration is imperfect in the 0.3–0.8 band; prompt injection can move the verdict&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Rules&lt;/strong&gt; are where to start and where to put anything that is actually a policy: requests over N tokens go to the long-context model, requests that need a tool go to the model that calls tools reliably, enterprise-tier users always get the frontier model. Rules can't judge difficulty from content, but they never surprise you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Embedding similarity&lt;/strong&gt; is what &lt;a href="https://github.com/aurelio-labs/semantic-router" rel="noopener noreferrer"&gt;Semantic Router&lt;/a&gt; does: you define each route with a handful of example utterances, the library embeds them, and at request time it embeds the query and returns the closest route (or &lt;code&gt;None&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;It is fast and needs no training loop, and it is the right tool when your traffic is intent-shaped — billing vs support vs sales. Its blind spot is that it measures &lt;em&gt;topic&lt;/em&gt;, not &lt;em&gt;difficulty&lt;/em&gt;. "Why was I charged twice" and "reconcile these three invoices against the contract amendment" are both billing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Learned routers&lt;/strong&gt; fix that by training a model to predict the thing you care about: RouteLLM predicts, from preference data, whether the weak model's answer would be preferred; Hybrid LLM predicts query difficulty against a quality target; NVIDIA's prefill-activation router predicts per-model correctness from internal activations. They have the strongest published results and the highest setup cost — you need labeled or preference data, and the router is a model that drifts when your traffic or your model pair changes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Calibrated decision models&lt;/strong&gt; are the newest option and the one most relevant to teams who want a learned-style signal without a training project.&lt;/p&gt;

&lt;p&gt;A model like &lt;a href="https://aiengineerinsights.com/blog/what-is-jev" rel="noopener noreferrer"&gt;Jev&lt;/a&gt; takes the request plus a label set you define at call time (simple / complex, or small / medium / frontier) and returns one typed Choice with a confidence trained to track accuracy.&lt;/p&gt;

&lt;p&gt;The confidence is the point: it lets you route on a threshold instead of a guess. We come back to the practical rules for that in the Jev section below.&lt;/p&gt;

&lt;p&gt;A fifth pattern deserves a name because it is a different &lt;em&gt;topology&lt;/em&gt;, not a different signal: the &lt;strong&gt;cascade&lt;/strong&gt;. Instead of predicting difficulty up front, FrugalGPT's cascade calls the cheap model first, scores the answer with a separate checker, and escalates only if the score is low.&lt;/p&gt;

&lt;p&gt;Cascades don't need a difficulty predictor, but they pay the cheap call on every request and add its latency to the hard ones. Routers and cascades compose well: route the obvious cases, cascade the ambiguous ones.&lt;/p&gt;

&lt;h2&gt;
  
  
  RouteLLM vs Semantic Router vs NVIDIA vs Jev vs LiteLLM vs OpenRouter: what each actually routes on
&lt;/h2&gt;

&lt;p&gt;The names below get thrown into the same "LLM router" bucket in search results and Reddit threads. They are not interchangeable. Read the third column: it tells you which signal from the previous section each one is built on, and therefore what it can't see.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;RouteLLM (LMSYS / UC Berkeley)&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;What it is:&lt;/strong&gt; Learned router, open source&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Routes on:&lt;/strong&gt; Predicted win-rate of a strong vs weak model, trained on human preference data&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Worth knowing:&lt;/strong&gt; Paper reports cost reductions of over 2× in certain cases without quality loss; routers transferred when the strong/weak pair was swapped&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Semantic Router (Aurelio Labs)&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;What it is:&lt;/strong&gt; Embedding router, open source&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Routes on:&lt;/strong&gt; Similarity between the query embedding and example utterances per route&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Worth knowing:&lt;/strong&gt; Returns a route name or None; positioned as a 'superfast decision-making layer' that avoids waiting on an LLM generation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;NVIDIA LLM Router blueprint&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;What it is:&lt;/strong&gt; Reference architecture&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Routes on:&lt;/strong&gt; v2: intent routing with a small Qwen 1.7B model, or 'auto-routing' with CLIP embeddings plus a trained network&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Worth knowing:&lt;/strong&gt; v2 only returns a model recommendation (it does not proxy the call); the repo now carries a deprecation notice pointing to NeMo Switchyard&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Jev (TypeSafe AI)&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;What it is:&lt;/strong&gt; Calibrated decision model, API&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Routes on:&lt;/strong&gt; A Choice you define at request time (e.g. simple / complex) with a calibrated confidence&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Worth knowing:&lt;/strong&gt; Used in LangChain's ModelRouterMiddleware with the instruction 'Choose the least costly model that can complete the task'&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;LiteLLM Router&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;What it is:&lt;/strong&gt; AI gateway / proxy, open source&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Routes on:&lt;/strong&gt; Deployment-level load balancing: simple-shuffle, latency-based, usage-based, least-busy, cost-based, custom&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Worth knowing:&lt;/strong&gt; Balances copies of the same model across providers and regions with cooldowns, fallbacks, retries; does not judge query difficulty&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;OpenRouter Auto Router (openrouter/auto)&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;What it is:&lt;/strong&gt; Hosted gateway feature&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Routes on:&lt;/strong&gt; A lightweight classifier assigns ~30 task types, then ranks models by what the OpenRouter community spent on that task over a trailing 7-day window&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Worth knowing:&lt;/strong&gt; Cost tiers (low → max), session stickiness, allowed/excluded model filters, graceful degradation to a default set&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Two of these deserve a closer look because they are the ones most often mistaken for a difficulty router. &lt;a href="https://docs.litellm.ai/docs/routing" rel="noopener noreferrer"&gt;LiteLLM's Router&lt;/a&gt; is a gateway: its documented strategies — simple-shuffle (the default, weighted by RPM/TPM), latency-based, usage-based, least-busy, cost-based, and custom — all pick a &lt;em&gt;deployment&lt;/em&gt; from a pool, and its reliability features are cooldowns, fallbacks, timeouts, and retries.&lt;/p&gt;

&lt;p&gt;It is excellent at that job. It will not notice that a request is easy.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://openrouter.ai/docs/features/model-routing" rel="noopener noreferrer"&gt;OpenRouter's Auto Router&lt;/a&gt; (&lt;code&gt;openrouter/auto&lt;/code&gt;) is a genuine model router, but note what it optimizes.&lt;/p&gt;

&lt;p&gt;Per its docs, "a fast, lightweight classifier assigns each prompt one of ~30 fine-grained task types," then the router "looks up which models the OpenRouter community actually spends on over a trailing 7-day window" for that task, filtered by the cost tier you choose (&lt;code&gt;low&lt;/code&gt; through &lt;code&gt;max&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;That is routing on &lt;em&gt;task type and market behavior&lt;/em&gt;, not on the difficulty of your specific request — useful if you want a sensible default per task without building anything, less useful if your goal is "send 70% of my support traffic to the small model." It also keeps a conversation on the model it landed on and degrades to a default set if classification is unavailable.&lt;/p&gt;

&lt;p&gt;The NVIDIA blueprint is a reminder that this space moves fast: the &lt;a href="https://github.com/NVIDIA-AI-Blueprints/llm-router" rel="noopener noreferrer"&gt;llm-router repo&lt;/a&gt; went from v1 (task/complexity classifiers) to v2 (intent routing via a small Qwen model, or CLIP embeddings plus a trained network, returning recommendations only) and now carries a deprecation notice pointing at NeMo Switchyard. Treat any specific router product as replaceable; the signal taxonomy is what lasts.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do you build an LLM router? A pattern that holds up in production
&lt;/h2&gt;

&lt;p&gt;The router you ship is almost never one signal. It is rules first, then one content signal, then a threshold with a safe default, then logging. In order:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Define the tiers and what "acceptable" means per tier.&lt;/strong&gt; Two tiers is enough to start (small, frontier). Write down the quality bar — the eval score or human rubric — that a small-tier answer must clear. Without this, "the router works" is unfalsifiable.
2.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Pull every hard constraint into rules.&lt;/strong&gt; Context length, tool requirements, tenant policy, regulated content. These never go through a model; they run first and short-circuit.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Pick one content signal for the rest.&lt;/strong&gt; Intent-shaped traffic: embeddings. Difficulty-shaped traffic with labeled data: a learned router.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Difficulty-shaped traffic with no labels and a deadline: a calibrated decision model. Start with the one you can ship this week; you can swap it later because the interface — request in, tier out — doesn't change.&lt;br&gt;
4.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Threshold with the expensive model as the default.&lt;/strong&gt; Route to the small tier only when the signal is confidently "easy." Everything uncertain goes up. A router that errs toward the frontier model costs you a little money; one that errs toward the small model costs you quality you may not notice for weeks.&lt;br&gt;
5.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Feed the router structured fields, not just raw text.&lt;/strong&gt; Token count, turn count, whether retrieval found anything, user tier. It gives rules something to act on and shrinks the surface an attacker can write into.&lt;br&gt;
6.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Log the decision, the confidence, the tier, and the outcome.&lt;/strong&gt; This log is your eval set, your drift detector, and the training data for a learned router later.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Put a gateway behind each tier.&lt;/strong&gt; Retries, fallbacks, cooldowns, provider keys — the LiteLLM job.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If the small tier is down, the router's decision should fall through to the frontier tier, not fail.&lt;/p&gt;

&lt;p&gt;The shape of it as pseudocode. &lt;strong&gt;Illustrative only&lt;/strong&gt; — the control flow, not real SDK calls:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# ILLUSTRATIVE PSEUDOCODE — control flow only, not real API syntax.
&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;route&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# 1. hard rules run first, in plain code
&lt;/span&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tokens&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;SMALL_CTX&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;needs_tools&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tier&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;enterprise&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;FRONTIER&lt;/span&gt;

    &lt;span class="c1"&gt;# 2. one content signal — swap the implementation, keep the interface
&lt;/span&gt;    &lt;span class="n"&gt;decision&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;router&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;decide&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;simple or complex?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="nb"&gt;input&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nf"&gt;structured&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;          &lt;span class="c1"&gt;# fields your code assembled
&lt;/span&gt;        &lt;span class="n"&gt;labels&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;simple&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;complex&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;label&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;confidence&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="c1"&gt;# 3. threshold with the expensive model as the default
&lt;/span&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;label&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;simple&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;confidence&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.9&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;SMALL&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;FRONTIER&lt;/span&gt;                          &lt;span class="c1"&gt;# the uncertain middle goes UP, not down
&lt;/span&gt;
&lt;span class="n"&gt;tier&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;route&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;answer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;gateway&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;complete&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tier&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;     &lt;span class="c1"&gt;# retries / fallbacks live in the gateway
&lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;tier&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;SMALL&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="nf"&gt;passes_check&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;answer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;gateway&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;complete&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;FRONTIER&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# optional cascade on failure
&lt;/span&gt;&lt;span class="nf"&gt;log_outcome&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tier&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice the asymmetry in step 3. The threshold is on "simple," not on "complex." That is deliberate: it encodes the safe default into the control flow, so a badly calibrated or drifting router degrades into "slightly more expensive," never into "silently worse." If the router is doing its job inside an &lt;a href="https://aiengineerinsights.com/blog/what-are-ai-agents" rel="noopener noreferrer"&gt;agent loop&lt;/a&gt;, the same decision point is also where you'd gate risky tool calls — a different question to the same cheap decision layer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where does Jev fit as an LLM router?
&lt;/h2&gt;

&lt;p&gt;Routing by complexity is the headline use case TypeSafe and LangChain give for &lt;a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev" rel="noopener noreferrer"&gt;Jev&lt;/a&gt;, so it is worth being precise about what it adds and what it doesn't.&lt;/p&gt;

&lt;p&gt;Jev is a non-autoregressive decision model: you send the request plus a label set, it returns one Choice (up to 255 labels) and a confidence in a single forward pass, with no text. It is trained with RLCD — Reinforcement Learning for Calibrated Decisions — so the confidence is meant to track how often it is right.&lt;/p&gt;

&lt;p&gt;Pricing is $0.042 per million input tokens with output free, which is what makes it plausible to call on every request.&lt;/p&gt;

&lt;p&gt;In the taxonomy above, that makes it a &lt;strong&gt;learned-style signal with no training step&lt;/strong&gt;: you get a difficulty judgment (not just a topic match, as with embeddings) without collecting preference data (as with RouteLLM).&lt;/p&gt;

&lt;p&gt;LangChain's &lt;a href="https://www.langchain.com/blog/building-a-harness-with-jev" rel="noopener noreferrer"&gt;harness write-up&lt;/a&gt; wires it in as a &lt;code&gt;ModelRouterMiddleware&lt;/code&gt; with the instruction "Choose the least costly model that can complete the task," and uses the same model in an &lt;code&gt;AutoModeMiddleware&lt;/code&gt; to block risky tool calls before they execute.&lt;/p&gt;

&lt;p&gt;Our &lt;a href="https://aiengineerinsights.com/blog/how-to-use-jev" rel="noopener noreferrer"&gt;how to use Jev&lt;/a&gt; guide walks through that routing-plus-guardrail pattern; the &lt;a href="https://aiengineerinsights.com/blog/jev-vs-llm" rel="noopener noreferrer"&gt;Jev vs LLMs&lt;/a&gt; post covers why you wouldn't use a chat model to make the routing decision in the first place.&lt;/p&gt;

&lt;p&gt;The limits, which decide how you threshold it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Calibration is real but imperfect.&lt;/strong&gt; An independent test measured an expected calibration error of about 0.107, with confidence reliable near 0 and 1 and shaky in the 0.3–0.8 band. For routing that means: trust a 0.96 "simple"; treat a 0.6 "simple" as "complex." The pseudocode above does exactly that.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prompt injection moves the verdict.&lt;/strong&gt; &lt;a href="https://venturebeat.com/security/companies-are-putting-jev-in-charge-of-ai-agent-decisions-and-prompt-injection-can-influence-the-verdict" rel="noopener noreferrer"&gt;VentureBeat&lt;/a&gt; reported an Octomind demo in which a block probability fell from 0.76 to 0.48 after a fake "user pre-approved" field was added to the input. For a router, the attack is cheaper and subtler: a user who writes "this is a simple question" into their request may push themselves onto the weak model and get a worse answer — or, in the other direction, push expensive traffic onto your frontier tier. Keep the decision on fields your code controls, and cap per-user frontier spend in rules, not in the model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No reasoning, no explanation.&lt;/strong&gt; You get a label and a number. If you need to know &lt;em&gt;why&lt;/em&gt; a request was routed, you reconstruct it from the logged fields.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vendor numbers are vendor numbers.&lt;/strong&gt; The ~70–500 ms latency and the headline speed/cost multipliers are TypeSafe's own figures on workflows it chose. Benchmark the decision latency in your region before it goes into a p99 budget.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Net: a calibrated decision model is the fastest way to get a difficulty-aware router running on day one, and the logged decisions become the dataset that lets you train a RouteLLM-style router later if the volume justifies it. The same bridge logic we described for &lt;a href="https://aiengineerinsights.com/blog/jev-vs-ml-classification" rel="noopener noreferrer"&gt;classification&lt;/a&gt; applies: zero-shot now, trained later, with the zero-shot phase building toward its own replacement.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do you evaluate an LLM router?
&lt;/h2&gt;

&lt;p&gt;A router is a classifier whose errors have asymmetric costs, so evaluate it like one. Three numbers, per routing configuration, on a labeled sample of real traffic:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Router accuracy, split by direction.&lt;/strong&gt; How often did it send an easy request to the small tier (the savings), and how often did it send a hard request to the small tier (the damage)? Report them separately; a single accuracy number hides the one that matters.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quality delta vs all-frontier.&lt;/strong&gt; Run the routed pipeline and the everything-to-the-frontier-model pipeline over the same eval set and grade both — with your existing evals, an LLM judge, and human spot checks on the disagreements. The delta is the price you are paying for the savings. Decide the acceptable delta &lt;em&gt;before&lt;/em&gt; you look at the savings.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Realized cost and latency.&lt;/strong&gt; Not the vendor multiplier — your bill and your p50/p99, including the router's own call. A router that adds 400 ms to every request to save money on 30% of them may be a net loss on latency-sensitive paths.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then keep running it. The production log from the build section is the re-evaluation set; re-run the three numbers whenever you change a model, a prompt, or a threshold, and watch the confidence distribution for drift.&lt;/p&gt;

&lt;p&gt;If you already have an eval harness for RAG or agents — the kind we describe in &lt;a href="https://aiengineerinsights.com/blog/rag-evaluation-metrics" rel="noopener noreferrer"&gt;RAG evaluation metrics&lt;/a&gt; — the router eval slots into it as one more configuration to compare.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common LLM routing mistakes
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Using a chat LLM as the router.&lt;/strong&gt; "Rate this query's difficulty from 1 to 10" costs an autoregressive call per request, returns a number that is generated text rather than a calibrated probability, and can take longer than the small-model answer it was supposed to save. Use a cheap signal.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Thresholding toward the cheap model.&lt;/strong&gt; If "complex" needs 0.9 confidence to reach the frontier model, every uncertain request gets the weak answer. Flip it: "simple" needs the confidence; everything else goes up.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Buying a gateway and calling it a router.&lt;/strong&gt; Load balancing across deployments of the same model will not reduce frontier spend. Check which signal the product routes on.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Routing on topic when the problem is difficulty.&lt;/strong&gt; Embedding routers are fast and useful, but an intent label doesn't tell you whether the small model can handle this instance of that intent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Letting user text control the decision.&lt;/strong&gt; Any model-based router can be nudged by what the user writes. Rules on fields you control — tier, spend caps, tool requirements — are the parts an attacker can't talk past.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Shipping without the quality delta.&lt;/strong&gt; Savings are visible on the bill the same week; quality loss shows up as churn a quarter later. Measure both on day one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not logging confidence.&lt;/strong&gt; A log of labels without confidences cannot tell you whether the router is drifting or whether the threshold is in the right place. Store the number.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is LLM routing?
&lt;/h3&gt;

&lt;p&gt;LLM routing is a cheap decision made before an expensive model call: for each incoming request, a router picks which model (or model tier) should answer it, based on difficulty, intent, cost, latency, or policy. The goal is to send the easy majority of traffic to a small, cheap model and reserve the frontier model for the hard minority, so you cut cost and latency without a visible drop in quality.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is the difference between an LLM router and an AI gateway?
&lt;/h3&gt;

&lt;p&gt;An AI gateway (LiteLLM, Portkey, OpenRouter's core proxy) sits in front of many providers and handles keys, rate limits, retries, fallbacks, and load balancing across deployments of a model — it decides which copy of a model serves a request.&lt;/p&gt;

&lt;p&gt;An LLM router decides which model should answer at all, based on the content of the request. Many products do both; the routing signal (rules, embeddings, a learned router, or a calibrated decision model) is the part that determines quality and cost.&lt;/p&gt;

&lt;h3&gt;
  
  
  How much does LLM routing actually save?
&lt;/h3&gt;

&lt;p&gt;Published results, on the authors' own benchmarks: RouteLLM (Ong et al., 2024) reports cost reductions of over 2× in certain cases without compromising response quality; Hybrid LLM (Ding et al., 2024) reports up to 40% fewer calls to the large model with no drop in response quality; FrugalGPT (Chen, Zaharia, Zou, 2023) reports matching the best individual LLM with up to 98% cost reduction using cascades and related tricks. Your number depends on how much of your traffic is genuinely easy, so measure it on your own logs before putting it in a budget.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is RouteLLM?
&lt;/h3&gt;

&lt;p&gt;RouteLLM is an open-source framework and paper from LMSYS and UC Berkeley (Ong et al., 2024) for training routers that dynamically choose between a stronger and a weaker LLM at inference time. The routers are trained on human preference data (plus data augmentation) to predict when the weaker model's answer would be preferred, and the paper reports that trained routers kept working even when the strong and weak models were swapped at test time.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can I use Jev as an LLM router?
&lt;/h3&gt;

&lt;p&gt;Yes — routing by complexity is one of the use cases TypeSafe and LangChain document for it. You ask Jev a Choice question (simple vs complex, or small / medium / frontier) and get back a label plus a calibrated confidence in one non-autoregressive pass, with no training data.&lt;/p&gt;

&lt;p&gt;The practical rules: route to the cheap model only on high-confidence 'simple', send every mid-range confidence (roughly 0.3–0.8) to the stronger model by default, and keep attacker-controlled text out of the fields the decision hinges on, because prompt injection has been shown to move Jev's verdicts.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do I evaluate an LLM router?
&lt;/h3&gt;

&lt;p&gt;Build a labeled eval set of real requests where you already know which tier answers acceptably, then measure three things for each routing configuration: router accuracy (how often it picks the cheapest acceptable tier), the quality delta versus sending everything to the frontier model (graded by your existing evals or an LLM judge with human spot checks), and the realized cost and latency. Log every production decision with its confidence and outcome so you can re-run that evaluation as traffic and models drift.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2406.18665" rel="noopener noreferrer"&gt;Ong et al. 2024 — RouteLLM: Learning to Route LLMs with Preference Data (arXiv)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2305.05176" rel="noopener noreferrer"&gt;Chen, Zaharia, Zou 2023 — FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance (arXiv)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2404.14618" rel="noopener noreferrer"&gt;Ding et al. 2024 — Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing (arXiv)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2603.20895" rel="noopener noreferrer"&gt;Varshney et al. 2026 — LLM Router: Rethinking Routing with Prefill Activations (arXiv)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/aurelio-labs/semantic-router" rel="noopener noreferrer"&gt;Aurelio Labs — semantic-router (GitHub)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/NVIDIA-AI-Blueprints/llm-router" rel="noopener noreferrer"&gt;NVIDIA AI Blueprints — LLM Router (GitHub)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.litellm.ai/docs/routing" rel="noopener noreferrer"&gt;LiteLLM — Router: load balancing, fallbacks, retries (docs)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://openrouter.ai/docs/features/model-routing" rel="noopener noreferrer"&gt;OpenRouter — Model routing and the Auto Router (docs)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.langchain.com/blog/building-a-harness-with-jev" rel="noopener noreferrer"&gt;LangChain — Building a harness with Jev (ModelRouterMiddleware, AutoModeMiddleware)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev" rel="noopener noreferrer"&gt;TypeSafe AI — Introducing System One models and Jev&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://venturebeat.com/security/companies-are-putting-jev-in-charge-of-ai-agent-decisions-and-prompt-injection-can-influence-the-verdict" rel="noopener noreferrer"&gt;VentureBeat — Companies are putting Jev in charge of AI agent decisions, and prompt injection can influence the verdict&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>python</category>
    </item>
  </channel>
</rss>
