<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Emil</title>
    <description>The latest articles on DEV Community by Emil (@igidov).</description>
    <link>https://dev.to/igidov</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4106315%2F48e2d2ea-4d0e-49d9-9891-594b3ce91fd3.png</url>
      <title>DEV Community: Emil</title>
      <link>https://dev.to/igidov</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/igidov"/>
    <language>en</language>
    <item>
      <title>91.2% lower LLM inference cost — and the failed benchmark I almost buried</title>
      <dc:creator>Emil</dc:creator>
      <pubDate>Tue, 15 Sep 2026 08:30:00 +0000</pubDate>
      <link>https://dev.to/igidov/912-lower-llm-inference-cost-and-the-failed-benchmark-i-almost-buried-5cl8</link>
      <guid>https://dev.to/igidov/912-lower-llm-inference-cost-and-the-failed-benchmark-i-almost-buried-5cl8</guid>
      <description>&lt;p&gt;&lt;em&gt;What building an LLM cost router, finding problems in my own benchmarks, and a corrected 100-query experiment taught me about the infrastructure layer between AI applications and model providers.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The expensive model problem
&lt;/h2&gt;

&lt;p&gt;You ship a feature. It calls a large LLM. It works beautifully, so you move on.&lt;/p&gt;

&lt;p&gt;Months later, someone looks at the bill and notices that a surprising amount of traffic is still going through that expensive model.&lt;/p&gt;

&lt;p&gt;“What’s my account number?”&lt;/p&gt;

&lt;p&gt;“Define photosynthesis.”&lt;/p&gt;

&lt;p&gt;“What does NDA stand for?”&lt;/p&gt;

&lt;p&gt;None of these questions needed the same model as a difficult reasoning task.&lt;/p&gt;

&lt;p&gt;But without routing, they often get one anyway.&lt;/p&gt;

&lt;p&gt;The model isn't necessarily the problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The lack of a routing layer is.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I built CARDIAC-PURR to address that problem.&lt;/p&gt;

&lt;p&gt;I'm an attorney by training, not an engineer, so I made an early decision: I wasn't going to trust my intuition, and I wasn't going to trust somebody else's benchmark either.&lt;/p&gt;

&lt;p&gt;I was going to measure it myself.&lt;/p&gt;

&lt;p&gt;That turned out to be much harder than I expected.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I built
&lt;/h2&gt;

&lt;p&gt;CARDIAC-PURR sits between an application and its LLM providers.&lt;/p&gt;

&lt;p&gt;The basic idea is simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Application
    ↓
CARDIAC-PURR
    ↓
Model tier
    ↓
LLM provider
    ↓
Response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Instead of sending every request to the most capable model, the router tries to determine which tier is appropriate for that request.&lt;/p&gt;

&lt;p&gt;The design spans multiple providers and multiple model tiers.&lt;/p&gt;

&lt;p&gt;The important economic idea is not “always use the cheapest model.”&lt;/p&gt;

&lt;p&gt;That would just trade compute cost for answer quality.&lt;/p&gt;

&lt;p&gt;The objective is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use a lower-cost model when the request can be handled there, and use a more capable tier when it cannot.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;There is also a safety path for cases where the initially selected tier produces an inadequate or truncated response.&lt;/p&gt;

&lt;p&gt;Some of the implementation is proprietary, so I'm not going to turn this article into a reverse-engineering document.&lt;/p&gt;

&lt;p&gt;What I can publish is how the system behaved, what the benchmarks showed, and where the benchmarks themselves failed.&lt;/p&gt;

&lt;h2&gt;
  
  
  I wanted a benchmark I could challenge
&lt;/h2&gt;

&lt;p&gt;The first version of the benchmark covered multiple providers and model tiers.&lt;/p&gt;

&lt;p&gt;That was important because pricing is different across providers, and a routing strategy that looks impressive with one provider may look very different with another.&lt;/p&gt;

&lt;p&gt;But benchmarking a router isn't just a matter of collecting dollar figures.&lt;/p&gt;

&lt;p&gt;I wanted to answer two different questions:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Did the router choose the intended model tier?&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;What did those routing decisions cost?&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Those are related, but they are not the same measurement.&lt;/p&gt;

&lt;p&gt;A router can be very good at choosing tiers and still have poor economics.&lt;/p&gt;

&lt;p&gt;It can also save money while making bad routing decisions.&lt;/p&gt;

&lt;p&gt;That distinction became increasingly important as the investigation progressed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then the benchmark started teaching me things I didn't expect
&lt;/h2&gt;

&lt;p&gt;During earlier testing, I saw unexpectedly large execution differences when the same underlying requests passed through different application frameworks.&lt;/p&gt;

&lt;p&gt;One example involved an agent framework.&lt;/p&gt;

&lt;p&gt;A simple question could arrive at the control plane wrapped in task instructions, output requirements, and other framework-generated context.&lt;/p&gt;

&lt;p&gt;The user's task had not become harder.&lt;/p&gt;

&lt;p&gt;The input seen by the router had.&lt;/p&gt;

&lt;p&gt;That exposed a routing edge case.&lt;/p&gt;

&lt;p&gt;We fixed it.&lt;/p&gt;

&lt;p&gt;But fixing the router wasn't enough.&lt;/p&gt;

&lt;p&gt;The investigation also led us to audit the benchmark itself.&lt;/p&gt;

&lt;p&gt;And that's where things got uncomfortable.&lt;/p&gt;

&lt;h2&gt;
  
  
  A benchmark can be wrong and still look right
&lt;/h2&gt;

&lt;p&gt;Our measurement harness had correct timestamps, but one part of the timing accounting was incomplete.&lt;/p&gt;

&lt;p&gt;An interval between provider dispatch stages was not being attributed correctly to a named latency component.&lt;/p&gt;

&lt;p&gt;In one example, the missing interval represented almost all of the observed time.&lt;/p&gt;

&lt;p&gt;The timestamps were real.&lt;/p&gt;

&lt;p&gt;The total request time was real.&lt;/p&gt;

&lt;p&gt;The &lt;em&gt;explanation of that time&lt;/em&gt; was wrong.&lt;/p&gt;

&lt;p&gt;A benchmark doesn't become trustworthy merely because the numbers look plausible.&lt;/p&gt;

&lt;p&gt;We corrected the accounting and added a validation layer for subsequent experiments.&lt;/p&gt;

&lt;p&gt;The principle was simple:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The benchmark itself needs acceptance criteria.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The first 100-query economic benchmark failed
&lt;/h2&gt;

&lt;p&gt;Once the routing and measurement issues had been addressed, I ran a paired benchmark using &lt;strong&gt;100 randomly selected technical queries&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The same 100 queries went through both arms, with cache bypassed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Arm A&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;CARDIAC-PURR routing enabled:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;model=auto&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Arm B&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;LARGE-tier reference&lt;/strong&gt;:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;model=large&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;I expected the router to reduce cost.&lt;/p&gt;

&lt;p&gt;The first result went the other way.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;CARDIAC-PURR&lt;/th&gt;
&lt;th&gt;LARGE-tier reference&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Total observed cost&lt;/td&gt;
&lt;td&gt;$0.868122&lt;/td&gt;
&lt;td&gt;$0.802010&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Difference&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;8.24% higher&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;CARDIAC-PURR had lost.&lt;/p&gt;

&lt;p&gt;I kept the result.&lt;/p&gt;

&lt;p&gt;More importantly, I investigated it.&lt;/p&gt;

&lt;p&gt;The problem was that the first run had left an output-budget variable uncontrolled. Different provider behavior interacted with the escalation path, causing some requests to be truncated and enter a more expensive recovery path.&lt;/p&gt;

&lt;p&gt;That made the first run useful as a diagnostic artifact.&lt;/p&gt;

&lt;p&gt;It did &lt;strong&gt;not&lt;/strong&gt; make it a clean economic comparison.&lt;/p&gt;

&lt;p&gt;So I did not use it as the headline result.&lt;/p&gt;

&lt;h2&gt;
  
  
  The corrected benchmark
&lt;/h2&gt;

&lt;p&gt;I reran the same 100-query cohort.&lt;/p&gt;

&lt;p&gt;The key change was simple:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;max_tokens=2048&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;was explicitly applied to both arms.&lt;/p&gt;

&lt;p&gt;The rest of the A/B structure and cache-bypass condition stayed the same.&lt;/p&gt;

&lt;p&gt;The corrected result was:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;CARDIAC-PURR&lt;/th&gt;
&lt;th&gt;LARGE-tier reference&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Total observed model cost&lt;/td&gt;
&lt;td&gt;$0.350698&lt;/td&gt;
&lt;td&gt;$3.997920&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Average cost/query&lt;/td&gt;
&lt;td&gt;$0.003507&lt;/td&gt;
&lt;td&gt;$0.039979&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost difference&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;91.2% lower&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cascades&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Retries&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That &lt;strong&gt;91.2%&lt;/strong&gt; figure is the result I am comfortable publishing.&lt;/p&gt;

&lt;p&gt;But it needs a precise definition:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;On this 100-query technical workload, CARDIAC-PURR produced 91.2% lower observed model inference cost than the realized LARGE-tier reference arm.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It is not a claim about total application cost.&lt;/p&gt;

&lt;p&gt;It is not a claim about TCO.&lt;/p&gt;

&lt;p&gt;And it is not a promise that every workload will save 91.2%.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where did the savings come from?
&lt;/h2&gt;

&lt;p&gt;The tier distribution tells most of the story.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Routed tier&lt;/th&gt;
&lt;th&gt;Queries&lt;/th&gt;
&lt;th&gt;CARDIAC-PURR cost&lt;/th&gt;
&lt;th&gt;Matched LARGE-reference cost&lt;/th&gt;
&lt;th&gt;Reduction&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;SMALL&lt;/td&gt;
&lt;td&gt;93&lt;/td&gt;
&lt;td&gt;$0.163543&lt;/td&gt;
&lt;td&gt;$3.715975&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;95.6%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MEDIUM&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;$0.065385&lt;/td&gt;
&lt;td&gt;$0.156775&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;58.3%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LARGE&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;$0.121770&lt;/td&gt;
&lt;td&gt;$0.125170&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2.7%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Ninety-three of the 100 requests stayed in SMALL.&lt;/p&gt;

&lt;p&gt;On those requests, the matched observed cost was 95.6% lower than the LARGE-tier reference.&lt;/p&gt;

&lt;p&gt;But the three requests that actually went to LARGE are just as interesting.&lt;/p&gt;

&lt;p&gt;They cost almost the same in both arms.&lt;/p&gt;

&lt;p&gt;That means the result wasn't simply:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Avoid the expensive model."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The router selected LARGE when LARGE was needed.&lt;/p&gt;

&lt;p&gt;The savings came primarily from not using LARGE when it wasn't needed.&lt;/p&gt;

&lt;p&gt;That is the economic behavior I was trying to measure.&lt;/p&gt;

&lt;h2&gt;
  
  
  One detail I don't want to hide
&lt;/h2&gt;

&lt;p&gt;The LARGE reference arm was not literally one identical provider for every request.&lt;/p&gt;

&lt;p&gt;Of the 100 realized reference calls:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;98 resolved to Gemini 2.5 Pro&lt;/li&gt;
&lt;li&gt;1 resolved to DeepSeek V4 Pro&lt;/li&gt;
&lt;li&gt;1 resolved to Command A&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;CARDIAC-PURR itself resolved to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;93 Gemini 2.5 Flash-Lite&lt;/li&gt;
&lt;li&gt;4 Gemini 2.5 Flash&lt;/li&gt;
&lt;li&gt;3 Gemini 2.5 Pro&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the 91.2% figure compares CARDIAC-PURR with the &lt;strong&gt;realized LARGE reference arm&lt;/strong&gt;, rather than pretending that every reference request was physically served by one identical provider path.&lt;/p&gt;

&lt;p&gt;I would rather disclose what actually happened than simplify the table until it tells a cleaner story.&lt;/p&gt;

&lt;h2&gt;
  
  
  There is another baseline, and it answers a different question
&lt;/h2&gt;

&lt;p&gt;The benchmark also produced a synthetic always-LARGE baseline of $1.048460.&lt;/p&gt;

&lt;p&gt;That is not a third live benchmark arm.&lt;/p&gt;

&lt;p&gt;It is a per-request calculation estimating what each request would have cost if it had been sent to the LARGE-tier model at the provider selected for that request.&lt;/p&gt;

&lt;p&gt;Against that baseline, CARDIAC-PURR was 66.6% lower.&lt;/p&gt;

&lt;p&gt;So:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;91.2% lower&lt;/strong&gt; = realized LARGE-tier reference arm&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;66.6% lower&lt;/strong&gt; = synthetic always-LARGE baseline&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those numbers should not be mixed.&lt;/p&gt;

&lt;p&gt;They answer different questions.&lt;/p&gt;

&lt;h2&gt;
  
  
  What about quality?
&lt;/h2&gt;

&lt;p&gt;This benchmark measured economics.&lt;/p&gt;

&lt;p&gt;It did not measure answer quality on these same 100 requests.&lt;/p&gt;

&lt;p&gt;That's important because cost savings without quality would be a bad trade.&lt;/p&gt;

&lt;p&gt;We have separate prior evaluation results in the &lt;strong&gt;96–100% range&lt;/strong&gt;, depending on the evaluation set and methodology.&lt;/p&gt;

&lt;p&gt;I am deliberately not attaching those results to the 91.2% number.&lt;/p&gt;

&lt;p&gt;They are separate experiments.&lt;/p&gt;

&lt;p&gt;The next benchmark needs to answer the question that matters most:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Does CARDIAC-PURR maintain answer quality while making these cheaper routing decisions on the same workload?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Until that is measured on the same cohort, I don't think it is responsible to claim otherwise.&lt;/p&gt;

&lt;h2&gt;
  
  
  And what about the failure path?
&lt;/h2&gt;

&lt;p&gt;The corrected benchmark produced &lt;strong&gt;zero cascades and zero retries&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That's good for the cleanliness of the cost comparison.&lt;/p&gt;

&lt;p&gt;But it also tells us what the experiment did not test.&lt;/p&gt;

&lt;p&gt;We measured successful tier selection.&lt;/p&gt;

&lt;p&gt;We did not exercise the economics of a request that starts in SMALL, fails or becomes inadequate, and then escalates.&lt;/p&gt;

&lt;p&gt;So this benchmark does not prove that the recovery path behaves optimally under failure.&lt;/p&gt;

&lt;p&gt;That is another experiment.&lt;/p&gt;

&lt;p&gt;Similarly, routing overhead was not separately measured as part of this economic result.&lt;/p&gt;

&lt;p&gt;That should be measured formally alongside inference cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  The framework investigation changed how I think about the problem
&lt;/h2&gt;

&lt;p&gt;The agent-framework episode ended up being useful for a reason that goes beyond that particular framework.&lt;/p&gt;

&lt;p&gt;A router sitting underneath an application does not necessarily receive a clean user question.&lt;/p&gt;

&lt;p&gt;It receives whatever the application sends.&lt;/p&gt;

&lt;p&gt;That can include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;framework instructions&lt;/li&gt;
&lt;li&gt;task context&lt;/li&gt;
&lt;li&gt;tool-related information&lt;/li&gt;
&lt;li&gt;output-format requirements&lt;/li&gt;
&lt;li&gt;generated memory or state&lt;/li&gt;
&lt;li&gt;other application-level scaffolding&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The underlying user request may remain simple while the actual prompt structure becomes much more complicated.&lt;/p&gt;

&lt;p&gt;That means a routing layer cannot be designed and calibrated only around direct API calls.&lt;/p&gt;

&lt;p&gt;It has to work with the way real AI applications construct requests.&lt;/p&gt;

&lt;p&gt;That's where the control plane idea becomes important.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why routing belongs at the infrastructure layer
&lt;/h2&gt;

&lt;p&gt;A lot of AI architecture still looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Application
    ↓
Model
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That works when there is one model and one provider.&lt;/p&gt;

&lt;p&gt;It gets less comfortable when an organization has multiple models, multiple providers, different cost policies, failure handling, and increasingly complex AI applications.&lt;/p&gt;

&lt;p&gt;The architecture starts looking more like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Application
    ↓
Agent / Framework
    ↓
AI Control Plane
    ↓
Model / Provider
    ↓
Execution
    ↓
Recovery / Observability / Cost
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At that point, model routing is no longer just a convenience function inside an application.&lt;/p&gt;

&lt;p&gt;It affects infrastructure economics.&lt;/p&gt;

&lt;p&gt;It affects provider selection.&lt;/p&gt;

&lt;p&gt;It affects failure behavior.&lt;/p&gt;

&lt;p&gt;It affects observability.&lt;/p&gt;

&lt;p&gt;And it has to understand the traffic arriving from the application layer.&lt;/p&gt;

&lt;p&gt;That is the problem space CARDIAC-PURR is being built around.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I got wrong
&lt;/h2&gt;

&lt;p&gt;This is probably the part I value most in the whole exercise.&lt;/p&gt;

&lt;p&gt;I was too quick to trust an early benchmark because its output looked reasonable.&lt;/p&gt;

&lt;p&gt;I initially attributed a large latency anomaly too directly to routing behavior.&lt;/p&gt;

&lt;p&gt;The benchmark's timing accounting contained a defect.&lt;/p&gt;

&lt;p&gt;The first economic experiment omitted an output-budget control that mattered.&lt;/p&gt;

&lt;p&gt;And the later investigation showed that fixing one routing bug did not explain every observed latency difference.&lt;/p&gt;

&lt;p&gt;In other words, the system wasn't wrong in only one place.&lt;/p&gt;

&lt;p&gt;The router could be wrong.&lt;/p&gt;

&lt;p&gt;The measurement could be wrong.&lt;/p&gt;

&lt;p&gt;And the interpretation of the measurement could be wrong.&lt;/p&gt;

&lt;p&gt;That's why I became much more conservative about what I was willing to call a benchmark result.&lt;/p&gt;

&lt;p&gt;The earlier multi-framework measurements remain useful as development and investigation artifacts, but I no longer treat all of their historical performance or savings numbers as authoritative measurements of the corrected system.&lt;/p&gt;

&lt;p&gt;That is a much less exciting claim.&lt;/p&gt;

&lt;p&gt;I think it is also a more honest one.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I still need to test
&lt;/h2&gt;

&lt;p&gt;There is a lot left to do.&lt;/p&gt;

&lt;p&gt;The current result needs to be tested across different workloads, not just technical questions.&lt;/p&gt;

&lt;p&gt;Quality and cost need to be measured together on the same requests.&lt;/p&gt;

&lt;p&gt;Provider diversity needs to be tested deliberately.&lt;/p&gt;

&lt;p&gt;Failure and escalation economics need a dedicated experiment.&lt;/p&gt;

&lt;p&gt;The system needs to be tested on real multi-step agentic workflows rather than mostly single-turn requests.&lt;/p&gt;

&lt;p&gt;Routing overhead needs to be measured formally alongside model inference cost.&lt;/p&gt;

&lt;p&gt;The objective isn't to find another impressive percentage.&lt;/p&gt;

&lt;p&gt;It is to find out &lt;strong&gt;where the system stops working&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That is a much harder problem.&lt;/p&gt;

&lt;p&gt;It is also the more interesting one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bigger lesson
&lt;/h2&gt;

&lt;p&gt;I started with a straightforward economic problem:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Why should every LLM request cost as much as the hardest request in the system?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That led to a router.&lt;/p&gt;

&lt;p&gt;The router led to benchmarks.&lt;/p&gt;

&lt;p&gt;The benchmarks exposed bugs.&lt;/p&gt;

&lt;p&gt;The bugs forced a deeper look at the measurement process.&lt;/p&gt;

&lt;p&gt;And the investigation eventually led to a broader conclusion:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Model routing is becoming an infrastructure problem.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Once applications use multiple models and providers, the decision about which model handles which request sits at the intersection of application behavior, model capability, provider economics, reliability, and operational policy.&lt;/p&gt;

&lt;p&gt;That is bigger than a helper function.&lt;/p&gt;

&lt;p&gt;It's a control-plane problem.&lt;/p&gt;

&lt;p&gt;The corrected benchmark gives us one concrete result:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;91.2% lower observed model inference cost on this 100-query technical workload versus the realized LARGE-tier reference.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's meaningful.&lt;/p&gt;

&lt;p&gt;But the number is only useful because the process around it is visible: the bad result, the investigation, the corrections, and the limitations are part of the result too.&lt;/p&gt;

&lt;p&gt;I built CARDIAC-PURR because I believed LLM routing needed a better infrastructure layer.&lt;/p&gt;

&lt;p&gt;I'm publishing the benchmark because I think that belief should be tested, not merely asserted.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Trust is something a benchmark has to earn.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Disclosure:&lt;/strong&gt; I’m the founder of CARDIAC-PURR. This article describes our own engineering work and benchmark results. CARDIAC-PURR is a product in development. I’m intentionally not disclosing proprietary routing mechanisms or implementation details.&lt;/p&gt;

&lt;p&gt;Emil Igidov&lt;br&gt;
Kraków, Poland&lt;br&gt;
Founder, CARDIAC-PURR · Patent-Pending Inventor&lt;/p&gt;

&lt;p&gt;Learn more about CARDIAC-PURR: &lt;a href="https://www.cardiac-purr.com" rel="noopener noreferrer"&gt;https://www.cardiac-purr.com&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>infrastructure</category>
      <category>llm</category>
      <category>performance</category>
    </item>
    <item>
      <title>I Built CARDIAC-PURR: What 900 Real API Calls Taught Me About LLM Cost Routing</title>
      <dc:creator>Emil</dc:creator>
      <pubDate>Thu, 03 Sep 2026 04:26:52 +0000</pubDate>
      <link>https://dev.to/igidov/cardiac-purr-i-built-an-llm-cost-router-heres-what-100-questions-per-provider-across-9-3li9</link>
      <guid>https://dev.to/igidov/cardiac-purr-i-built-an-llm-cost-router-heres-what-100-questions-per-provider-across-9-3li9</guid>
      <description>&lt;p&gt;You ship a feature, it calls GPT-4, it works beautifully, you move on. Six months later, someone checks the logs and — surprise — half your traffic is "what's my account number" going through a model smart enough to pass the bar exam. This isn't a one-off mistake, either — it's the default outcome of not routing at all. Support-queue traffic is mostly simple, but "mostly simple" tends to get shipped through whatever model handled the hard case that made you reach for a big model in the first place. The model's not the problem. The lack of a router is.&lt;/p&gt;

&lt;p&gt;I built a router to fix that. I'm an attorney by training, not an engineer, so instead of trusting my own intuition (or published benchmarks), I decided to measure it properly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What I Built&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;CARDIAC-PURR sits between your app and your LLM providers, monitoring each query and deciding whether it actually needs a large model or would do fine with a cheaper one. I tested it against 9 providers, each with its own small→medium→large split:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Provider&lt;/th&gt;
&lt;th&gt;Small&lt;/th&gt;
&lt;th&gt;Medium&lt;/th&gt;
&lt;th&gt;Large&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Anthropic&lt;/td&gt;
&lt;td&gt;claude-haiku-4-5-20251001&lt;/td&gt;
&lt;td&gt;claude-sonnet-4-6&lt;/td&gt;
&lt;td&gt;claude-opus-4-6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI&lt;/td&gt;
&lt;td&gt;gpt-4.1-nano&lt;/td&gt;
&lt;td&gt;gpt-4.1-mini&lt;/td&gt;
&lt;td&gt;gpt-4.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Google&lt;/td&gt;
&lt;td&gt;gemini-2.5-flash-lite&lt;/td&gt;
&lt;td&gt;gemini-2.5-flash&lt;/td&gt;
&lt;td&gt;gemini-2.5-pro&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Azure OpenAI&lt;/td&gt;
&lt;td&gt;gpt-4.1-nano&lt;/td&gt;
&lt;td&gt;gpt-4.1-mini&lt;/td&gt;
&lt;td&gt;gpt-4.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mistral&lt;/td&gt;
&lt;td&gt;mistral-small-latest&lt;/td&gt;
&lt;td&gt;mistral-medium-latest&lt;/td&gt;
&lt;td&gt;mistral-large-latest&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek&lt;/td&gt;
&lt;td&gt;deepseek-v4-flash&lt;/td&gt;
&lt;td&gt;deepseek-v4-flash†&lt;/td&gt;
&lt;td&gt;deepseek-v4-pro&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cohere&lt;/td&gt;
&lt;td&gt;command-r7b-12-2024&lt;/td&gt;
&lt;td&gt;command-r-plus-08-2024&lt;/td&gt;
&lt;td&gt;command-a-03-2025&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Grok&lt;/td&gt;
&lt;td&gt;grok-4.3§&lt;/td&gt;
&lt;td&gt;grok-4.3§&lt;/td&gt;
&lt;td&gt;grok-4.3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen&lt;/td&gt;
&lt;td&gt;qwen-turbo&lt;/td&gt;
&lt;td&gt;qwen-plus&lt;/td&gt;
&lt;td&gt;qwen-max&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;† DeepSeek's small and medium tiers point at the identical model (&lt;code&gt;deepseek-v4-flash&lt;/code&gt;) — there's no true 3-way split for DeepSeek, just flash vs. pro. That sounds like a limitation, and I originally treated it as one, but it turns out to be the reason DeepSeek posts the highest savings of any provider tested: flash is dramatically cheaper than pro, and 92% of traffic never touches pro at all. Full explanation, including why an earlier 46.0% estimate for DeepSeek was wrong, is below.&lt;/p&gt;

&lt;p&gt;§ This resolves an open flag from an earlier draft — I'd noted that xAI retired the standalone &lt;code&gt;grok-4-1-fast&lt;/code&gt; endpoints on May 15, 2026 and worried the small/medium/large split might secretly be hitting one model. Turns out that's exactly what's happening, confirmed directly from the router's own config: all three Grok tiers call &lt;code&gt;grok-4.3&lt;/code&gt;. The only lever available is a &lt;code&gt;reasoning_effort&lt;/code&gt; parameter (none/low/high) — a compute-budget knob on one model, not a routing decision between differently-priced models. Details and what this does to Grok's savings ceiling are below.&lt;/p&gt;

&lt;p&gt;Tier selection happens before inference, based on query complexity — no waiting for a model response, no trained classifier, no embedding model.&lt;/p&gt;

&lt;p&gt;Every query gets a complexity score against calibrated thresholds (&lt;code&gt;self.c_target&lt;/code&gt;, currently 0.600 for the LARGE-tier gate). Most queries score nowhere near a boundary — the 100-question set was intentionally realistic, not adversarial, and the accuracy numbers below reflect that. Queries that land close to a threshold are exactly where I'd expect the router to be least confident, and it's also where the escalation safety net matters most: if a borderline SMALL-tier pick produces a weak or truncated response, that's caught and escalated rather than silently returned. I don't yet have a clean breakdown of accuracy specifically on near-boundary queries in this run — that's on my list for the next pass.&lt;/p&gt;

&lt;p&gt;There's also a safety net: if the tier it picked gives a truncated or weak-looking response, the router escalates to a bigger model. That's a fallback, not the main path.&lt;/p&gt;

&lt;p&gt;Under the hood, there are 15 distinct cost-saving mechanisms and 5 layers of failure/quality prevention (traffic protection, security, cost control, data durability, observability) working together. Patent-pending, so I won't break them down in detail here — but the numbers below are what all of them add up to in production.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How I Tested It&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I've been running this in production since March. For this post, I wanted a clean, controlled comparison: same 100 questions, asked to nine providers, so the numbers below are apples-to-apples.&lt;/p&gt;

&lt;p&gt;I ran 100 queries per provider — Anthropic, OpenAI, Google, Azure OpenAI, Mistral, DeepSeek, Cohere, Grok, and Qwen. Real money spent. Real API calls. Same 100 questions across all nine providers.&lt;/p&gt;

&lt;p&gt;The queries were intentionally unglamorous: 75% factual definitions ("What does NDA stand for?"), 17% explanatory ("Describe the main steps in..."), 8% complex reasoning. This is what actual support queues look like. The full 100-question set, labeled by expected tier and vertical, is linked at the bottom of this post — if you think my distribution is wrong for your workload, don't take my word for it; run it yourself.&lt;/p&gt;

&lt;p&gt;I measured two things separately because mixing them is how benchmarks lie:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Did the router pick the right tier? (routing accuracy)&lt;/li&gt;
&lt;li&gt;Did accuracy hold up across domains, not just in aggregate? (vertical accuracy — legal, healthcare, finance, IT)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;The Numbers&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A note before the table: with n=100 per provider, 100% accuracy doesn't mean zero future misclassifications — it means the 95% confidence interval is [96.4%, 100%]. I'm expanding sample size (see "What Comes Next"), and in the meantime the full methodology below is reproducible — run it on your own data and see where your numbers land.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Run date: July 12, 2026 — router commit &lt;code&gt;57175b7&lt;/code&gt;, which fixed a warmup-accounting bug that had previously forced the first ~47 post-warmup queries to the LARGE tier regardless of actual complexity. This is a clean rerun; the numbers below supersede any earlier pass.&lt;/em&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Provider&lt;/th&gt;
&lt;th&gt;Accuracy&lt;/th&gt;
&lt;th&gt;Savings&lt;/th&gt;
&lt;th&gt;Fail %&lt;/th&gt;
&lt;th&gt;Errors&lt;/th&gt;
&lt;th&gt;Vertical (Legal/Health/Finance/IT)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;88.3%†&lt;/td&gt;
&lt;td&gt;0.0%&lt;/td&gt;
&lt;td&gt;0/100&lt;/td&gt;
&lt;td&gt;100/100/100/100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;86.3%&lt;/td&gt;
&lt;td&gt;0.0%&lt;/td&gt;
&lt;td&gt;0/100&lt;/td&gt;
&lt;td&gt;100/100/100/100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;84.9%&lt;/td&gt;
&lt;td&gt;0.0%&lt;/td&gt;
&lt;td&gt;0/100&lt;/td&gt;
&lt;td&gt;100/100/100/100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Azure OpenAI&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;84.9%&lt;/td&gt;
&lt;td&gt;0.0%&lt;/td&gt;
&lt;td&gt;0/100&lt;/td&gt;
&lt;td&gt;100/100/100/100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Google&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;84.4%&lt;/td&gt;
&lt;td&gt;0.0%&lt;/td&gt;
&lt;td&gt;0/100&lt;/td&gt;
&lt;td&gt;100/100/100/100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Anthropic&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;78.8%&lt;/td&gt;
&lt;td&gt;0.0%&lt;/td&gt;
&lt;td&gt;0/100&lt;/td&gt;
&lt;td&gt;100/100/100/100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mistral&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;77.7%&lt;/td&gt;
&lt;td&gt;0.0%&lt;/td&gt;
&lt;td&gt;0/100&lt;/td&gt;
&lt;td&gt;100/100/100/100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cohere⚠&lt;/td&gt;
&lt;td&gt;98.0%&lt;/td&gt;
&lt;td&gt;73.9%‡&lt;/td&gt;
&lt;td&gt;2.0%&lt;/td&gt;
&lt;td&gt;2/100&lt;/td&gt;
&lt;td&gt;96/100/96/100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Grok&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;30.9%§&lt;/td&gt;
&lt;td&gt;0.0%&lt;/td&gt;
&lt;td&gt;0/100&lt;/td&gt;
&lt;td&gt;100/100/100/100&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;† DeepSeek is the best performer in this benchmark, and the reason flips an earlier, wrong estimate: a prior pass reported 46.0% from a coarse price-ratio table, not real dollars. The actual number, computed from real per-query costs against DeepSeek's LARGE-tier price, is &lt;code&gt;1 - $0.0059/$0.0505 = 88.3%&lt;/code&gt;. Full mechanism explained in the DeepSeek/Grok section below.&lt;/p&gt;

&lt;p&gt;‡ Cohere's medium and large tiers are priced identically ($12.50/MTok combined), so there's no discount available between those two tiers by design — that structural fact caps Cohere's best-case savings below the other providers' ceiling regardless of errors. Separately, and unrelated to pricing: Cohere hit 98.0% accuracy (95% CI: 93.0–99.4%) because of two transient read timeouts — one on a legal MEDIUM-tier query, one on a financial LARGE-tier query. In both cases the router picked the correct tier and the underlying API call timed out before returning a response; these are provider/network-side failures, not routing misclassifications.&lt;/p&gt;

&lt;p&gt;§ Grok is the lowest-savings provider in this run — structural, not a routing weakness. Full explanation below.&lt;/p&gt;

&lt;p&gt;⚠ Cohere is the only provider with a nonzero error rate in this run (2/100 — real timed-out API calls, distinct from cascades, which were zero for every provider this run). Every other provider, including Anthropic, came back completely clean: 100% accuracy, 0 errors, no vertical dips.&lt;/p&gt;

&lt;p&gt;*Savings % isn't comparable across providers as a ranking — it reflects each provider's own small-vs-large price ratio, not routing quality. Routing accuracy is the fair cross-provider comparison, and on that metric 8 of 9 providers hit 100%.&lt;/p&gt;

&lt;p&gt;All queries were pre-labeled into expected tiers (small/medium/large) based on response complexity. Routing accuracy compares the router's decision against these labels.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Methodology: labeling, baseline, and confidence intervals.&lt;/strong&gt; Given the actual query mix (75% small / 17% medium / 8% large), a naive router that always guesses "small" gets 75% accuracy for free, with zero routing logic — that's the honest baseline to beat. (Random guessing among three tiers gives 33%, the more commonly cited comparison, but it's the weaker one given this distribution.)&lt;/p&gt;

&lt;p&gt;Per-provider 95% confidence intervals (Clopper-Pearson): providers at 100/100 sit at [96.38%, 100.00%]; Cohere at 98/100 sits at [93.0%, 99.4%].&lt;/p&gt;

&lt;p&gt;8 of 9 providers hit 100% routing accuracy against my 96% target; only Cohere (98%) came in under it, and that's down to two transient timeouts, not a misclassification — detailed above and again under Reliability. Anthropic's savings (78.8%) are the lowest among the clean-run providers, though, and that's worth explaining on its own:&lt;/p&gt;

&lt;p&gt;Claude's small tier (Haiku) has a 256-token output cap. When you ask medical or legal questions that need longer answers, Haiku hits that cap. The router detects this on the response and escalates to a larger tier. This is not a bug — it's the router correctly identifying when a cheap tier won't work.&lt;/p&gt;

&lt;p&gt;But here's exactly what "escalation" does and doesn't guarantee. If the small-tier answer scores as weak or truncated, the router retries once against the large-tier model. If that retry also fails — a provider timeout, rate limit, or outage — the error isn't propagated as a hard failure. It's caught and logged, and the router falls back to returning the pre-escalation answer rather than a 500, with the event tagged in response metadata (&lt;code&gt;escalation_failed: true&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;That's graceful degradation, not a guarantee that escalation always produces a better answer. An earlier draft of this post said "100% cascade recovery," which overstated it — the accurate claim is: escalation is a single retry, fully observable, that never turns into a hard failure on our end, but can't promise the retry itself succeeds when it depends on an upstream provider.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why Grok saves 30.9%, not 78%+ like Anthropic&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This one's worth a full explanation rather than a footnote, because the root cause surprised me. Every other provider in this benchmark gets its savings from calling a genuinely smaller, cheaper-per-token model for easy queries. Grok doesn't have that option right now: all three of its tiers — small, medium, and large — call the same model, &lt;code&gt;grok-4.3&lt;/code&gt;. The only lever the router has is a &lt;code&gt;reasoning_effort&lt;/code&gt; parameter (none/low/high), which is a compute-budget knob on one model, not a routing decision between differently-priced models.&lt;/p&gt;

&lt;p&gt;The measured per-tier cost ratios (relative to LARGE = 1.0) make this concrete:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Provider&lt;/th&gt;
&lt;th&gt;SMALL ratio&lt;/th&gt;
&lt;th&gt;MEDIUM ratio&lt;/th&gt;
&lt;th&gt;LARGE ratio&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Anthropic&lt;/td&gt;
&lt;td&gt;0.052&lt;/td&gt;
&lt;td&gt;0.55&lt;/td&gt;
&lt;td&gt;1.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Grok&lt;/td&gt;
&lt;td&gt;0.605&lt;/td&gt;
&lt;td&gt;0.924&lt;/td&gt;
&lt;td&gt;1.0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Anthropic's small tier (Haiku) costs 5.2% of the large (Opus) because it's a fundamentally smaller model. Grok's small tier costs 60.5% of the large because it's the same model just told to think less — fewer reasoning tokens burned before answering, same per-token price. That gap in ratios is the entire explanation for why Grok tops out at 30.9% savings while Anthropic gets 78.8%+. It's an honest, structurally-limited ceiling given how xAI currently exposes Grok, not a routing weakness on my end. Getting Grok into Anthropic's range would need xAI to expose an actual smaller, cheaper model for the tier — not just an effort knob on grok-4.3.&lt;/p&gt;

&lt;p&gt;While investigating this, I also found and fixed two real bugs. Neither explains the 30.9% ceiling itself — they were quietly suppressing accuracy and dollar-separation before the fix, independent of the structural issue above:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;_GROK_EFFORT_MAP&lt;/code&gt; was defined at function scope instead of module scope in &lt;code&gt;universal_http_client_v20.py&lt;/code&gt; — it was intermittently unreachable.&lt;/li&gt;
&lt;li&gt;The router's LARGE-tier gate used a hardcoded &lt;code&gt;c_current &amp;lt; 0.65&lt;/code&gt; threshold instead of the calibrated &lt;code&gt;self.c_target&lt;/code&gt; (0.600), which made LARGE unreachable for queries scoring between 0.600 and 0.649.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Both are fixed and spot-checked — 4/4 verification queries now route correctly with real dollar separation between tiers ($0.00037 → $0.00175 → $0.0045). The numbers in this post reflect the fixed router.&lt;/p&gt;

&lt;p&gt;The same identical-model pattern exists on DeepSeek (see the model table above), which is what prompted checking it too — but DeepSeek's story turned out to be the opposite of Grok's. Full explanation is in the main results table footnote above: DeepSeek's shared small/medium model is so much cheaper than the large tier that it posts 88.3% savings, the highest of any provider in this benchmark, once measured from real per-query dollars rather than an earlier coarse price-ratio estimate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Latency (Where I Lost Sleep)&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Routing decision time (pre-inference): measured in isolation over 3,000 samples — p50 0.42ms, p95 0.68ms, p99 0.97ms, max 1.85ms. Pure in-process computation, no I/O, so it stays sub-millisecond in the typical case.&lt;/p&gt;

&lt;p&gt;But end-to-end latency (routing + provider response), same run as the results above:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Provider&lt;/th&gt;
&lt;th&gt;P50&lt;/th&gt;
&lt;th&gt;P95&lt;/th&gt;
&lt;th&gt;P99&lt;/th&gt;
&lt;th&gt;Router Overhead&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Qwen&lt;/td&gt;
&lt;td&gt;880ms&lt;/td&gt;
&lt;td&gt;4.4s&lt;/td&gt;
&lt;td&gt;5.0s&lt;/td&gt;
&lt;td&gt;14.9ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI&lt;/td&gt;
&lt;td&gt;835ms&lt;/td&gt;
&lt;td&gt;2.4s&lt;/td&gt;
&lt;td&gt;3.7s&lt;/td&gt;
&lt;td&gt;15.0ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Azure OpenAI&lt;/td&gt;
&lt;td&gt;1.2s&lt;/td&gt;
&lt;td&gt;2.3s&lt;/td&gt;
&lt;td&gt;3.5s&lt;/td&gt;
&lt;td&gt;16.7ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Google&lt;/td&gt;
&lt;td&gt;717ms&lt;/td&gt;
&lt;td&gt;4.7s&lt;/td&gt;
&lt;td&gt;5.4s&lt;/td&gt;
&lt;td&gt;15.0ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Anthropic&lt;/td&gt;
&lt;td&gt;1.8s&lt;/td&gt;
&lt;td&gt;21.2s&lt;/td&gt;
&lt;td&gt;46.2s&lt;/td&gt;
&lt;td&gt;16.7ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mistral&lt;/td&gt;
&lt;td&gt;730ms&lt;/td&gt;
&lt;td&gt;4.4s&lt;/td&gt;
&lt;td&gt;5.5s&lt;/td&gt;
&lt;td&gt;15.6ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Grok&lt;/td&gt;
&lt;td&gt;1.0s&lt;/td&gt;
&lt;td&gt;5.5s&lt;/td&gt;
&lt;td&gt;7.2s&lt;/td&gt;
&lt;td&gt;15.5ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cohere&lt;/td&gt;
&lt;td&gt;3.7s&lt;/td&gt;
&lt;td&gt;11.7s&lt;/td&gt;
&lt;td&gt;16.8s&lt;/td&gt;
&lt;td&gt;622.7ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek&lt;/td&gt;
&lt;td&gt;1.6s&lt;/td&gt;
&lt;td&gt;9.2s&lt;/td&gt;
&lt;td&gt;22.9s&lt;/td&gt;
&lt;td&gt;15.9ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Router overhead itself is a tight 14.9–16.7ms across 8 of 9 providers — genuinely small and consistent, which is what I'd hoped for. Cohere's 622.7ms is the outlier, and this run I can point to exactly why: it's not steady-state routing cost; it's the retry overhead from the same two timeout events described above. Two calls that had to time out and retry pull the aggregate proxy-overhead number up dramatically even though 98 of Cohere's 100 calls behaved normally.&lt;/p&gt;

&lt;p&gt;Anthropic's P99 is the ugliest number in the table — 46.2 seconds. Worth explaining rather than hand-waving: router overhead for Anthropic was 16.7ms this run, same as every other clean provider, so the router itself isn't the delay. Anthropic's IT-vertical queries also had the longest average output in the run (309.9 tokens vs. 65–100 for most other providers on the same vertical), and the LARGE-tier prompts in the query set are genuinely demanding — think cross-border data-sharing risk under Schrems II, or designing zero-trust architecture for a 2,000-person org. That's Opus doing real reasoning work, not an infrastructure stall — though I don't have per-query timing saved from this run to point to the single slowest call directly, which is a gap in the benchmark script I'm fixing next.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reliability (The Part I Care About)&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;2 query errors out of 900 total calls (100 per provider × 9 providers, 0.22%) — both on Cohere, both read timeouts (one MEDIUM-tier legal query, one LARGE-tier financial query), 0 across the other 8 providers&lt;/li&gt;
&lt;li&gt;Memory footprint: baseline RSS 36.8–45.4 MB, peak RSS 38.0–46.6 MB across providers — stable, no leak pattern across the run&lt;/li&gt;
&lt;li&gt;Cascade rate: 0.00% for all 9 providers this run — no query required escalation beyond its initially assigned tier&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Separately, an earlier QA pass (May 2026) ran 751 automated tests across unit, integration, and prompt-injection suites — 751/751 passing at the time. Happy to share the suite breakdown if useful; I didn't want to pad this post with test-framework details.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What Happens When Things Break&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Clean benchmark numbers are half the story — what matters is what happens when a provider goes down mid-write. Three decisions I'd stand behind: Redis calls fail open (semantic cache and rate limiter degrade gracefully instead of taking the router down with them); every billing write that can't reach Postgres falls back to an fsync'd on-disk log and rehydrates automatically once the database is back, because "the database hiccuped" should never quietly become "we lost a customer's usage record"; and the provider circuit breaker is shared across worker processes, so one outage gets detected once, not four times. All of it is chaos-tested, not just unit-tested — killing Redis and Postgres mid-run, SIGTERM under load, injection attempts while infrastructure is already degraded. Happy to go deeper on any of this in the comments.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where the Router Struggles (Honestly)&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In this 9-provider run, the cascade rate was 0.0% for every provider, across all four verticals (legal, healthcare, finance, IT). That's not a gap in instrumentation; the field is populated and genuinely zero this time. It doesn't mean the mechanism is broken or unused, though. Archived test records show the cascade logic firing exactly as designed elsewhere: small-tier Anthropic queries truncated at the 256-token ceiling, correctly flagged as structurally weak, escalated to the large-tier model, with the escalation reflected in both the response metadata (&lt;code&gt;cascaded: true&lt;/code&gt;) and a real, higher cost for that specific query. The mechanism works. This particular 9-provider run's query/tier mix just didn't happen to trigger it.&lt;/p&gt;

&lt;p&gt;For real cascade examples, I'm pulling from an earlier, separate 4-provider deep-dive run, where cascading did occur:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Anthropic, on "What does the HIPAA Privacy Rule require from covered entities?" — the small-tier response was insufficient and the router cascaded to a larger tier, producing a 384-token response at $0.00625 (vs. an average $0.00055 for a small-tier query on Anthropic).&lt;/li&gt;
&lt;li&gt;Azure OpenAI, on "What does contraindicated mean?" — same pattern, cascading to an 80-token response at $0.000436.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Both fell into the same category: clinical or regulatory terminology needing a qualifying clause to answer correctly. In that run, only Anthropic cascaded on the HIPAA question and only Azure cascaded on "contraindicated" — Google and OpenAI both had zero cascades on the same two queries. The failure mode is systematic (a query category, not noise); which provider hits it on a given day is not — and in the newer, larger run above, it didn't happen at all.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Economic Reality&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;CARDIAC-PURR runs on a savings-share model: no per-request fee, no platform fee, and nothing owed in any month where verified savings are $0.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How "verified savings" actually works, and why this is different from the benchmark numbers above.&lt;/strong&gt; Every month, we measure two things: your &lt;strong&gt;baseline&lt;/strong&gt; (your average LLM/API spend over the trailing 90 days on the same workloads) and your &lt;strong&gt;current&lt;/strong&gt; spend with the router turned on. Verified savings = baseline − current, based on your actual provider invoices plus our usage analytics, and agreed with you before anything is billed. That's the only number we price against.&lt;/p&gt;

&lt;p&gt;One distinction matters here: this uses a different baseline than the benchmark section above. The benchmark's "Baseline Cost (Always-LARGE)" is a synthetic comparison — what the same query set would have cost if every query hit the large-tier model — used to measure routing quality in a controlled test. Your actual billing baseline is your own real, historical spend, not a synthetic always-large figure. The two will generally differ, sometimes substantially, depending on how much oversized-model usage you already have in production.&lt;/p&gt;

&lt;p&gt;Pricing is a savings-share model that scales automatically with verified monthly savings: free up to 15K requests/month, then 75–82% of verified savings stays with you across all paid tiers, with a monthly cap so there's never an open-ended bill regardless of scale. Full tier breakdown is on the site — it aligns our incentives with yours (if we don't reduce your spend, we don't earn) and gives finance teams a predictable ceiling.&lt;/p&gt;

&lt;p&gt;Actual payback depends on your real trailing-90-day baseline, provider mix, and volume — not the benchmark numbers above. Calculate it against your own invoices. (There's also an async batch mode for latency-tolerant queries, roughly +5–6 points on Anthropic and Azure in earlier, not-yet-reverified testing.)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How This Compares to Other Routers, Mechanically&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;People will ask about OpenRouter and Martian, so here's the technical distinction rather than a pricing pitch. OpenRouter's default Auto Router is a meta-model call: your prompt goes to a model that decides which of dozens of downstream models should handle it, then forwards the request — an extra inference step, with the meta-model's own latency and judgment in the loop, tuned for output quality across providers rather than a specific cost target. Martian's router works differently again — it's built on what they call Model Mapping, converting model internals into a more interpretable form so it can predict a candidate model's expected quality and cost for a given prompt without necessarily running it.&lt;/p&gt;

&lt;p&gt;CARDIAC-PURR's routing decision happens before any inference call, using a calibrated complexity score against fixed thresholds (&lt;code&gt;c_target&lt;/code&gt;) rather than a second model call or an internals-based prediction. No meta-model latency, no black-box quality prediction — a deterministic score, a tier decision, then one inference call to the chosen tier. The tradeoff is the mirror image of both approaches: no meta-model means faster, more predictable routing decisions, but the classifier is calibrated against my own benchmark distribution, not a general-purpose prompt-quality predictor — which is exactly why the per-provider tier-pairing table above matters more here than a claimed "best model" headline would.&lt;/p&gt;

&lt;p&gt;I haven't run a head-to-head benchmark against either of them, and I'd rather say that plainly than imply a comparison I haven't done.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What the Benchmark Doesn't Measure&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Production at scale (I tested 100 questions per provider, not 1M)&lt;/li&gt;
&lt;li&gt;Provider errors and outages at higher volume — 0.22% here, two timed-out calls on 1 of 9 providers, isn't the same as what a real outage looks like under sustained load&lt;/li&gt;
&lt;li&gt;Your specific query distribution (mine was mostly 75/17/8; yours might be different)&lt;/li&gt;
&lt;li&gt;What happens when models change (testing was conducted over several months, with the most recent full benchmark run completed today; new models will behave differently)&lt;/li&gt;
&lt;li&gt;Correctness of the underlying model's answer, at any tier. The router decides which model answers a query — it doesn't verify the answer itself. For high-stakes domains (legal, medical, financial), that verification step is still on you; a "correctly routed" query and a "correct answer" are different things, and I'd treat the vertical-accuracy numbers above as a tier-selection metric, not an answer-quality guarantee.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;How to Use This&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The router is available as:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Managed platform: &lt;a href="https://www.cardiac-purr.com/PLATFORM/pricing.html" rel="noopener noreferrer"&gt;https://www.cardiac-purr.com/PLATFORM/pricing.html&lt;/a&gt; — click through, get an API key immediately, free tier needs no credit card&lt;/li&gt;
&lt;li&gt;Self-hosted Docker: runs in your own VPC/on-prem, your own Postgres, your own API keys&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Source code isn't open (patent-pending) — more on why at the end.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Compliance &amp;amp; Production Status&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;0.22% overall error rate in testing (2 of 900 total calls, 100 per provider — both Cohere read timeouts), but caveat emptor in production at higher volume&lt;/li&gt;
&lt;li&gt;HIPAA BAA available for self-hosted&lt;/li&gt;
&lt;li&gt;GDPR Data Processing Agreement ready&lt;/li&gt;
&lt;li&gt;U.S. Provisional Patent 64/005,834 (May 24, 2026)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why I'm Publishing This&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I'm tired of reading benchmark posts that hide the methodology. "We tested on our internal queries" tells you nothing about whether it applies to your traffic. "75% simple, 17% explanatory, 8% complex across 4 domains, same 100 questions per provider, across 9 providers" — that, you can actually check your own workload against.&lt;/p&gt;

&lt;p&gt;The code runs in Docker. The methodology is reproducible. If you disagree with my query distribution, run your own benchmark.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What Comes Next&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I'm expanding the per-provider query count beyond 100 for tighter confidence intervals — Cohere's CI is the widest right now, so that's first. I'm waiting for customer feedback on what breaks in production.&lt;/p&gt;

&lt;p&gt;If you want to try it, the Free tier is $0/month forever, up to 15K requests/month, no credit card required — not a trial, a structural zero under that threshold. Use your own API keys (your data, your keys, your infrastructure).&lt;/p&gt;

&lt;p&gt;Questions about "why not open-source" are fair — short answer: solo founder, patent-pending, I need the runway before I can afford to give away the moat. That changes once the patent grants or the moat shifts elsewhere. Source access for compliance audits is available under NDA in the meantime.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Emil Igidov&lt;/strong&gt;&lt;br&gt;
Kraków, Poland&lt;br&gt;
Founder, CARDIAC-PURR · Patent-Pending Inventor. &lt;br&gt;
Learn more about CARDIAC-PURR: &lt;strong&gt;&lt;a href="https://www.cardiac-purr.com" rel="noopener noreferrer"&gt;https://www.cardiac-purr.com&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>architecture</category>
    </item>
  </channel>
</rss>
