<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Merriban</title>
    <description>The latest articles on DEV Community by Merriban (@merriban).</description>
    <link>https://dev.to/merriban</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4076828%2Fdd5d6b6e-10d5-4145-9a0f-78ae61c3a356.png</url>
      <title>DEV Community: Merriban</title>
      <link>https://dev.to/merriban</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/merriban"/>
    <language>en</language>
    <item>
      <title>I modeled the Jevons paradox for TypeSafe's Jev. My first model was wrong, twice.</title>
      <dc:creator>Merriban</dc:creator>
      <pubDate>Thu, 08 Oct 2026 11:00:00 +0000</pubDate>
      <link>https://dev.to/merriban/i-modeled-the-jevons-paradox-for-typesafes-jev-my-first-model-was-wrong-twice-3fh4</link>
      <guid>https://dev.to/merriban/i-modeled-the-jevons-paradox-for-typesafes-jev-my-first-model-was-wrong-twice-3fh4</guid>
      <description>&lt;p&gt;When TypeSafe AI launched Jev in September, they explained the name: it's a nod to William Stanley Jevons, who argued in 1865 that more efficient steam engines would mean &lt;em&gt;more&lt;/em&gt; coal burned, not less. TypeSafe expects intelligence to follow the same path.&lt;/p&gt;

&lt;p&gt;That raised a question I couldn't find answered anywhere: if AI decisions get radically cheaper, does total AI energy go down, or up?&lt;/p&gt;

&lt;p&gt;So I built an interactive scenario model: &lt;a href="https://merriban.github.io/jevons-jev/" rel="noopener noreferrer"&gt;https://merriban.github.io/jevons-jev/&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;Jev doesn't generate text. It takes a state plus typed questions (choice, score, yes/no) and returns decisions with probabilities. Per decision, it's much cheaper than asking an LLM: across the eight LLM configurations in TypeSafe's own workflow evals, the price ratio works out to roughly 77x (geometric mean).&lt;/p&gt;

&lt;p&gt;The model has five knobs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;s&lt;/code&gt;: share of AI inference energy spent on decision-type tasks&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;r&lt;/code&gt;: energy per decision, Jev vs. LLM&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;q&lt;/code&gt;: price per decision, Jev vs. LLM&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;c&lt;/code&gt;: share of decisions that still need an LLM call alongside Jev&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;eps&lt;/code&gt;: how strongly decision volume responds to cost (price elasticity)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The decision share of energy goes from &lt;code&gt;s·E0&lt;/code&gt; to &lt;code&gt;s·E0·(r+c)·(q+c)^(−eps)&lt;/code&gt;. The rest stays put.&lt;/p&gt;

&lt;p&gt;One honest problem up front: nobody has published Jev's energy use in watt-hours, as far as I could find. So &lt;code&gt;r&lt;/code&gt; is reconstructed from three indirect proxies (price, latency, token count), and two of them come from the same vendor eval. The page labels every number as cited, derived, or assumed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mistake 1: I priced complementarity wrong
&lt;/h2&gt;

&lt;p&gt;My first version grew decision volume as &lt;code&gt;q^(−eps)&lt;/code&gt;, as if every decision only cost Jev's price. Then it charged a full LLM call to the fraction &lt;code&gt;c&lt;/code&gt; of decisions that still need one.&lt;/p&gt;

&lt;p&gt;Those two things can't both be true. If a decision still calls the LLM, it didn't get 77x cheaper, so its demand can't explode as if it did.&lt;/p&gt;

&lt;p&gt;The effect was not small. With the defaults (&lt;code&gt;eps = 0.5&lt;/code&gt;, well below the Jevons threshold), the page showed energy rising 55.8%. A backfire, produced entirely by a modeling inconsistency. After the fix, the same settings show a 12.5% &lt;em&gt;drop&lt;/em&gt;. The share of Monte Carlo samples ending in backfire went from 93% to 62%.&lt;/p&gt;

&lt;p&gt;The fix also made the math cleaner. If energy tracks price (&lt;code&gt;r = q&lt;/code&gt;), the decision energy ratio is &lt;code&gt;(q+c)^(1−eps)&lt;/code&gt;. Backfire happens if and only if &lt;code&gt;eps &amp;gt; 1&lt;/code&gt;, whatever &lt;code&gt;c&lt;/code&gt; is.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mistake 2: I measured rebound against the wrong baseline
&lt;/h2&gt;

&lt;p&gt;Rebound is the share of expected savings that demand growth eats back. I computed expected savings as &lt;code&gt;1 − r&lt;/code&gt;, as if every decision stopped calling the LLM. But the &lt;code&gt;c&lt;/code&gt; share calls it anyway: that's not a behavioral response, it's the cost of the new setup.&lt;/p&gt;

&lt;p&gt;Against the right baseline, &lt;code&gt;1 − (r + c)&lt;/code&gt;, the default scenario's rebound drops from 57% to 38%. That moved it from "partial rebound" to "efficiency wins".&lt;/p&gt;

&lt;h2&gt;
  
  
  Why my tests didn't catch either one
&lt;/h2&gt;

&lt;p&gt;This is the part I keep thinking about. The project has a JS model and an independent Python reimplementation written from the methodology text, checked against each other on thousands of random points. Both agreed perfectly, both times.&lt;/p&gt;

&lt;p&gt;Of course they did. The bugs weren't in the code, they were in the spec. Differential testing proves two implementations match a description; it says nothing about whether the description makes economic sense. What caught both errors was reading a result and asking: "does this contradict the model's own threshold?"&lt;/p&gt;

&lt;h2&gt;
  
  
  The data had problems too
&lt;/h2&gt;

&lt;p&gt;The project was built with Claude Code, in an environment that couldn't reach most source websites. So I re-opened every primary source by hand. That found:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The first price ratio used only two of the eight LLM configurations in TypeSafe's eval table. The cheapest one with nearly identical accuracy made the upper bound of &lt;code&gt;q&lt;/code&gt; about 9x higher.&lt;/li&gt;
&lt;li&gt;A "40–400x cheaper" range that circulates in secondary coverage isn't in TypeSafe's own launch post. Their own peak claim is 444.6x.&lt;/li&gt;
&lt;li&gt;An energy-per-query estimate I had labeled GPU-only actually includes server and data-center overhead.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  So what's the answer?
&lt;/h2&gt;

&lt;p&gt;The robust result isn't a probability, it's a threshold. At the default settings, total AI inference energy rises above its 2025 baseline only if demand elasticity for AI decisions exceeds about 0.97.&lt;/p&gt;

&lt;p&gt;The Monte Carlo share (~62%) mostly reflects the elasticity range I chose (0.1 to 2.0), and the page says so. Nobody has measured that elasticity for AI decisions. That's the number to watch.&lt;/p&gt;

&lt;p&gt;It's a scenario tool, not a forecast. The model, every source, and the verification report are on GitHub: &lt;a href="https://github.com/Merriban/jevons-jev" rel="noopener noreferrer"&gt;https://github.com/Merriban/jevons-jev&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you find an error, open an issue. Two have already been found, by me. I'd rather the third one come from you.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>datascience</category>
      <category>opensource</category>
      <category>sustainability</category>
    </item>
  </channel>
</rss>
