When TypeSafe AI launched Jev in September, they explained the name: it's a nod to William Stanley Jevons, who argued in 1865 that more efficient steam engines would mean more coal burned, not less. TypeSafe expects intelligence to follow the same path.
That raised a question I couldn't find answered anywhere: if AI decisions get radically cheaper, does total AI energy go down, or up?
So I built an interactive scenario model: https://merriban.github.io/jevons-jev/
The setup
Jev doesn't generate text. It takes a state plus typed questions (choice, score, yes/no) and returns decisions with probabilities. Per decision, it's much cheaper than asking an LLM: across the eight LLM configurations in TypeSafe's own workflow evals, the price ratio works out to roughly 77x (geometric mean).
The model has five knobs:
-
s: share of AI inference energy spent on decision-type tasks -
r: energy per decision, Jev vs. LLM -
q: price per decision, Jev vs. LLM -
c: share of decisions that still need an LLM call alongside Jev -
eps: how strongly decision volume responds to cost (price elasticity)
The decision share of energy goes from s·E0 to s·E0·(r+c)·(q+c)^(−eps). The rest stays put.
One honest problem up front: nobody has published Jev's energy use in watt-hours, as far as I could find. So r is reconstructed from three indirect proxies (price, latency, token count), and two of them come from the same vendor eval. The page labels every number as cited, derived, or assumed.
Mistake 1: I priced complementarity wrong
My first version grew decision volume as q^(−eps), as if every decision only cost Jev's price. Then it charged a full LLM call to the fraction c of decisions that still need one.
Those two things can't both be true. If a decision still calls the LLM, it didn't get 77x cheaper, so its demand can't explode as if it did.
The effect was not small. With the defaults (eps = 0.5, well below the Jevons threshold), the page showed energy rising 55.8%. A backfire, produced entirely by a modeling inconsistency. After the fix, the same settings show a 12.5% drop. The share of Monte Carlo samples ending in backfire went from 93% to 62%.
The fix also made the math cleaner. If energy tracks price (r = q), the decision energy ratio is (q+c)^(1−eps). Backfire happens if and only if eps > 1, whatever c is.
Mistake 2: I measured rebound against the wrong baseline
Rebound is the share of expected savings that demand growth eats back. I computed expected savings as 1 − r, as if every decision stopped calling the LLM. But the c share calls it anyway: that's not a behavioral response, it's the cost of the new setup.
Against the right baseline, 1 − (r + c), the default scenario's rebound drops from 57% to 38%. That moved it from "partial rebound" to "efficiency wins".
Why my tests didn't catch either one
This is the part I keep thinking about. The project has a JS model and an independent Python reimplementation written from the methodology text, checked against each other on thousands of random points. Both agreed perfectly, both times.
Of course they did. The bugs weren't in the code, they were in the spec. Differential testing proves two implementations match a description; it says nothing about whether the description makes economic sense. What caught both errors was reading a result and asking: "does this contradict the model's own threshold?"
The data had problems too
The project was built with Claude Code, in an environment that couldn't reach most source websites. So I re-opened every primary source by hand. That found:
- The first price ratio used only two of the eight LLM configurations in TypeSafe's eval table. The cheapest one with nearly identical accuracy made the upper bound of
qabout 9x higher. - A "40–400x cheaper" range that circulates in secondary coverage isn't in TypeSafe's own launch post. Their own peak claim is 444.6x.
- An energy-per-query estimate I had labeled GPU-only actually includes server and data-center overhead.
So what's the answer?
The robust result isn't a probability, it's a threshold. At the default settings, total AI inference energy rises above its 2025 baseline only if demand elasticity for AI decisions exceeds about 0.97.
The Monte Carlo share (~62%) mostly reflects the elasticity range I chose (0.1 to 2.0), and the page says so. Nobody has measured that elasticity for AI decisions. That's the number to watch.
It's a scenario tool, not a forecast. The model, every source, and the verification report are on GitHub: https://github.com/Merriban/jevons-jev
If you find an error, open an issue. Two have already been found, by me. I'd rather the third one come from you.
Top comments (0)