For five articles I've been arguing that agents get captured one layer at a time: distribution (#45), the model (#46), identity (#47), access (#48), the harness (#49), and the meter (#50). Each time the answer sounded the same — own the layer, or at least own an exit from it. There was always an unspoken asterisk: owning the runtime wasn't really possible. You could own your prompts, your keys, your workflow — but the thing that actually does the work, the model, lived on someone else's silicon.
That asterisk just got smaller.
The runtime was the last rented layer
Everything downstream of the GPU was yours to control. Everything upstream — the weights, the inference, the tokens-per-second — was rented by the second. And renting the runtime meant renting the terms of the runtime: the price (which you don't set), the deprecation date (which you don't choose), the usage policy (which can change while your product is live), and the latency (which you can't fix).
This is why the "just self-host it" advice rang hollow for years. Two blockers killed it:
- Capability. Open weights lagged. If the model you needed was closed, "self-host" wasn't an option, it was a hobby.
- Feasibility. Even when weights were open, the hardware math was brutal. A 125B model at usable speed meant a rack, not a desktop.
Both blockers are the same blocker, and it is mostly gone.
125B at 100 tokens/sec on one consumer card
The top story this week was a demo, not a launch: running a 125-billion-parameter open model on a single RTX 4090 and getting 100 tokens per second. Read that again — a model class that used to imply a datacenter, streaming at a speed that is interactive, on a GPU a serious freelancer already owns.
This is the moment "self-host" stops being ideology and starts being arithmetic. If a sovereign runtime can hit 100T/s on hardware you can buy, then the runtime is no longer the layer you're forced to rent. It becomes the layer you choose to rent when renting is cheap — the same way you use a cloud for storage but keep a local backup.
And it arrives the same week two smaller stories make the same point from the other direction:
- "Turn off Apple Intelligence and get your disk space back." A platform shipped an AI feature no one asked for, and the top comment was about how to remove it. When the runtime is rented, the only exit is the off switch — and sometimes not even that.
- "Why don't more developers 'use the platform'?" Because "the platform" is a set of terms, not a technology. Developers avoid platforms they can't leave. The ones they adopt are the ones with a real exit ramp.
Exit ramps are the whole game. A sovereign runtime is just the exit ramp with the least friction, because the exit is the weights on your own disk.
What "sovereign runtime" actually requires
Owning the runtime is not one decision. It's four, and they can be scored on a cost curve:
1. Weights you can obtain and store. Open weights are necessary but not sufficient — a runtime you can't download isn't sovereign. Audit: if the license lets the provider revoke your access, you don't have the weights, you have access to the weights.
2. Hardware with a path to break-even. The question isn't "does it run?" — it's "how many tokens until the card pays for itself?" On a 4090 at 100T/s, that number is now small enough to calculate instead of speculate.
3. A serving path you can operate. Inference engines, quantization, batching, context management. This is where most teams stall — not on the model, on the plumbing. It's also the part that makes a service worth paying for, which is the interesting part.
4. A fallback, not a replacement. Sovereignty doesn't mean never touching a cloud. It means the cloud is a choice with a price, not a dependency without one. Keep one closed model for the frontier cases you can't run locally; keep everything else on metal you own.
Why this matters most for cross-border operators
If you run a thin-margin, unattended business across time zones, the rented runtime is your worst dependency:
- Your supplier's price is your cost of goods — and you don't set it, and it changes without a change in your code.
- Your supplier's deprecation is your roadmap — a model you built on can go away while your integration is live.
- Your supplier's jurisdiction is your compliance risk — data travels to wherever the rented runtime lives.
- Your supplier's outage is your outage — no SLA means no remedy.
A runtime you own converts those four unbounded risks into one bounded cost: capex. That's the trade — and for the first time, the capex side of the ledger is getting small enough to make.
The takeaway
The "capture" story has a happy ending that sounds like a load order. Own your distribution or rent it cheap. Own your model identity and your fallback. Own your harness. Put a ceiling on your meter. And now: own a runtime good enough that the cloud has to compete for your attention instead of setting your terms.
You don't have to run everything locally. You just have to run something locally well enough that "we'll move it in-house" is a sentence you can say and mean.
Four articles ago I wrote "stop renting your distribution." This one is the sequel: stop assuming the runtime is always rented. The demo this week says that assumption is now a choice.
Top comments (0)