DEV Community

JE Ramos
JE Ramos

Posted on

You are renting the decisions that set your AI bill.

Frontier models reached your engineers before any standard for AI-driven development did. The teams adopted them, the work changed, and nobody wrote down how this company ships software now. Most of us did it in that order.

So the harness decides instead. It picks the model, decides how long to think, when to retry, and when to stop. Nobody in your company approved any of that, and nobody in your company can change it.

That is where the money goes. Reasoning models bill their thinking at output rates, and on a complex task the thinking can be 70 to 85% of the output bill. You never see those tokens. You pay for every one of them. The frontier price has not moved to rescue you either: GPT-4 was $60 per million output tokens in 2023, and list price for Claude Fable 5 today is $50, with GPT-5.6 Sol at $30. Smarter, not cheaper.

The labs made the models smarter and handed the engineering responsibility back to you. They are teaching us to be lazy and charging us for the privilege.

So the decision in front of you is not which model. It is whether you build and own your agentic loops or keep renting them.

Build when AI-driven development is becoming how you ship. Then the routing, the spend cap, the stop condition and the exit are yours, they have an owner, and they get reviewed like any other engineering decision.

Rent when it is genuinely peripheral, and then hold it to a supplier standard: a cost per completed task you can quote, and a second path you have actually tested.

Because we have been here before with other vendors and we know how it ends. Too much dependency on one company is a bad thing, and an enterprise at scale always has a contingency plan. That second path is real now. Moonshot shipped Kimi K3 with open weights in July, listed at $3 in and $15 out, and independent testing reported it finishing a task for roughly half what Opus 4.8 costs. It carries a different kind of bill: Cursor built Composer 2 on Moonshot's model and drew a House Committee investigation into enterprise use of Chinese AI. A contingency plan is never free, which is why it belongs on a board agenda and not in procurement.

Cost will always be a problem. Every business wants high margins, and that means lower cost and higher throughput. But cost on its own decides nothing. If the output outweighs what it costs, spend. If the quality is good enough that the customer comes back, the expensive version is the cheaper one.

Every enterprise is in a different state of AI adoption, and there is no reason left to experiment. The blueprints exist. But the decision still has to be made by someone.

So who is making it in your company right now, your engineers or the harness?

Top comments (0)