DEV Community

Cover image for The $40 Question: When Execution Gets Cheap, Imagination Is the Constraint
Cleber de Lima
Cleber de Lima

Posted on

The $40 Question: When Execution Gets Cheap, Imagination Is the Constraint

The most valuable work in your company right now is on nobody's backlog, and your operating model has no way to fund it.

A backlog is a list of work somebody already knew how to ask for, and for twenty years that was its strength. Execution was the expensive part of software, so the discipline that mattered was keeping a clean, prioritized list of known work and grinding it down. Everything about how we run engineering organizations, the ceremonies, the metrics, the budgets, is built on the assumption that the list is where the value lives. That assumption has quietly inverted. The work that can be written as a ticket is exactly the work that has become cheap, and the work that matters most now never reaches a ticket at all.

The cleanest evidence I have seen for this is an experiment Mitchell Hashimoto published in June. He ran ordinary backlog work, the "implement this feature" kind of task, through three models head to head: GLM-5.1 finished in minutes for under a dollar, GPT-5.5 in a couple of minutes for about $1.50, and Claude Fable 5 churned for 40 minutes and cost $9. All three, in his words, "produced equally acceptable final results." Nine times the price for parity. His operational conclusion about that layer is the argument of Loop Engineering in one line: interactively babysitting an agent on this kind of work is nonsense; you write loops and parallelize instead.

Then he did something no backlog would ever have produced. He let Fable 5 churn for two hours, at a cost of $40, on optimizing a SwiftUI-layout resolver he had written himself in Go. It brought performance down from microsecond to nanosecond scale, an order of magnitude he says he could not reach himself, though he had to claw back some changes the model had overfit to Apple Silicon. His verdict: "Still, very worth it." Notice what that task was not. It was not on any backlog. No product manager prioritized it, no sprint contained it, no roadmap would ever have surfaced it. It became a task only because one of the most respected engineers in the industry suspected his own code had headroom and was allowed to spend $40 finding out.

Put the two halves side by side and the shape of the change appears. On the work everyone already knows how to ask for, the cheap model tied the expensive one, and that is a fact about the task, not about the models. "Implement this feature" is a shape the whole industry has learned to ask for; the prompts are shared, the playbooks are public, and that is precisely where models have converged. When a cheap model ties an expensive one on your backlog, it does not mean the expensive model has no edge. It means your backlog contains no work that needs one. The result its own author could not reach came from the other kind of work: the question that existed nowhere until someone asked it.

So the question that matters is not which model to route to or how much to spend on tokens. It is this:

who in your organization may spend $40 of frontier model time today, on a question that sits on nobody's backlog, without asking permission?

If the honest answer is nobody, or three people with director titles, then your binding constraint is no longer cost, and it is not capability either.

It is imagination, and that is the one input no vendor sells you.

The rest of this article calls this kind of question a $40 question: a question that exists on no backlog, posed to a frontier model at a price that looks large next to a routine call and trivial next to an engineer-week, in search of a result you did not know was possible. Nobody assigns it. A ticket system will never produce it.

And the number is an anecdote, a reference to the experiment, not a threshold: the same question might cost four dollars or four hundred, and nothing in this article asks you to match the figure. What defines the question is that somebody was free to ask it and to spend frontier model tokens to get the best answer.

Two honest caveats before the rest of the article builds on this. It is an informal test, not a controlled benchmark, and Hashimoto's own conclusion is narrower than the use I am making of it: he would reserve frontier models "for targeted, surgical analysis and work," not for daily driving.

The wider generalization, that execution is commoditizing and the differentiator becomes the capacity to pose unbacklogged questions, belongs to Nate B. Jones, whose essay Model Routing Is Table Stakes built the argument on top of that thread. Credit where it is due.

But the pattern did not stay a single anecdote. Sergey Karayev reported in the same thread that he ran his own version: Fable 5, GPT-5.5 and Opus 4.8 in parallel on a real production memory-leak ticket, and Fable 5 was the only one that found the root cause. And the opposite reading was voiced right there too, by a named industry figure. Bindu Reddy argued that "cost per successful task is the only metric in AI that will matter," so just route to cheap.

That is a serious position held by serious people, and I agree 100 percent on the metric; I have put my own thoughts down on token waste and cost per feature. Still, I think it is exactly half right, and the half it misses is where your next advantage lives.

When we went looking across the estate at Betsson, we found AI applications that people had already built for themselves, without asking anyone. The reflex in most companies is to shut that down. We put information security guardrails around it instead, kept the permission to experiment explicit, and moved the ones that worked into production. The questions were already there, sitting in the heads of the people who had context nobody else had. Nobody had to teach them to be curious about their own systems. What the organization actually controlled was whether asking was allowed, and back then asking cost nothing but time. The version of that permission that matters now has a price on it, which is why it needs a budget line and a written boundary rather than a tolerant manager.

The machine answers. It does not ask.

Ask a model anything and it answers. It does not tell you when you are asking the wrong question, and it will not, on its own, bring you the one you should have asked instead. Answering needs compute, instructions and a right-sized model for the task, which is why answering became cheap.

Asking, however, needs things that do not ship with any subscription:

  • context to suspect something more is possible
  • a person who cares what the answer would change
  • the standing to spend money on a suspicion

So the backlog is the ceiling on what your tooling can return: a ten times faster answering machine only reaches the bottom of the same list sooner, a saving every competitor with the same tools is booking too.

It is the electrification story again, motors bolted onto steam-era layouts, decades of real savings before anyone redesigned the building around what the new motors made possible.

And nearly everyone is still bolting.

EY's 2025 survey of 15,000 employees across 29 countries, as Duane Forrester documented, found 88 percent using AI at work and only 5 percent using it in ways that change what they produce.

There are two readings of that 5 percent, and I think both are true: people are not allowed to ask bigger questions, and people cannot yet judge what these models can do. That judgment is a new skill, one Hashimoto himself names in a follow-up to the same thread.

The two fail together.

Permission without capability produces an allowance nobody spends; capability without permission produces a private hobby that never reaches the backlog.

That is why the program below has two parts: first the people who should be spending, then an engine for the questions nobody owns.

One condition before any of this pays: you must be able to check the answer.

The one-day migration of Stripe's 50-million-line Ruby codebase, vendor-reported and unaudited, is really a story about the years of test coverage and review systems Stripe built before the model arrived; pointed at an unprepared codebase, the same model produces 50 million lines of changes nobody can approve. Imagination without verification is not an advantage, it is a liability.

The constraint is real, and a flat cap is the wrong shape for it

Now the pushback I would raise myself, and it is the strongest one.

Organizations are clamping frontier spend, not loosening it. Forbes reported that Microsoft and Uber both blew through their 2026 AI coding budgets within months, and Uber's answer was to cap each employee at $1,500 a month.

What this article asks of a CFO who just lived through that is not more spend and not an open tap. It is a reasonable, ring-fenced experimentation budget with one unusual rule: inside it, your people put questions to your best models without a request form.

The reconciliation is the one Token Economics made for the bill as a whole: a blind cap creates token anxiety and stops spending exactly where spending pays, so caps should be bounded by the value the spend produces, not set as one flat number across every kind of work.

What actually blew those budgets was uncontrolled agentic execution, long autonomous runs and frontier models grinding through routine work a cheap model ties on. Gate that hard; the consoles to do it now ship from every vendor. And that gating is where the money for asking comes from: every trivial task moved off a frontier model onto a right-sized one releases spend, and the experimentation budget is nothing more than a slice of that saving, reinvested.

Framed that way, the CFO is not being asked for new money; the routing policy pays for the questions.

Asking is different. A question is bounded by its own nature, one person, one run, one verifiable answer, so the right control is a budget, not an approval. Set the experimentation money aside in advance, and inside that boundary let people ask with no approval step and no justification ticket, because the moment a question needs a business case it stops being the kind of question this article is about.

What makes a problem worth frontier money

If the answer is not "spend more," it has to be "spend where it changes the outcome." That means being able to tell a frontier problem from a backlog task before the money leaves. Four tests, and in my reading of the evidence a question needs all four.

There is no public playbook for it. This is the first half of Hashimoto's test restated as a filter: on any shape of work the industry already knows how to ask for, the cheap model will tie. Frontier capability shows up on the hard tail: the root cause nobody found, the optimization everyone assumed was at its limit, the combination of two things that no tutorial puts together. Reddy's point holds perfectly for everything else.

You own the deepest context on it. Hashimoto could pose his question because he wrote that layout engine. This is why you cannot hire your way out of the constraint. A newly hired AI visionary arrives with imagination and none of your context, and imagination only fires when it sits next to context. Your context is spread across the people who do the work, which is where the questions have to come from.

You can verify the answer. A nanosecond benchmark, a reproduced memory leak, a test suite that either goes green or does not. This is the Stripe condition. Without it you cannot tell a breakthrough from a confident hallucination, and the run is not an investment, it is a story.

You believe the ceiling is fixed. The best candidates are the beliefs your organization holds without evidence: this service cannot go faster, this migration is too expensive to attempt, this reconciliation has to stay manual. Those beliefs were formed when execution was expensive. Most of them have never been retested.

The negative test matters just as much, because it is where the money burns. If the request is "implement this feature," route it down and do not think about it again. A frontier model doing backlog work is the most expensive way to buy a result you could have had for a dollar, and it teaches the organization exactly the wrong lesson about what the expensive model is for.

First, the field: make the right spend easy

The people closest to your systems already hold questions they have never been allowed to price, so the program starts with them, not with tooling. Give named people the experimentation allowance, funded out of the routing savings.

Money set aside in advance never waits on an approval. Answer the question from the top of this article with actual names, and spread them by context, not seniority.

Then teach people to spend a run the way the model is built to be used, because most engineers still treat a frontier model like fast autocomplete. Anthropic's own guidance for working with Fable-class models reads like a manual for the $40 question. Bring it work you assumed was impossible, not work you already know how to do. Start from an idea and let the model interview you: "before you start, ask me everything you need to know to get this right." Brief it like a colleague, purpose and context over rule lists, and give it something concrete to judge against, a benchmark, a failing test, an earlier draft.

Try answering by voice. In my experience, talking produces a denser, deeper answer and hands the model richer context.

Delegate the outcome rather than the steps, because a step-by-step prompt limits the model to the steps you already thought of. Run it at high effort, let it run long, and when it returns, ask where a number came from before you trust it.

The discipline per run is the asker's own, not an approver's. One line before the money moves: what you believe is impossible, and how you will know if it is not, held against the four tests above. A number on both sides of every run.

Judge the run on what it proved, not on whether the code shipped; Hashimoto clawed back changes overfit to one chip and still called the run very worth it, because what he bought was the knowledge that the headroom existed.

The field is also where the next askers are trained, because you cannot imagine with capabilities you have never touched. The Velocity Trap carried the Stanford finding of a 13 percent relative decline in employment for early-career engineers in AI-exposed roles; read it as a supply problem, because those are the people who would have been posing your $40 questions in 2030. So juniors hold the allowance on the same terms as everyone else, and one item in every cycle of work stays hands-on, chosen for what it teaches, booked against training rather than delivery.

Then, the engine: frontier problem identification

The field practice has a built-in limit: it only surfaces what somebody already suspects. The questions with no owner, the ones that live between teams or under a KPI everyone stopped watching, never make anyone's list. So here is the experiment I would put on the table next. What about a frontier problem identification engine: a set of always-on loops on right-sized models that scan the signals your company already produces and emit candidate frontier questions, never answers. Point one loop at customer service history and ask what single feature would remove the most recurring pain from the customer experience. Point one at site performance telemetry and let it hunt the speed improvements users are silently paying for. Point one at your business KPIs and ask which metric has been flat for years, and what assumption keeps it flat. Point one at the codebase, held against a new target architecture definition, and let it list the refactor candidates a real modernization would start with.

What turns these loops from noise generators into an engine is context, and this is where the knowledge base this series keeps arguing for stops being a nice-to-have. A loop reading complaints without knowing your product, your margins, or your roadmap produces trivia. The same loop reading them against your knowledge base, the ontology of what the product is, what each KPI means, what has already been tried, produces candidates a person with context can recognize in seconds. Your Coding Agents Are Drowning in Context made the negative case, that wrong context is paid for twice; the engine is the positive one, because owned, curated context is exactly what lets cheap models find frontier problems. And the engine itself runs on execution-tier tokens; the frontier budget stays reserved for the survivors.

Loops that discover are not speculation. Andrej Karpathy's autoresearch kept about twenty unbacklogged improvements out of roughly 700 experiments run over about two days, as he described it. Google's Co-Scientist, published in Nature, generates hypotheses for scientists to test, and FutureHouse's Robin proposed the hypothesis behind a candidate treatment for a leading cause of irreversible blindness. All of them stand on crisp verification, which is why the engine only works over streams you can measure.

The governance is the same as the field's, at larger scale. The engine proposes; a person with context qualifies the candidates against the four tests; a frontier model runs the survivors long; every result lands in the same log, dead ends included. The dead ends matter more than they look.

Start small and measure hard. One loop over one stream, for one or two teams, this quarter; a monthly thirty-minute review that promotes winners onto the real backlog, where the proven workflow moves down to cheap models and the money comes back. The measure that proves or kills the whole program inside two quarters is promoted findings: questions that changed what the execution layer was building. Zero after two quarters means shrink the allowance, retune the engine, or accept that your constraint was somewhere else.

Start, Stop, Continue

For executives. Start by answering the $40 question test with names, then widen the list until it reaches the people with the deepest context rather than the highest titles. Stop treating every dollar of frontier spend as waste to be routed away. What wrecked other people's budgets was routine work running on expensive models, which is a different problem with a different fix, and the fix for it should not also gate the asking. Continue the cost discipline and the investment in verification, since a frontier answer you cannot check is worth nothing.

For engineers. Start a personal list of the things you believe are impossible in systems you own, and spend your allowance there rather than on the sprint board. Stop asking frontier models for work a cheap model ties on, and stop waiting for a ticket system to grant permission it was never designed to give. Continue doing some of the work by hand, because the feel you build there is the raw material of your next good question.

The Strategic Takeaway

Execution savings plateau, because every competitor gets the same tools and the same playbook. That layer converges by design. What compounds is different in kind: the stock of questions your people know how to pose, the map of your own jagged frontier that every run leaves behind, the workflows harvested from the answered ones, and the judgment forming in the people who will ask next year's. Marc Andreessen made the point from the investor's chair that models commoditize and costs fall, so the moat is never the model. For operators I would sharpen it: the moat is not your routing policy either. Two companies with identical tools and identical spend will diverge on one variable, which is how many unbacklogged questions they were structurally able to ask.

So run the test on your own organization. Who could have run Hashimoto's experiment this morning, on your own code, without asking anyone? If the honest answer is nobody, you have built a very efficient answering machine and put no one in a position to steer it.

Top comments (0)