DEV Community

Cover image for From T-Shirt Sizes to Token Quotes?
Remo H. Jansen
Remo H. Jansen

Posted on

From T-Shirt Sizes to Token Quotes?

I'll start with my position, because the rest of the article depends on it.

I'm against estimation. I have been for a long time. My team has already stopped estimating and stopped doing sprint planning, and I have no plans to bring either back. One thing is left: when upper management asks how big something is, we still answer with a t-shirt size.

Even that has started to feel wrong.

In an earlier article I argued that coding agents are finally killing Scrum. In another I argued that when code gets cheap, verification gets expensive. This article follows from both. If teams no longer estimate, and producing code is no longer the constraint, what should we say when the organisation still asks "how big is this?"

I'm not going to try to bring estimation back to teams. My question is narrower: if an organisation still needs a size signal, what should that signal be?

My hypothesis is that token quotes could replace t-shirt sizes.

I want to be careful about how much confidence that deserves. This is a hypothesis, not a finding. I don't have a dataset of token quotes compared with actual results, and I haven't run the experiment. Some of what follows rests on arguments I've made before and still stand by, mainly that verification has become the bottleneck. Some of it is speculation about a metric nobody has tested yet, including me. I'll try to keep the two apart, and I've included a section on how to test the idea and what would show it to be wrong. Any numbers in this article are illustrative.

I should also be clear about context. I'm describing a team that does most of its implementation work with coding agents, in an organisation that tolerates working without sprint commitments. Plenty of teams don't work like that, and I come back to where this does and doesn't apply near the end.

Estimation never really worked

We all know the problems.

The planning fallacy says we consistently underestimate how long things will take, even when we know about the planning fallacy. Hofstadter's Law says it always takes longer than you expect, even when you take Hofstadter's Law into account. Decades of software projects have confirmed both.

Story points and t-shirt sizes were meant to fix this by moving away from time. Size work relative to other work, the reasoning went, and stop pretending you can predict hours. Then velocity came along and turned points back into time. "We do 30 points a sprint, so this epic will take four sprints." Relative estimation became absolute estimation again with an extra step.

The #NoEstimates movement was right in principle. But as I argued in the Scrum article, organisations lean toward practices they can follow over principles they have to internalise. "Don't estimate" is a principle. Organisations still need something to put in the box, so the box got filled with points and t-shirts.

It's worth asking what management actually needs from that box. Sometimes it really is a date: a contract, a regulatory deadline, a launch tied to a marketing campaign. This article has nothing useful to say about those cases. But in my experience, most of the time the request for a date is standing in for something else. What they need is a relative size signal. Is this bigger or smaller than that? By roughly how much? Which items are outliers we should discuss before committing? They use it to prioritise, to sequence, and to notice when something is much larger than expected.

A size signal is a reasonable thing to want. The problem is how we've been producing it.

Why t-shirt sizes feel wrong now

A t-shirt size is a human's intuition about how much human effort a piece of work will take. Someone who knows the codebase looks at a ticket, compares it with past work, and says "that's an M."

That made sense when human implementation effort was what drove delivery. It no longer is.

With coding agents, implementation time collapses, and it also gets noisy. The same ticket might take an agent four minutes or forty, depending on how much context it needs to find, how many dead ends it explores, and how clearly the intent was expressed. Wall-clock time stops being a useful signal.

More importantly, the constraint has moved. As I argued in When Code Gets Cheap, Verification Becomes Expensive, producing code and establishing that it is correct are two different activities. Agents have made the first dramatically cheaper. They haven't done the same for the second. Or, in the terms of my TPS article, production now outpaces inspection.

So a t-shirt size is now partly an intuition about a quantity that matters less than it used to. (Not entirely: a good t-shirt size often quietly includes review and risk, and that's a point I'll come back to, because token quotes don't.) It still costs real team time to produce: someone has to read the ticket, think about it, and agree with whoever else is in the room. And it can never be checked. Nobody can measure whether something "really was" an L. There's no feedback loop, so the estimates never get better.

A signal that costs human time, increasingly measures the wrong thing, and can't improve looks like a bad deal. The question is whether there's a better one.

The token quote

A token quote is an estimated range of tokens attached to an issue or ticket before work starts. An agent generates it in a pre-flight step: it reads the ticket, explores the relevant parts of the repository, and returns a range along with its reasoning. No code is written.

The output might look like this (illustrative):

Token quote: 180K (P50) – 420K (P90)
Reasoning: Touches the billing module and two API contracts.
           Billing has strong types and a schema; the notification
           service it integrates with does not, and will need exploration.
Confidence: medium. The acceptance criteria don't specify retry behaviour.
Enter fullscreen mode Exit fullscreen mode

That's all there is to it. The useful part is what this signal is and isn't.

Why tokens and not dollars

When I first thought about this idea, the obvious move was to convert tokens into money. Tokens have a price, so why not report "this ticket will cost about $20"?

I think that would be a mistake.

A dollar figure reads as the cost of the work. If management sees $20 next to a ticket, they'll anchor on $20. They'll forget the engineer who reviews the change, the product owner who validates it, the time spent on integration, coordination and deployment, and the operational cost of running it afterwards. The token cost is a small fraction of the total cost of a change. Showing it in currency invites exactly the wrong conclusion: that software has become almost free to build.

Tokens avoid that trap because they're deliberately abstract. "1.5M tokens" doesn't look like a budget line, and nobody will mistake it for the cost of a feature. But anyone can see that 1.5M tokens is much bigger than 10K tokens. The unit works as a relative size signal without pretending to be anything else.

That is exactly what t-shirt sizes were trying to be. The aim is for token quotes to keep the useful property (relative size) and drop the misleading one (a false sense of precision about cost). Whether they really capture the right relative size is the open question I come back to below.

Why token quotes might be better than t-shirt sizes

They cost the team nothing. An agent produces the quote. There's no meeting, no planning poker, no refinement session. This is the main reason the idea doesn't bother me as someone who is against estimation: no human time goes into it. It's a by-product, not a ceremony.

They can be checked against reality. This is the property t-shirt sizes and story points never had. Issue trackers already record when work starts and when it's done, so every quote can be compared with the ticket's actual cycle time, with nobody filling in a timesheet. For the first time, a size signal can be checked against an outcome that was measured, not one somebody estimated.

They're on a continuous scale. A t-shirt size has five buckets. A token quote can tell an outlier apart from something that is merely large.

If management is attached to the familiar vocabulary, the labels can be kept and derived from token bands instead of gut feel. For example (illustrative thresholds only):

Label Token quote (P50)
S < 50K
M 50K – 250K
L 250K – 1M
XL ≥ 1M

The labels stay. The meeting that produced them goes.

The assumption everything rests on

All of this depends on one assumption I can't yet back up: that a token quote is a meaningful proxy for the size management cares about. Below, I'll pin "size" down as the ticket's end-to-end cycle time.

There are good reasons to doubt it. Tokens measure how much work the agent did. They don't measure how much work the humans will do. A one-line change to an authentication flow might cost a few thousand tokens and still need a careful security review. A large, mechanical rename across a codebase might cost a million tokens and need almost no review. If management reads token quotes as "how big is this for us", those two cases will mislead them in opposite directions.

A t-shirt size from an experienced engineer usually includes that risk and review effort, even if nobody says so. A token quote doesn't. That's a real loss, and it's the strongest argument against the whole idea.

My guess is that across a typical backlog, token quotes and end-to-end cycle time are correlated well enough to be useful as a relative signal, with the exceptions being exactly the high-risk, low-volume changes that should get human attention anyway. But that is a guess. It's the first thing to test, and if it turns out to be wrong, the token quote is a poor replacement for a t-shirt size.

Calibrating against the right thing

At first, the quote is just the agent's guess. I don't know how good that guess will be. The agent has one advantage over a human gut feel: it actually reads the code it would need to change. It also has clear disadvantages: it lacks the organisational context, history and tacit knowledge an experienced engineer brings, and models are not known for being well calibrated about their own uncertainty. The one thing in its favour that isn't in doubt is that it's free.

The obvious way to improve the guess would be to compare quoted tokens with actual tokens, then eventually train a model on (ticket → actual tokens). That would probably work, in the narrow sense that the model would get good at predicting token consumption. But it would answer the wrong question. A perfectly calibrated prediction of how many tokens an agent will spend tells you nothing about whether token consumption is what management should care about. You'd be validating the measuring instrument, not what it measures.

The comparison that matters is between the token quote and the issue's cycle time: how long the ticket actually took from starting work to being done, including review, rework, waiting and verification. Cycle time is the outcome the size signal is supposed to stand in for. If tokens and cycle time move together, the quote is a useful relative size signal. If they don't, it isn't, however accurately it predicts its own token usage.

There's an apparent irony here: I'm against estimating time, and I'm proposing to validate against time. The difference is that cycle time is measured, not forecast. Nobody is asked to predict it. It's recorded by the tracker as a side effect of the work, and it captures exactly what the token count misses, which is the human side: review, verification, back-and-forth on requirements.

Cycle time is a noisy target, though. It includes queueing, reprioritisation, people on holiday, and tickets left blocked over a weekend. Using it well needs a consistent definition (for example, from "in progress" to "deployed"), and possibly excluding time spent explicitly blocked. Even then, expect a correlation, not a precise prediction.

Actual token usage is still worth recording, but as a diagnostic, not as the validation target. A large gap between quoted and actual tokens tells you something about the agent's pre-flight step. It doesn't tell you whether the quote means anything.

The long-term step, and I want to be clear that this is aspirational, would be to train your own model on historical (token quote + ticket features → cycle time) data. Plausible features include the modules touched, how explicit those modules are, how clear the ticket is, whether a prototype is attached, and which model or agent did the work. At that point the token quote stops being the answer and becomes one input among several for predicting relative size. I don't know how well this would work. It's a direction, not a recipe.

If the approach works at all, it might produce signals that are more useful than the number itself. These are also untested.

Going over the quote is a reason to reassess. When an agent goes well past its P90, something may be wrong: an ambiguous requirement, missing context, or a hidden dependency. In the TPS article I described this kind of signal as an Andon cord: it draws attention to the problem, which isn't the same as shutting the line down. Some tasks legitimately go over, especially exploratory ones where the point is to find out what's there. So a breach should trigger reassessment, not an automatic stop. The agent pauses at a sensible checkpoint and summarises what it has found and why the work is bigger than expected. A human then decides whether to continue with a revised quote, split the ticket, clarify the requirement, or abandon the approach. Tickets that are openly exploratory can be marked as such and given a wider range from the start. If there's a hard cap at all, it should be set well above the P90 as protection against runaway loops, not used as the main control.

Explicitness becomes measurable. In From Rigidity to Explicitness I argued that constraints such as types, schemas and contracts act as context compression. They tell the agent what the system means, so it doesn't have to infer it. Implicit codebases force agents to spend tokens rediscovering intent. Tokens per change, broken down by module, could become an indicator of where a codebase is hard to work in: a kind of tech debt you can put a number on. Interpreting it would need care, though. Some modules are expensive because they're messy, and others because the domain is genuinely complex.

Spec quality becomes visible. Clear tickets, and especially tickets that come with a working prototype, should produce tighter quotes than vague ones. The cost of ambiguity becomes something you can see.

One caveat on comparisons: token counts are not directly comparable across models. Tokenisers differ, and so does how efficiently models reason. Historical comparisons need to record which model produced the numbers, or be normalised to a reference model.

Don't estimate verification. Invest in reducing it.

This is where I want to put most of the weight. It's also the part of the article that doesn't depend on token quotes at all. If the token-quote idea turns out to be wrong, the argument in this section still holds, because it rests on my earlier articles, not on an untested metric.

Once you accept that tokens are only part of the cost of a change, the obvious next step is to estimate the rest. If tokens size the machine side, why not add a verification estimate for the human side: reviewer-hours, risk classes, approval tiers?

I think that would be a mistake. It would bring planning poker back, aimed at a quantity that is even harder to predict than implementation effort. How long it takes to establish that a change is correct depends on things that are hard to know in advance: what the reviewer finds, which edge cases turn up, how the change interacts with everything around it.

More fundamentally, estimating verification only predicts the cost of the bottleneck. Automating verification reduces it.

I want to be precise about how far that goes. Automation doesn't eliminate verification. It takes over the parts that can be made mechanical: type errors, contract violations, regressions in covered behaviour, invalid states the system can't represent. A substantial part of verification can't be handed over that way. Someone still has to judge whether the requirement was the right one, whether the system behaves the way users actually need, whether a change opens a security hole no test was written for, and whether there are operational, legal or product risks that never show up in a pipeline. Those judgements remain human, and in an agentic workflow they become a larger share of the remaining work.

The point of automation isn't to remove humans from verification. It's to stop spending scarce human attention on the mechanical checks, so it's available for the judgements only humans can make.

That difference matters more than it might seem. As I argued in the verification article, implementation cost is mostly paid once, but verification cost is recurrent. We pay it when we build a feature, and again whenever we change it, refactor it, migrate its data, upgrade its dependencies, or let an agent modify it. Any investment that lowers the cost of verification keeps paying off on every future change. Effort spent estimating verification has to be spent again for every ticket, and it doesn't make the next change any cheaper to verify.

So, for teams in a position like mine, I'd suggest that the time that used to go into estimation is better spent on making verification scale. I'm not claiming this is right for every team. If your organisation genuinely depends on forecasts, you may not be able to free that time up. But where you can, this is the trade I'd make. In practice, that means:

  • Making invalid states unrepresentable. Strong types, schemas, database constraints and explicit state machines turn some classes of bugs into things the system cannot express, so those particular mistakes don't need checking again.
  • Contracts at boundaries. Contract tests between services and modules mean integration correctness is checked mechanically rather than rediscovered in review.
  • A fast, tiered verification pipeline. Cheap mistakes should be caught cheaply, by parsers, type checkers and static analysis, before expensive tests or human attention are needed.
  • Prototypes as requirements. As I argued in the Scrum article, a working prototype from the person who understands the problem validates intent before the build starts. That reduces, though it doesn't remove, the "is this what you meant?" verification later.

None of this needs a forecast. All of it makes the mechanical part of verification cheaper and faster, for humans and for agents, and leaves more human attention for the part that can't be automated.

Bigger changes cost more to verify, now and later

There's one relationship between size and verification worth stating explicitly, though it doesn't need its own estimate.

A large change is harder to verify today: there's more to review, more interactions to consider, and more ways for something subtle to slip through. It also leaves a larger surface of behaviour that has to be re-verified whenever related code changes. Size adds to verification cost twice: once now, and again with every future change nearby.

That gives the token quote a second possible use. A large quote isn't a reason to estimate more carefully. It's a prompt to split the work into smaller changes, each of which is cheaper to verify and leaves less behind to re-verify. The quote doesn't need to predict verification cost to be useful here. It only needs to flag size early, while splitting is still cheap, with the caveat from earlier that a small quote doesn't mean a change is low-risk.

How to test this, and what would prove it wrong

Because the idea is untested, the honest way to adopt it is as an experiment, not a policy. Fortunately it's a cheap experiment. Token quotes cost no team time, so they can run alongside existing t-shirt sizes for a few months without changing anything else.

Over that period, I'd want to answer four questions:

  1. Do token quotes track cycle time? Is there a useful correlation between the quote and the ticket's measured cycle time? Do tickets in a higher token band reliably take longer end to end? This is the key assumption from earlier and the question that matters most. If the correlation is weak, token quotes are measuring agent effort and nothing management needs.
  2. Where do they diverge? Look at the tickets where quote and cycle time disagree most. If the outliers are mostly high-risk, review-heavy changes, as I'd expect, that's a known limitation that can be flagged. If they're random, the signal is noise.
  3. Do they agree with experienced engineers? Run t-shirt sizes and token quotes side by side, and check which one correlates better with cycle time. If the t-shirt size usually wins, the human intuition is capturing something the tokens miss.
  4. Do they change decisions? Did management prioritise, sequence or split work any differently because of the quotes? A signal nobody acts on isn't worth producing, however cheap it is.

Notably absent from this list: "does the agent predict its own token usage accurately?" That's worth monitoring, but a quote can be perfectly calibrated against tokens and still useless as a size signal.

If the answers come back mostly negative, the conclusion should be to keep t-shirt sizes, or to drop sizing altogether, rather than to keep adjusting the token quote until it looks useful.

Caveats

There are ways this can go wrong, and they're worth naming.

Goodhart's Law. If teams are judged on how well they match their token quotes, the quotes will stop being informative. They're a sizing signal for management, not a performance metric. The moment they become a target, they become story points again.

Tokens are not cost. I'll repeat this because it's the easiest mistake to make. A token quote says nothing about labour, review, coordination or operations. It shouldn't be converted to currency for reporting, however tempting that is.

Non-determinism. The same ticket, run twice, will use different numbers of tokens. That's why a quote is a range, not a point.

Model drift. New models change token efficiency, sometimes dramatically. Historical baselines need to be revisited when the underlying model changes.

No new ceremony. If token quotes ever turn into a meeting, with people debating whether a ticket is really 200K or 400K, the approach has failed. The whole point is that no human time goes into producing them.

Where this doesn't apply

The argument assumes a particular setting, and it's worth being explicit about where that setting doesn't hold.

  • Real external deadlines. Fixed-price contracts, regulatory dates and coordinated launches need forecasts in time. Token quotes don't provide those, and they aren't meant to.
  • Heavy cross-team dependencies. When several teams have to sequence work against each other, some form of time-based planning is often unavoidable.
  • Teams where agents do little of the work. If most implementation is still done by hand, human effort is still the main driver, and a t-shirt size may well remain the better signal.
  • High-risk domains. Where most changes need careful human review whatever their size, such as safety-critical, financial or security-sensitive systems, the gap between agent effort and total effort is at its widest. Token quotes would be at their least informative exactly where the stakes are highest.

Outside those cases, I think the idea is worth trying. Inside them, I'd be cautious.

Summary

There are two separate claims in this article, and they deserve different levels of confidence.

The first is one I'm fairly confident about, because it builds on arguments I've made before. In agentic development, verification is the constraint. Effort that makes verification cheaper keeps paying off on every future change, and effort spent forecasting it doesn't. Where a team has the freedom to choose, I'd put the time into verification automation rather than estimation.

The second is a hypothesis. If an organisation still needs a size signal, agent-generated token quotes might be a better one than t-shirt sizes. They cost the team nothing, they can be checked against measured cycle time, and expressing them in tokens rather than dollars avoids presenting a partial cost as the whole cost. Whether they actually track the size that matters is an open question, and the way to answer it is to run them alongside existing sizes and look at the data.

Conclusion

In my experience, when management asks for estimates, what it most often needs is a sense of relative size, so it can prioritise and spot outliers. We've met that need with rituals that cost team time, measure human effort, and can never be checked.

Token quotes could meet the same need at no cost to the team, in a unit that can be measured against reality. I think they're worth an experiment. I don't yet know whether they'll survive one.

Either way, they only answer the easy question. The hard question in agentic development isn't "how big is this?" It's "when can we trust it?" That one can't be answered by estimating harder. It's answered by making verification cheaper.

If you need a size, try tokens, and measure whether they work. If you need speed, invest in verification. Neither should require an estimation meeting.

Top comments (2)

Collapse
 
sameerqaisar17 profile image
Sameer Qaiser •

The line that hit me: "Estimating verification only predicts the cost of the bottleneck. Automating verification reduces it."

I'm three weeks into Python — 27 articles published. I don't have a team, I don't estimate tickets, and I don't use tokens in any way that matters. So this is coming from someone with no skin in the game.

But here's what your argument made me think about.

I've been writing about AI and learning for the last few weeks. And I keep running into the same pattern: people use AI to produce more output, but they don't invest in checking whether the output is correct. They measure how fast they can generate, not how much they can trust what they generated.

That's the same mistake as estimating the bottleneck. You're measuring the thing that isn't the constraint.

For me, as a beginner, verification is the whole game. I can generate code in seconds. But I can't tell if the code is right until I run it, read it, and understand it. And that step — the verification step — is where all my learning happens.

Your distinction between "predicting the cost" and "reducing the cost" applies to learning too. I can predict how long it'll take me to learn a new Python concept. Or I can build something that forces me to actually understand it. The first is estimation. The second is investment.

I don't know if token quotes will replace t-shirt sizes. But your framing of the actual problem — verification is the constraint, not implementation — is something I'm going to carry into my own work.

Great post. Saving it.

Collapse
 
baumgaerben profile image
baumgaerben •

Estimation theater is the real tax — not the estimates themselves. Spent years watching teams burn sprint capacity arguing whether something is a 3 or a 5, then missing the deadline anyway because the "2" hid a legacy integration nobody touched in three years.

What actually moved the needle for us: slicing work until each piece fits in 1-2 days max, then forecasting via throughput (items done per week) instead of velocity points. No planning poker, no t-shirt sizes, no Fibonacci debates. Just "can we ship this slice by Thursday?" and historical data doing the forecasting.

The hard part isn't the math — it's teaching product to accept "we'll know more after we ship the first slice" as a valid answer. Once they see working software weekly instead of a Gantt chart that's wrong by month two, the trust builds itself.

Curious if you've tried throughput-based forecasting or if the article goes a different direction — found it via LabAgent, site: labagent .tech