DEV Community

Cover image for Your coding agent plan is a subsidy
Shrawan Saproo for Nasiko Labs

Posted on

Your coding agent plan is a subsidy

The monthly fee tells you what access costs. It doesn’t tell you what your work consumes.

At Nasiko, we’re building around a question every team using agents should be able to answer: what did this work consume, and what did it accomplish? Coding subscriptions can make that connection easy to overlook. The monthly fee is predictable, even when the work behind it changes.

You’re halfway through fixing a bug with your coding agent when you hit the usage limit. It’s made a few changes, but the tests still fail. You’ll have to wait for the reset to continue.

You know there’s work left, but it’s harder to tell what used up the allowance. The agent spent time reading the codebase, running tests, and trying fixes. The conversation got longer as it went. Without a breakdown, you’re left guessing which of those contributed most.

The monthly bill doesn’t help. It’s the same amount you paid last month, regardless of how the work went.


That predictability makes experimentation easier. We can explore an unfamiliar codebase without pricing every question. But it also separates the work from its consumption.

A subscription prices access under a particular set of rules. It does not tell us what the same work will cost under a different billing arrangement.

To understand that exposure, we need to connect consumption to the tasks that caused it. That connection matters while we are waiting for a reset. It matters more when we deploy the workflow, expand its use, or have to explain its budget.

In June 2026, SemiAnalysis bought subscriptions across OpenAI and Anthropic’s plan tiers and ran long coding tasks until they exhausted their weekly limits. It estimated maximum monthly API-equivalent usage of roughly $14,000 for the $200 ChatGPT tier and $8,000 for the $200 Claude tier. SemiAnalysis’s methodology and reported results.


Those figures describe consumption valued at API retail prices. They are not vendor costs.

SemiAnalysis modeled subscription gross margins assuming a 75% API gross margin. Its published break-even thresholds were approximately 5.7% utilization for the top OpenAI tier and 10% for the top Anthropic tier. Those thresholds concern modeled subscription gross margin, not either company’s overall profitability. Reported margin estimates.

The arithmetic reproduces the thresholds: at that assumption, $14,000 in API-priced consumption implies $3,500 in serving costs, and $200 divided by $3,500 is about 5.7%. The Claude example implies $2,000 in serving costs and a 10% threshold.

These are stress-test estimates under a stated margin assumption, not typical customer bills or guaranteed allowances. Their usefulness is the distance they expose between the subscription price and the economics of metered access.

For a team making a budget, that distance is a question to investigate. Our own workload might sit well below those extremes. It might also change substantially once a successful experiment becomes something we run every day.


Caching explains why the comparison needs care.

A coding agent repeatedly sends overlapping context: instructions, tool definitions, repository content, and conversation history. Prompt caching lets the provider reuse processing of matching prefixes.

Requesty’s analysis of production gateway traffic reported a 92% input-token cache hit rate for Claude Code in April 2026. That observation belongs to its dataset, but the consequence is clear: much of the input in those workloads was reused context, making the raw token count a poor guide to processing cost. Requesty’s research.

A subscription can therefore look generous partly because the workload is efficient to serve. Under the subscription, included cache reads do not create separate charges on our bill. Under API pricing, reads are discounted but still metered, while cache writes, new input, and output are billed at their own rates. Anthropic’s caching documentation.

A useful comparison preserves those categories. Pricing every token as uncached input inflates the estimate. Treating every cached token as free understates it.

Heavy interactive users can still get excellent value from subscriptions. Caching can also keep an API workload affordable. The efficiency may survive a move; the subscription entitlement depends on the destination's rules.

file-1

We encounter that boundary in ordinary engineering work.

An experiment becomes a service. Instead of supervising each run, we schedule the agent to process a queue. Jobs overlap, retries accumulate, and sessions begin with different context.

A team standardizes its tooling. Its enterprise agreement has different allowances or metering from the individual plans developers used during evaluation.

Or a provider changes what access includes.


Anthropic’s proposed Agent SDK billing split shows how unsettled these boundaries remain. The company announced a separate credit arrangement for SDK and claude -p usage, then paused it on June 15. Its current notice states that the affected usage continues to draw from subscription limits and that the announced separate credit is unavailable. The proposed split should not be treated as an implemented change. Anthropic’s current notice.

GitHub provides an implemented example. Usage-based AI Credits billing went live on June 1, with included allowances and controls for additional spending. Organizations can cap spending. GitHub’s release announcement.

The transition included temporary monthly allowances of $30 for Business and $70 for Enterprise during June, July, and August, compared with the announced standard allowances of $19 and $39. With that promotional period over, teams need to distinguish changes in consumption from changes in what their plan covers. GitHub’s billing announcement.

We do not need a prediction about the end of subscriptions. We already have a practical reason to understand our workloads independently of them.


The organizational consequences extend beyond the invoice.

In May, Windows Central, citing The Verge, reported that Microsoft’s Experiences + Devices division was expected to move from Claude Code to GitHub Copilot CLI by the end of June. The reporting described financial considerations alongside Microsoft’s interest in shaping its own tooling around internal repositories, security requirements, and workflows. Windows Central’s report.

That example shows how developer preference sits alongside budget, integration, and governance in tool selection. But the same problem appears in much smaller teams.


Imagine a team introducing automated pull-request reviews. Spending rises. The lead can see usage by developer, but cannot tell whether the increase came from broader review coverage, repeated attempts on the same failing job, or longer context carried into every request.

A blanket cap would control the total while leaving those questions unanswered. It might interrupt useful reviews and preserve the wasteful ones. Before changing the budget, the team needs to identify what the additional consumption accomplished.

That makes attribution as much an engineering concern as a finance concern.
A billing export can identify spending by account, project, key, and sometimes user or model. Those dimensions help allocate expenses. Explaining them requires execution context.

A shared key might serve a migration assistant, an incident investigation, and a nightly review job. A personal key might span several repositories. A single session might contain productive work followed by a long, unsuccessful detour.


The unit we pay for and the unit we make decisions about are different. Providers meter requests and tokens. Engineering teams decide whether a migration, investigation, or review was worth doing.

Consider a retry loop. The provider can accurately report the tokens consumed for each attempt. Execution context reveals that the attempts belonged to the same unresolved task. A task identifier, linked requests, and completion status let us distinguish an expensive success from an expensive failure.

Tool calls require similar care. A local test command may not have a model charge itself, but its output increases the charge of the next model request. Recording the sequence helps explain the cost without pretending that every tool invocation has an independently measurable token bill.

file-kv


Existing telemetry supplies some of this evidence. Claude Code supports OpenTelemetry exports for usage, costs, tool activity, and sessions. The work is to collect it consistently and connect it to the repositories, tasks, owners, and outcomes our organization understands. Claude Code’s monitoring documentation.

That gives us a more productive conversation with decision-makers. We can show which workflows consumed more, whether they completed successfully, and where people still had to intervene. We can investigate repeated failures without treating heavy usage as evidence of poor judgment.

Token counts alone cannot establish value. An expensive session that resolves a production incident may be worth far more than dozens of inexpensive sessions whose changes are discarded. Measurement makes that tradeoff visible. It does not make the judgment for us.

We can begin with three concrete steps this week.

  1. Price our own last week at published API rates. Export the available usage records. Separate uncached input, cache writes, cache reads, and output by model, then apply the corresponding rates. Label the result “API-equivalent estimate” and keep actual subscription payments alongside it. If records are incomplete, identify the gap and start collecting them now.

  2. Connect consumption to sessions and tasks. Record session and request identifiers, repository, task, model, token categories, retries, and tool events. Preserve parent-child relationships when agents delegate work. Start with one workflow and verify that a costly run can be traced back to something an engineer recognizes. Reconcile the records against provider totals before using them for budget decisions.

  3. Price the path out of the current plan. Identify which workflows would require API billing under their intended deployment. Estimate costs using observed cache behavior, expected volume, concurrency, and retries. Pair the estimate with an outcome measure, such as accepted changes or incidents resolved. That gives us a basis for deciding what is worth funding when the billing arrangement changes.


We built Nasiko to bring agent execution and consumption into the same view, so teams can investigate the runs behind their spending. Our open-source platform records execution traces and collects token usage and costs from telemetry, with session reporting for supported coding agents. The Nasiko repository includes the telemetry setup and a Docker quickstart that gets the platform running in minutes once prerequisites are in place.

The next time a session stops at a reset window, we may still choose to wait. A subscription may remain the best deal for the way we work.

But we should be able to explain what happened before the interruption: which task consumed the allowance, what it accomplished, and what comparable metered execution would cost.

The subscription’s limits, eligibility, and future terms are determined by the provider.

Our record of the work should be ours.

Top comments (0)