<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: code&amp;cosmos</title>
    <description>The latest articles on DEV Community by code&amp;cosmos (@codeandcosmos).</description>
    <link>https://dev.to/codeandcosmos</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4175682%2F3f306c61-37d4-446d-85b5-92ff5c364c49.png</url>
      <title>DEV Community: code&amp;cosmos</title>
      <link>https://dev.to/codeandcosmos</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/codeandcosmos"/>
    <language>en</language>
    <item>
      <title>The Developer's Guide to AI Tokenomics: Why Your Coding Agent Can Burn Through Credits in Minutes</title>
      <dc:creator>code&amp;cosmos</dc:creator>
      <pubDate>Sat, 10 Oct 2026 16:59:34 +0000</pubDate>
      <link>https://dev.to/codeandcosmos/the-developers-guide-to-ai-tokenomics-why-your-coding-agent-can-burn-through-credits-in-minutes-m0m</link>
      <guid>https://dev.to/codeandcosmos/the-developers-guide-to-ai-tokenomics-why-your-coding-agent-can-burn-through-credits-in-minutes-m0m</guid>
      <description>&lt;p&gt;&lt;em&gt;You never came close to your context limit, yet your team's AI credit balance dropped sharply. Here's the real math behind agent loops, hidden context, reasoning tokens and prompt caching.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;A developer or QA automation engineer sits down with a new task. The organization has rolled out AI coding assistants (Cursor, GitHub Copilot, Claude Code, or an internal agent harness) and given everyone a monthly credit quota.&lt;/p&gt;

&lt;p&gt;You ask the assistant to fix a few flaky tests, generate some mock data and debug an integration timeout. Then you grab a coffee.&lt;/p&gt;

&lt;p&gt;When you get back, the usage dashboard shows a big chunk of the month's quota already gone.&lt;/p&gt;

&lt;p&gt;Your first reaction is disbelief: "The UI says my context is only at 45,000 tokens out of 1 million. How did a few minutes of work cost this much?"&lt;/p&gt;

&lt;p&gt;AI pricing isn't broken. Tools have moved from single-turn completions to autonomous agent loops, and most engineers were never shown how token consumption actually works under the hood.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note on numbers:&lt;/strong&gt; Prices, credit conversions and token counts below are illustrative. They vary widely by vendor, plan, model and tool. Check your own provider's pricing page and usage dashboard for real figures.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  1. The core misconception: context window vs. cumulative usage
&lt;/h2&gt;

&lt;p&gt;Keep two different measurements separate:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Context window capacity&lt;/strong&gt; is the size of the desk: how much text the model can hold at one moment.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cumulative token consumption&lt;/strong&gt; is the water meter: everything sent to and generated by the model across every call.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A "1M context window" means the model can hold roughly that many tokens in a single request without failing. You aren't billed for the size of the desk. You're billed every time paper is slid onto it.&lt;/p&gt;

&lt;p&gt;If you load 50,000 tokens of code into a conversation and go back and forth 10 times, the model re-reads that 50,000-token foundation on each turn (prompt caching softens this, covered in section 5).&lt;/p&gt;

&lt;p&gt;The UI may show &lt;code&gt;Current context: 58,000 / 1,000,000&lt;/code&gt;, while the billing meter has already recorded more than half a million input tokens.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. The agent multiplier: one prompt is not one API call
&lt;/h2&gt;

&lt;p&gt;Early chat assistants were simple: you sent a prompt and got an answer.&lt;/p&gt;

&lt;p&gt;Today's coding assistants are agents. When you type something like &lt;em&gt;"Fix the failing payment test,"&lt;/em&gt; the model doesn't answer in one shot. It starts a loop: list files, read some of them, read the implementation, run the test, read the error, edit code, check the diff, run the test again.&lt;/p&gt;

&lt;p&gt;Each step is a separate model call, and each one appends its results to the growing history.&lt;/p&gt;

&lt;h3&gt;
  
  
  The accumulation math (illustrative)
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Step&lt;/th&gt;
&lt;th&gt;What gets added&lt;/th&gt;
&lt;th&gt;Input tokens sent that call&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Base instructions and your prompt&lt;/td&gt;
&lt;td&gt;15,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;File list&lt;/td&gt;
&lt;td&gt;17,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;File contents&lt;/td&gt;
&lt;td&gt;23,500&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Client implementation&lt;/td&gt;
&lt;td&gt;31,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Error log output&lt;/td&gt;
&lt;td&gt;35,500&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;Git diff&lt;/td&gt;
&lt;td&gt;42,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;164,000&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The final context is 42,000 tokens, but 164,000 input tokens were sent across six calls.&lt;/p&gt;

&lt;p&gt;Because each call resends everything before it, total input grows roughly with the &lt;strong&gt;square&lt;/strong&gt; of the number of steps. It isn't exponential, but it climbs fast. If the agent gets stuck in trial and error (edit, fail, read stdout, retry 8 to 12 times), a single innocent prompt can run into several hundred thousand tokens before you see a green check.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Hidden context: what's sent besides your prompt
&lt;/h2&gt;

&lt;p&gt;A 12-word prompt in your IDE is never just 12 tokens. Several things ride along with it.&lt;/p&gt;

&lt;h3&gt;
  
  
  A. System instructions and tool definitions
&lt;/h3&gt;

&lt;p&gt;To act as an agent, the model needs instructions on how to behave and descriptions of every tool it can call: terminal, file search, edit, browser and so on. Depending on the tool, that can run to thousands or even tens of thousands of tokens at the top of every request.&lt;/p&gt;

&lt;h3&gt;
  
  
  B. Editor context
&lt;/h3&gt;

&lt;p&gt;Many IDE assistants pull in extra context to guess what you mean: the current file, recently viewed files, selections, and sometimes snippets from open tabs or search results. How much is included depends heavily on the tool and its settings. It's rarely every open tab in full, but unrelated files can still add meaningful noise.&lt;/p&gt;

&lt;h3&gt;
  
  
  C. Giant terminal dumps
&lt;/h3&gt;

&lt;p&gt;When a test suite crashes, it's tempting to paste the whole terminal output. A full stack trace with warnings, deprecation notices and raw JSON can easily add thousands of tokens in one paste. In an agent loop it gets resent on every later step.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Pricing asymmetry: input, output and reasoning
&lt;/h2&gt;

&lt;p&gt;Tokens aren't all priced the same. The general pattern across major providers:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Token type&lt;/th&gt;
&lt;th&gt;Relative cost (rough pattern)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Cached input&lt;/td&gt;
&lt;td&gt;Cheapest, often a fraction of normal input&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fresh input&lt;/td&gt;
&lt;td&gt;Baseline&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output (generated text)&lt;/td&gt;
&lt;td&gt;Usually several times the input price&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reasoning ("thinking") tokens&lt;/td&gt;
&lt;td&gt;Usually billed as output&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Exact ratios differ by provider and model, so check the current pricing page rather than relying on any single table.&lt;/p&gt;

&lt;h3&gt;
  
  
  The hidden reasoning cost
&lt;/h3&gt;

&lt;p&gt;Reasoning models, and models with extended thinking turned on, generate internal deliberation before they answer. You may never see those tokens in the chat, but they're typically billed at the output rate. A hard architectural question can generate thousands of reasoning tokens on top of the visible answer.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. Prompt caching: the savior and the trap
&lt;/h2&gt;

&lt;p&gt;Resending huge histories would be ruinous, so major providers offer &lt;strong&gt;prompt caching&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  How it works
&lt;/h3&gt;

&lt;p&gt;When a model processes text, it builds internal key-value (KV) representations of the tokens. With caching, the provider stores those for a repeated &lt;strong&gt;prefix&lt;/strong&gt;. If your next request starts with exactly the same tokens, that part can be read from the cache at a large discount instead of being recomputed.&lt;/p&gt;

&lt;h3&gt;
  
  
  The catch: prefix matching
&lt;/h3&gt;

&lt;p&gt;Caching works from the start of the prompt forward, and caches expire after a period of inactivity, often minutes.&lt;/p&gt;

&lt;p&gt;If you change something near the top (edit the system prompt, insert a file above the conversation, flip a setting that changes early context), everything after that point misses the cache and is billed at the full input rate again.&lt;/p&gt;

&lt;p&gt;The practical lesson is to keep stable content at the top, append new content at the bottom, and avoid needlessly reshuffling early context.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. Do &lt;code&gt;SKILL.md&lt;/code&gt; or rules files cost more or less?
&lt;/h2&gt;

&lt;p&gt;Many teams package reusable instructions as &lt;code&gt;SKILL.md&lt;/code&gt;, &lt;code&gt;.cursorrules&lt;/code&gt; or similar rules files.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Turn cost:&lt;/strong&gt; if a tool always attaches a rules file, it adds its size to every request. If the tool loads a skill only when it's relevant, you pay for it only in those turns. Check how your tool behaves.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;System cost:&lt;/strong&gt; modular instructions beat one giant always-on prompt covering every framework, lint rule and convention your company has.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Iteration factor:&lt;/strong&gt; the most expensive tokens come from unnecessary retries. If a short skill file gives the agent the exact test setup it needs, it may succeed on the first pass instead of looping through five failed runs.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A thousand tokens of good instructions up front can save many times that in avoided retries.&lt;/p&gt;




&lt;h2&gt;
  
  
  7. Where tokens go: the heaviest activities in the dev and test lifecycle
&lt;/h2&gt;

&lt;p&gt;Some tasks are naturally cheap. Others turn into long agent loops, huge inputs or heavy reasoning. Here are the usual heavy hitters, roughly grouped by lifecycle stage, with why each one burns tokens and how to rein it in.&lt;/p&gt;

&lt;h3&gt;
  
  
  Planning and design
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Activity&lt;/th&gt;
&lt;th&gt;Why it's expensive&lt;/th&gt;
&lt;th&gt;How to contain it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;"Understand this whole codebase" or architecture walkthroughs&lt;/td&gt;
&lt;td&gt;The agent crawls directory trees and reads many files to build a mental model&lt;/td&gt;
&lt;td&gt;Point it at a few key modules or an existing architecture doc first&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Large refactor or migration planning (framework upgrades, monolith to services)&lt;/td&gt;
&lt;td&gt;Reads many files and often triggers long reasoning&lt;/td&gt;
&lt;td&gt;Plan one module or layer at a time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Requirements-to-design with a reasoning model&lt;/td&gt;
&lt;td&gt;Lots of hidden thinking tokens billed as output&lt;/td&gt;
&lt;td&gt;Use reasoning models only for genuinely hard tradeoffs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Development
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Activity&lt;/th&gt;
&lt;th&gt;Why it's expensive&lt;/th&gt;
&lt;th&gt;How to contain it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Open-ended bug hunting ("why is this broken?")&lt;/td&gt;
&lt;td&gt;Exploration plus repeated run, read, edit loops&lt;/td&gt;
&lt;td&gt;Reproduce first, then give the failing test, file and function&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-file feature generation&lt;/td&gt;
&lt;td&gt;Many reads and edits, and every step resends the growing history&lt;/td&gt;
&lt;td&gt;Break into small, scoped steps with explicit files&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Repo-wide changes (renames, API changes, dependency bumps)&lt;/td&gt;
&lt;td&gt;Touches many files, and each diff and check adds context&lt;/td&gt;
&lt;td&gt;Use scripted tools such as codemods, &lt;code&gt;sed&lt;/code&gt; or IDE refactors for the mechanical part&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fixing build or compile errors in a loop&lt;/td&gt;
&lt;td&gt;Long compiler output gets resent on every retry&lt;/td&gt;
&lt;td&gt;Fix the first error, trim output, cap retries&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long-running chats that mix many tasks&lt;/td&gt;
&lt;td&gt;Stale context from earlier work rides along&lt;/td&gt;
&lt;td&gt;Start a fresh session per task&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Testing and QA
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Activity&lt;/th&gt;
&lt;th&gt;Why it's expensive&lt;/th&gt;
&lt;th&gt;How to contain it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;"Make all the tests pass" loops&lt;/td&gt;
&lt;td&gt;Runs the full suite, reads huge output, edits, reruns, repeats&lt;/td&gt;
&lt;td&gt;Target one failing test at a time (&lt;code&gt;pytest path::test -x -q&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Generating tests for an entire module or service at once&lt;/td&gt;
&lt;td&gt;Reads lots of source and produces lots of output&lt;/td&gt;
&lt;td&gt;Generate per function or class, using a template or example test&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Debugging flaky or intermittent tests&lt;/td&gt;
&lt;td&gt;Many reruns, timing logs and guesswork&lt;/td&gt;
&lt;td&gt;Collect evidence first (seed, logs, failure rate), then ask&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;E2E and UI test debugging (Selenium, Playwright)&lt;/td&gt;
&lt;td&gt;Verbose logs, DOM dumps, screenshots, traces&lt;/td&gt;
&lt;td&gt;Share only the failing step, selector and trimmed error&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Large mock or test data generation&lt;/td&gt;
&lt;td&gt;Output tokens are the priciest kind&lt;/td&gt;
&lt;td&gt;Ask for a small sample plus a generator script&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Coverage-gap analysis over the whole repo&lt;/td&gt;
&lt;td&gt;Reads coverage reports and many source files&lt;/td&gt;
&lt;td&gt;Run the coverage tool yourself and paste only the uncovered lines&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Analyzing big CI logs&lt;/td&gt;
&lt;td&gt;Thousands of log lines per paste or read&lt;/td&gt;
&lt;td&gt;Filter to the failing job and step, using &lt;code&gt;grep&lt;/code&gt; or &lt;code&gt;tail&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LLM or RAG evaluation runs (for example LLM-as-judge scoring)&lt;/td&gt;
&lt;td&gt;Every test case is one or more model calls, so runs multiply quickly&lt;/td&gt;
&lt;td&gt;Start with a small dataset, cache results, and use cheaper judge models where possible&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Code review, security and maintenance
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Activity&lt;/th&gt;
&lt;th&gt;Why it's expensive&lt;/th&gt;
&lt;th&gt;How to contain it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Reviewing large PRs or full diffs&lt;/td&gt;
&lt;td&gt;Big diffs plus surrounding context&lt;/td&gt;
&lt;td&gt;Review file by file, or focus on risky areas&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Security or dependency audits across the repo&lt;/td&gt;
&lt;td&gt;Scans many files and lockfiles&lt;/td&gt;
&lt;td&gt;Run SAST and dependency scanners first, then ask the AI about specific findings&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Writing docs for a whole project&lt;/td&gt;
&lt;td&gt;Reads lots of code and outputs long text&lt;/td&gt;
&lt;td&gt;Document one module at a time from its public interface&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Incident or root-cause analysis on production logs&lt;/td&gt;
&lt;td&gt;Huge log volumes and repeated hypothesis loops&lt;/td&gt;
&lt;td&gt;Pre-filter by time window, service and error ID&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  The common pattern
&lt;/h3&gt;

&lt;p&gt;Almost every expensive activity combines one or more of these:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Wide input:&lt;/strong&gt; many files, big logs, full diffs or whole reports.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Many iterations:&lt;/strong&gt; retry loops, repeated test runs, trial and error.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Large output:&lt;/strong&gt; bulk tests, big datasets, long docs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Heavy reasoning:&lt;/strong&gt; hard problems sent to thinking models.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Cut any one of the four and costs drop. Cut two and they fall dramatically.&lt;/p&gt;




&lt;h2&gt;
  
  
  8. The developer and tester playbook: 5 rules for AI cost control
&lt;/h2&gt;

&lt;p&gt;You don't need a FinOps team to keep AI costs sane. You need some basic prompt hygiene.&lt;/p&gt;

&lt;h3&gt;
  
  
  Rule 1: Start fresh for each discrete task
&lt;/h3&gt;

&lt;p&gt;Long threads accumulate stale file dumps, old test output and dead ends. Some tools summarize or trim long threads automatically, but carried-over noise still costs tokens and can confuse the model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The rule:&lt;/strong&gt; start a new session for each separate ticket, bug or subtask. Don't drag debugging baggage into feature work.&lt;/p&gt;

&lt;h3&gt;
  
  
  Rule 2: Control what context goes in
&lt;/h3&gt;

&lt;p&gt;Closing unrelated tabs can help in editors that pull from open files, but explicit context matters more.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The rule:&lt;/strong&gt; attach exactly the files you want (for example with @-mentions) and state the scope: &lt;em&gt;"Only change &lt;code&gt;billing_service.py&lt;/code&gt; and &lt;code&gt;test_billing.py&lt;/code&gt;."&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Rule 3: Use explicit pointers instead of open-ended crawling
&lt;/h3&gt;

&lt;p&gt;Avoid prompts like &lt;em&gt;"Find out why the tests are failing and fix them."&lt;/em&gt; The agent has to explore directory trees, configs and many files, spending tokens the whole way.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Better:&lt;/strong&gt; run the test yourself, isolate the failure, and point to it:&lt;br&gt;
&lt;em&gt;"&lt;code&gt;test_refund_rounding&lt;/code&gt; in &lt;code&gt;tests/test_billing.py&lt;/code&gt; fails with this assertion. The logic is in &lt;code&gt;billing/refunds.py&lt;/code&gt;, &lt;code&gt;calculate_refund()&lt;/code&gt;. Fix the rounding."&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Rule 4: Right-size the model
&lt;/h3&gt;

&lt;p&gt;Don't use the heaviest reasoning model for mock JSON fixtures, regex or unit-test boilerplate.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Use fast, cheaper models for routine scaffolding, syntax and boilerplate.&lt;/li&gt;
&lt;li&gt;Save frontier reasoning models for multi-file bugs, concurrency issues and complex refactors.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Rule 5: Trim logs and output
&lt;/h3&gt;

&lt;p&gt;Don't paste 400 lines of terminal output. Keep the failing assertion, the few most relevant stack frames and the input that triggered it. When an agent runs commands itself, ask for quiet output, for example &lt;code&gt;pytest -q&lt;/code&gt; or &lt;code&gt;pytest -x --tb=short&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bonus:&lt;/strong&gt; cap retries. If the agent has failed the same fix three times, stop it, look yourself, and give it a sharper pointer.&lt;/p&gt;




&lt;h2&gt;
  
  
  Summary
&lt;/h2&gt;

&lt;p&gt;AI tools don't bill by the hour or by how clever the answer is. They bill by &lt;strong&gt;volume and iteration&lt;/strong&gt;: every input token, every generated token, every reasoning token, and every step of every loop.&lt;/p&gt;

&lt;p&gt;Treat an assistant like an all-knowing oracle and let it wander your repo, and agent loops will multiply your usage fast. Scope the task, point to exact files, trim noisy output, choose the right model, and let caching work for you, and you can keep the same engineering speed at a fraction of the cost.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://medium.com/@codecosmos/the-developers-guide-to-ai-tokenomics-why-your-coding-agent-can-burn-through-credits-in-minutes-47bc5b89388c" rel="noopener noreferrer"&gt;Medium&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>llm</category>
      <category>productivity</category>
    </item>
  </channel>
</rss>
