<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: zhx842htt</title>
    <description>The latest articles on DEV Community by zhx842htt (@zhx842htt).</description>
    <link>https://dev.to/zhx842htt</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4159528%2F098e42c8-cbea-404d-9952-c3ebe22cb89e.png</url>
      <title>DEV Community: zhx842htt</title>
      <link>https://dev.to/zhx842htt</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/zhx842htt"/>
    <language>en</language>
    <item>
      <title>How I Track My AI Coding Costs Per Project (and Bill Them to Clients)</title>
      <dc:creator>zhx842htt</dc:creator>
      <pubDate>Sat, 03 Oct 2026 10:37:04 +0000</pubDate>
      <link>https://dev.to/zhx842htt/how-i-track-my-ai-coding-costs-per-project-and-bill-them-to-clients-1m1c</link>
      <guid>https://dev.to/zhx842htt/how-i-track-my-ai-coding-costs-per-project-and-bill-them-to-clients-1m1c</guid>
      <description>&lt;p&gt;If you deliver software with Claude Code (or Cursor, or raw API keys), you probably know roughly what you spend on AI each month. What you almost certainly &lt;em&gt;can't&lt;/em&gt; answer is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Which project — or which &lt;strong&gt;client&lt;/strong&gt; — consumed that spend?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That question matters if any of your work is billed. Hours and licenses show up on invoices; AI costs currently hide inside a terminal. This post shows where the data actually lives, how to turn it into per-project dollars, and the tooling I built to make it one command.&lt;/p&gt;

&lt;h2&gt;
  
  
  The data was on your machine all along
&lt;/h2&gt;

&lt;p&gt;Claude Code writes every session to local JSONL files:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;~/.claude/projects/&amp;lt;project-slug&amp;gt;/&amp;lt;session-id&amp;gt;.jsonl
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each line is an event. Assistant messages carry the interesting part — a &lt;code&gt;message.usage&lt;/code&gt; object with token counts, the model name, plus the session's &lt;code&gt;cwd&lt;/code&gt; (which is how you know &lt;em&gt;which project&lt;/em&gt; the tokens belong to):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"assistant"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"timestamp"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-09-20T10:00:00.000Z"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"sessionId"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"session-a"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"cwd"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"/home/dev/clientA-api"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"message"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"claude-sonnet-4-20250514"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"usage"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"input_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"output_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;800&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"cache_creation_input_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;5000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"cache_read_input_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;20000&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Group by &lt;code&gt;cwd&lt;/code&gt;, sum tokens, multiply by prices — done? Almost. Three gotchas first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gotcha 1: naive token × price is wrong
&lt;/h2&gt;

&lt;p&gt;Most of the tokens in an agentic coding session never hit list input price. The big three buckets:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Bucket&lt;/th&gt;
&lt;th&gt;Typical price vs input&lt;/th&gt;
&lt;th&gt;Why it dominates&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;input_tokens&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1×&lt;/td&gt;
&lt;td&gt;the actual new context&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;cache_creation_input_tokens&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;~1.25×&lt;/td&gt;
&lt;td&gt;writing context to cache&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;cache_read_input_tokens&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;~0.1×&lt;/td&gt;
&lt;td&gt;re-reading cached context — &lt;strong&gt;the majority in long sessions&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Ignore the cache split and your estimate can be off by several times. The formula per message:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;billable = input×1 + cache_write×1.25 + cache_read×0.1   (× input price)
         + output×output_price
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Gotcha 2: duplicate lines
&lt;/h2&gt;

&lt;p&gt;Streaming writes can emit the same assistant message more than once. Dedupe on &lt;code&gt;message.id&lt;/code&gt; (unique per message) — or your costs inflate silently. A fingerprint of &lt;code&gt;timestamp + sessionId + token counts&lt;/code&gt; works for older formats without ids.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gotcha 3: prices drift
&lt;/h2&gt;

&lt;p&gt;Model pricing changes often enough that hardcoding numbers in a script you'll forget about is a trap. Keep a price table in a config file (&lt;code&gt;~/.devspend/prices.json&lt;/code&gt; overrides the built-ins) and treat every output as an &lt;em&gt;estimate&lt;/em&gt; until you calibrate it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Putting it together: one command
&lt;/h2&gt;

&lt;p&gt;I packed the above into a small zero-dependency CLI, &lt;strong&gt;devspend&lt;/strong&gt; (MIT, open source — I'm the author, to be clear):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx devspend projects
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;项目路径                          花费        输入tok       输出tok
/clients/acme-rewrite           $142.80      31.2M         1.8M
/clients/globex-mvp              $96.15      19.7M         1.1M
/side-projects/devspend          $12.40       2.1M         0.3M
─────────────────────────────────────────────────────────────
合计                             $251.35
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It reads only local files, sends nothing anywhere, and also does day/model breakdowns (&lt;code&gt;devspend report&lt;/code&gt;, &lt;code&gt;devspend models&lt;/code&gt;). The full source is short enough to audit in one sitting — which is the point; billing numbers should be verifiable.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;GitHub: &lt;a href="https://github.com/zhx842htt/devspend" rel="noopener noreferrer"&gt;zhx842htt/devspend&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Waitlist for the billing-grade features: &lt;a href="https://devspend.netlify.app" rel="noopener noreferrer"&gt;devspend.netlify.app&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The part that actually pays: billing
&lt;/h2&gt;

&lt;p&gt;Once spend is attributed per project, two defensible billing models emerge:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Rate-card billing (API users):&lt;/strong&gt; your token costs at list prices become a line item, like billable hours. Estimate-based, transparent, easy to defend.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost-pool allocation (subscription users):&lt;/strong&gt; your Claude Max / Cursor subscription is a fixed pool; allocate it across clients proportionally to usage. Arguably &lt;em&gt;more&lt;/em&gt; honest than rate-card, since it matches your real cash outflow.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Either way the win is the same: a cost that used to silently eat your margin becomes visible, discussable, and recoverable. The first freelancer conversation I'd love to hear more about is simply: &lt;em&gt;do your clients accept AI as a billable line today, or is everyone still rolling it into the rate?&lt;/em&gt; Genuine question — comments open.&lt;/p&gt;

&lt;h2&gt;
  
  
  Roadmap
&lt;/h2&gt;

&lt;p&gt;API-key usage ingest (OpenRouter/OpenAI/Anthropic) and subscription cost pools are next; client entities and invoice-grade exports after that. The CLI stays free and open — the billing workflow is what becomes Pro.&lt;/p&gt;

&lt;p&gt;If you've built your own tracking scripts, I'd genuinely like to compare notes.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devtools</category>
      <category>productivity</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
