DEV Community

Cover image for Hourly billing isn't dead if you can prove the hours
Jose Daniel Leon Ruiz
Jose Daniel Leon Ruiz

Posted on

Hourly billing isn't dead if you can prove the hours

There's a good debate going on about how freelancers should price work now that an AI agent does part of it. Mark Fulton's Freelance Developer Rates in 2026 makes the case for moving away from hourly billing, and a lot of it is right: you shouldn't invoice hours you no longer spend, and review and judgment are where the value is now.
But I think hourly billing is getting blamed for a different problem. The problem isn't the unit. It's that almost nobody actually knows their hours.

The timer was always the weak part

I bill clients by the hour. For years my "time tracking" was a timer I forgot to start, followed by a Friday afternoon reconstructing the week from memory and git log. Everyone I know who bills hourly does some version of this. The numbers that end up on the invoice are, charitably, estimates.
That was already shaky before AI. With an agent in the loop it gets worse, because the work doesn't look like typing anymore. It's writing a prompt, reading what comes back, running it, pushing back, reviewing a diff. None of that shows up as keystrokes, and it's exactly the "review and judgment" part everyone agrees is valuable.

The agent is already keeping the log

Here's the part that surprised me: the agent is a better timekeeper than I ever was.
Claude Code writes every session to ~/.claude/projects/. Codex writes to ~/.codex/sessions/. Each turn has a timestamp, the model, and the exact token counts. Your git history says what shipped and when. Put those together and you can rebuild where the time went, per project, without anyone starting a timer.
So instead of "stop billing hours", the option I'd add is: bill hours you can prove.

What "provable" means in practice

When I started turning those logs into timesheets, three things turned out to matter more than I expected.
1. Keep measured and estimated apart. Time rebuilt from agent sessions is measured: it has timestamps. Time inferred from commits alone (code written by hand, or with a tool that doesn't leave session logs) is an estimate. Those two should never be added into one number without a label. A client can accept an estimate; what breaks trust is finding out later that a "measured" number was a guess. And when you estimate, err on the short side, because someone is paying for it.
2. AI cost is a real line item, and it's easy to get wrong. If you're on a flat subscription, your real cost per project isn't the sum of tokens at list price, it's the fee split by how much each project used. And the raw logs have traps. In Claude Code transcripts the same turn can appear up to five times with identical usage, so summing rows inflates cost about 2.9x. In Codex, each event carries both a per-call and a running total, and summing the running totals gave me 6x on a real session. Codex also counts cache inside the input tokens, the opposite of Anthropic's format: not subtracting it turned $2.17 into $10.03. If you're going to pass AI cost through to a client, or just decide whether a project is profitable, those multipliers matter.
3. Let the client see the evidence. "I worked 42 hours" invites a negotiation. "Here are the 42 hours, and each block links to the commits it produced" mostly doesn't. The most useful thing I built was a read-only view I can send a client, where every number drills down to the work behind it.

So, hourly or value-based?

Honestly, both, depending on the client. Fixed-scope and outcome pricing are great when the scope is clear. But a lot of real client work is open-ended maintenance, "can you also look at…" requests, and retainers, and for those hours are still the fairest unit anyone has found. What changed is that we can finally measure them instead of reconstructing them on Friday.
And even if you move to value pricing, you still want the numbers. Knowing that a project took 30 measured hours and $40 of AI is how you find out whether your fixed price was a good idea.

What I built

I turned this into an open-source CLI called Estela. It reads your Claude Code, Codex and GitHub Copilot sessions plus your git history, and rebuilds your hours and AI cost per project and client. It runs entirely on your machine: no account, zero runtime dependencies, MIT licensed, and it never stores the content of your prompts.
npx estela setup
It marks measured vs estimated hours, splits a flat subscription across projects, can either absorb the AI cost or pass it through to the client's report, and publishes a read-only panel you can share. Code and docs: https://dub.sh/Qy9tqca
If you bill by the hour, I'd really like to know how you handle this today, and whether "measured vs estimated" is how you'd want to show time to a client.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.