DEV Community

Cover image for Cost per Task, Not Cost per Token: Build an LLM Cost Calculator in TypeScript
Bobby Hall Jr
Bobby Hall Jr

Posted on

Cost per Task, Not Cost per Token: Build an LLM Cost Calculator in TypeScript

Most LLM pricing pages answer one question.

What does a million tokens cost?

That's a useful number.

It's just not the number that shows up on the bill.

The bill is made of tasks. A task is a few turns, a few tool calls, a retry or two, and a prompt that gets resent every single time.

This month, the model launches started saying that out loud.

  • On September 21, SpaceXAI released Grok 4.7, "served at the same price and speed as Grok 4.6." Its docs list $2 input and $6 output per 1M tokens.
  • On September 22, Anthropic released Claude Opus 5.5. Its token prices are 20% lower than Opus 5. But Anthropic's headline number is different: "at default settings it will cost 40% less than Opus 5 on typical workloads."
  • On September 28, Anthropic released Claude Sonnet 5.5 at the same $2 / $10 as Sonnet 5, and wrote: "In our testing, it costs up to 30% less per task than its predecessor."
  • On September 29, at DevDay, OpenAI released GPT-6.1 Sol. OpenAI's pricing page lists it at $2 input and $10 output, a fifth of GPT-6 Astra's $10 / $50. The same day, GitHub added it to Copilot, saying it completed tasks "while using noticeably fewer tokens and steps."

Different companies.

Same pitch.

The price per token barely moved. The claim moved to the task.

That distinction matters, because it cuts both ways.

Artificial Analysis, a third party, compares Grok 4.7 (high) with Grok 4.6 (medium). Same token prices. It lists $2.73 versus $1.50 per task, because Grok 4.7 (high) used 66k output tokens per task and Grok 4.6 (medium) used 28k. Those are different effort settings, so it isn't a like-for-like model comparison. It is a clean example of the same price producing a different bill.

So let's stop comparing price lists and build a tiny cost-per-task calculator.

By the end, you will run one command:

npx tsx cost.ts
Enter fullscreen mode Exit fullscreen mode

No API key.

No SDK.

Just list prices, a task model, and arithmetic you can check.

Table of Contents

  1. Project Setup
  2. Step 1: A Price Table With Sources
  3. Step 2: Price One Turn
  4. Step 3: Model a Task
  5. Step 4: Cache Hit Rate and Latency
  6. Step 5: Compare Models per Task
  7. Step 6: Find the Break Even Point
  8. Step 7: Prompt Layout vs Model Choice
  9. Step 8: Look at the Tail
  10. Where It Breaks Down
  11. The Bigger Idea

Project Setup

You will need Node.js 18 or newer.

mkdir cost-per-task
cd cost-per-task

npm init -y
npm install --save-dev typescript tsx @types/node
Enter fullscreen mode Exit fullscreen mode

Save the following blocks, in order, as cost.ts.

Step 1: A Price Table With Sources

type Price = {
  input: number; // $ per 1M uncached input tokens
  cachedInput: number; // $ per 1M cache-hit input tokens
  cacheWrite: number | null; // $ per 1M tokens written to cache (null = not listed)
  output: number; // $ per 1M output tokens
  source: string;
  checked: string;
};

// Standard tier, short context. Checked by hand on Sep 30, 2026.
const PRICES: Record<string, Price> = {
  "claude-sonnet-5-5": {
    input: 2, cachedInput: 0.2, cacheWrite: 2.5, output: 10,
    source: "https://platform.claude.com/docs/en/about-claude/pricing",
    checked: "2026-09-30",
  },
  "claude-opus-5-5": {
    input: 4, cachedInput: 0.2, cacheWrite: 5, output: 20,
    source: "https://platform.claude.com/docs/en/about-claude/pricing",
    checked: "2026-09-30",
  },
  "gpt-6.1-sol": {
    input: 2, cachedInput: 0.1, cacheWrite: 2.5, output: 10,
    source: "https://developers.openai.com/api/docs/pricing",
    checked: "2026-09-30",
  },
  "grok-4.7": {
    input: 2, cachedInput: 0.5, cacheWrite: null, output: 6,
    source: "https://docs.x.ai/docs/models/grok-4.7",
    checked: "2026-09-30",
  },
};
Enter fullscreen mode Exit fullscreen mode

Every row carries a URL and the date I checked it.

Prices change.

A price without a source is a rumor with a decimal point.

Three details are easy to miss:

  • Cache reads are not priced the same everywhere. $0.10 on GPT-6.1 Sol, $0.20 on both Claude models, $0.50 on Grok 4.7.
  • Anthropic and OpenAI list a separate cache write price. The Grok 4.7 docs don't list one, so the code bills those tokens at the normal input rate.
  • These are short-context rates. OpenAI charges more above 272K input tokens and xAI above 200K. Our example stays well under both.

Sourced price table

Step 2: Price One Turn

type Turn = {
  freshInput: number; // input tokens billed at the full rate
  cacheWrite: number; // input tokens stored in the cache this turn
  cacheRead: number; // input tokens served from the cache
  output: number; // output tokens, including reasoning tokens
  latencyMs: number;
  retry: boolean;
};

function turnCost(t: Turn, p: Price): number {
  const writeRate = p.cacheWrite ?? p.input;

  return (
    t.freshInput * p.input +
    t.cacheWrite * writeRate +
    t.cacheRead * p.cachedInput +
    t.output * p.output
  ) / 1_000_000;
}

const taskCost = (turns: Turn[], p: Price) =>
  turns.reduce((sum, t) => sum + turnCost(t, p), 0);
Enter fullscreen mode Exit fullscreen mode

A turn is one model call.

It has four kinds of tokens, and each one has its own rate.

In this calculator, output includes reasoning tokens. Check how your provider reports them. They're usually the part that grows when you turn effort up.

Step 3: Model a Task

This is the important step.

// EXAMPLE INPUTS. Not measurements of any model.
type TaskShape = {
  systemTokens: number;
  taskTokens: number;
  toolTurns: number; // model calls a tool, gets a result back
  toolResultTokens: number;
  outputPerTurn: number;
  retryTurns: number[]; // turn indexes that failed and ran again
  cachePrefix: boolean; // is the start of the prompt stable?
  firstTokenMs: number;
  msPerOutputToken: number;
  toolMs: number;
};

function simulateTask(s: TaskShape): Turn[] {
  const turns: Turn[] = [];
  let context = s.systemTokens + s.taskTokens;
  let cached = 0;

  for (let i = 0; i <= s.toolTurns; i++) {
    const attempts = s.retryTurns.includes(i) ? 2 : 1;

    for (let a = 0; a < attempts; a++) {
      const read = s.cachePrefix ? cached : 0;
      const write = s.cachePrefix ? context - cached : 0;
      const isTool = i < s.toolTurns;

      turns.push({
        freshInput: s.cachePrefix ? 0 : context,
        cacheWrite: write,
        cacheRead: read,
        output: s.outputPerTurn,
        latencyMs:
          s.firstTokenMs +
          s.outputPerTurn * s.msPerOutputToken +
          (isTool ? s.toolMs : 0),
        retry: a > 0,
      });

      if (s.cachePrefix) cached = context;
    }

    context += s.outputPerTurn + s.toolResultTokens;
  }

  return turns;
}
Enter fullscreen mode Exit fullscreen mode

Everything in TaskShape is an example input. I made these numbers up to show the math. They are not measurements of any model.

The key line is this one:

context += s.outputPerTurn + s.toolResultTokens;
Enter fullscreen mode Exit fullscreen mode

Every tool call makes the next prompt longer.

The model re-reads the system prompt, the task, and every earlier result on every turn.

So input cost doesn't grow in a straight line with the number of turns. It grows faster.

A retry sends the same prompt again and pays for the output again.

With a stable prefix, most of that resent context comes from the cache. Without one, you pay full price for all of it, every turn.

How one task turns into a bill

Step 4: Cache Hit Rate and Latency

function cacheHitRate(turns: Turn[]): number {
  let read = 0;
  let total = 0;

  for (const t of turns) {
    read += t.cacheRead;
    total += t.freshInput + t.cacheWrite + t.cacheRead;
  }

  return total === 0 ? 0 : read / total;
}

const timeToDoneMs = (turns: Turn[]) =>
  turns.reduce((sum, t) => sum + t.latencyMs, 0);

function percentile(values: number[], p: number): number {
  const sorted = [...values].sort((a, b) => a - b);
  const index = Math.ceil((p / 100) * sorted.length) - 1;
  return sorted[Math.max(0, index)];
}
Enter fullscreen mode Exit fullscreen mode

Cache hit rate is cached input divided by all input.

timeToDoneMs adds up every turn, including tool time.

That's the latency a user actually feels.

Time to first token is one number inside it.

Step 5: Compare Models per Task

const exampleTask: TaskShape = {
  systemTokens: 6_000,
  taskTokens: 1_500,
  toolTurns: 6,
  toolResultTokens: 2_500,
  outputPerTurn: 800,
  retryTurns: [3],
  cachePrefix: true,
  firstTokenMs: 900,
  msPerOutputToken: 12,
  toolMs: 1_500,
};

const usd = (n: number) => `$${n.toFixed(4)}`;
const turns = simulateTask(exampleTask);

console.log("Step 5: same example task, every model");
console.log(`  turns: ${turns.length} (${turns.filter((t) => t.retry).length} retry)`);
console.log(`  cache hit rate: ${(cacheHitRate(turns) * 100).toFixed(1)}%`);

for (const [model, price] of Object.entries(PRICES)) {
  console.log(`  ${model.padEnd(18)} ${usd(taskCost(turns, price))} per task`);
}
Enter fullscreen mode Exit fullscreen mode

Output:

Step 5: same example task, every model
  turns: 8 (1 retry)
  cache hit rate: 80.4%
  claude-sonnet-5-5  $0.1546 per task
  claude-opus-5-5    $0.2869 per task
  gpt-6.1-sol        $0.1434 per task
  grok-4.7           $0.1489 per task
Enter fullscreen mode Exit fullscreen mode

Same task. Same token counts. Only the prices change.

Sonnet 5.5 and GPT-6.1 Sol have identical input and output prices. GPT-6.1 Sol still comes out cheaper here, because its cache reads cost half as much.

Grok 4.7 has the cheapest output. It doesn't win either, because this task is mostly cached input.

80% of the input came from the cache.

On a long agent task, the cache price can matter more than the output price.

Step 6: Find the Break Even Point

// How many times the baseline's tokens can `rival` spend
// on this task and still cost the same?
function breakEven(baseline: Price, rival: Price, turns: Turn[]) {
  return taskCost(turns, baseline) / taskCost(turns, rival);
}

console.log("\nStep 6: break-even token ratio vs claude-sonnet-5-5");

for (const [model, price] of Object.entries(PRICES)) {
  const ratio = breakEven(PRICES["claude-sonnet-5-5"], price, turns);
  console.log(`  ${model.padEnd(18)} ${ratio.toFixed(2)}x the tokens`);
}
Enter fullscreen mode Exit fullscreen mode

Output:

Step 6: break-even token ratio vs claude-sonnet-5-5
  claude-sonnet-5-5  1.00x the tokens
  claude-opus-5-5    0.54x the tokens
  gpt-6.1-sol        1.08x the tokens
  grok-4.7           1.04x the tokens
Enter fullscreen mode Exit fullscreen mode

This is the question the per-task claims are really about.

On this task shape, GPT-6.1 Sol can use 1.08x the tokens Sonnet 5.5 uses and still tie. Opus 5.5 has to use 0.54x the tokens to tie.

So if one model needs 30% fewer tokens, that gap is much bigger than any price difference in this table.

It also means a per-token price list can't tell you which model is cheaper. Only your own token counts can.

Step 7: Prompt Layout vs Model Choice

const noCache = simulateTask({ ...exampleTask, cachePrefix: false });
const sonnet = PRICES["claude-sonnet-5-5"];

console.log("\nStep 7: prompt layout vs model choice (claude-sonnet-5-5)");
console.log(`  stable prefix:   ${usd(taskCost(turns, sonnet))}`);
console.log(`  changing prefix: ${usd(taskCost(noCache, sonnet))}`);
console.log(`  changing prefix on grok-4.7: ${usd(taskCost(noCache, PRICES["grok-4.7"]))}`);
Enter fullscreen mode Exit fullscreen mode

Output:

Step 7: prompt layout vs model choice (claude-sonnet-5-5)
  stable prefix:   $0.1546
  changing prefix: $0.3424
  changing prefix on grok-4.7: $0.3168
Enter fullscreen mode Exit fullscreen mode

"Changing prefix" means something near the top of the prompt changes every turn. A timestamp in the system prompt is the classic one.

That breaks the cache. Every turn pays full input price.

On this example, it more than doubles the cost.

And switching to the model with the cheapest output doesn't win it back.

Fix the prompt layout before you switch models.

Stable prefix vs changing prefix

Step 8: Look at the Tail

// EXAMPLE INPUTS: ten runs of the same task, varying only
// how many tool turns and retries each run needed.
const runs: [number, number[]][] = [
  [3, []], [4, []], [4, []], [5, []], [5, [2]],
  [6, []], [6, [3]], [7, []], [9, [4, 7]], [14, [5, 9, 12]],
];

console.log("\nStep 8: ten example runs on gpt-6.1-sol");

const costs: number[] = [];
const times: number[] = [];

for (const [toolTurns, retryTurns] of runs) {
  const t = simulateTask({ ...exampleTask, toolTurns, retryTurns });
  costs.push(taskCost(t, PRICES["gpt-6.1-sol"]));
  times.push(timeToDoneMs(t) / 1000);
}

console.log(`  cost  p50 ${usd(percentile(costs, 50))}  p95 ${usd(percentile(costs, 95))}`);
console.log(`  time  p50 ${percentile(times, 50).toFixed(1)}s   p95 ${percentile(times, 95).toFixed(1)}s`);
Enter fullscreen mode Exit fullscreen mode

Output:

Step 8: ten example runs on gpt-6.1-sol
  cost  p50 $0.1245  p95 $0.3296
  time  p50 82.5s   p95 214.5s
Enter fullscreen mode Exit fullscreen mode

Averages hide the expensive runs.

The p95 run costs more than 2.5x the median here, and takes more than 2.5x as long.

That run needed 14 tool calls and 3 retries.

Budgets should be set on the tail, not on the demo.

To run everything:

npx tsx cost.ts
Enter fullscreen mode Exit fullscreen mode

I also wrote a small test, cost.test.ts, that checks these numbers against an independent hand calculation. All 7 checks pass.

Where It Breaks Down

This is a teaching calculator.

Here is what a real one needs.

The Token Counts Are Made Up

Everything in TaskShape is an example.

Real numbers come from your logs. The APIs return token usage with each response. Record it per task, not per request.

Vendor Per-Task Claims Use Vendor Tasks

"Up to 30% less per task" is Anthropic's measurement, on Anthropic's tests.

It might hold for your workload.

It's a hypothesis until you measure it.

Effort Settings Change Everything

The Artificial Analysis comparison above is the proof. Same price, different effort, almost twice the cost per task.

Treat effort as part of the model name when you compare.

Caching Rules Differ

Cache lifetimes, what counts as a hit, and whether caching is automatic or explicit all vary by provider. Anthropic, for example, prices 5-minute and 1-hour cache writes differently.

Our simulation assumes every stable prefix is a hit. Real hit rates are lower.

Latency Isn't Time to Done

Our latency is a formula with made-up constants.

Real latency has queues, rate limits, and tool calls that hang.

Prices Change

That's why every row has a checked date.

Put the price table in config, not in code, and re-check it when a model ships.

The Bigger Idea

When people compare models, they usually compare price lists.

I think that's the wrong unit.

A real cost model looks like this:

Task
  ↓
Turns        (tool calls + retries)
  ↓
Tokens       (fresh input, cache writes, cache reads, output)
  ↓
Prices       (sourced, dated)
  ↓
Cost per task  →  p50, p95, time to done
Enter fullscreen mode Exit fullscreen mode

The price table provides rates.

The task shape provides turns.

The cache provides discounts.

Retries provide surprises.

The logs provide evidence.

And the budget provides a limit.

Earlier today, in What Is an Agent Harness?, the harness had a step budget.

A step budget counts turns.

The next one should count dollars.

On September 19, in The Computer Is Becoming an API for AI, I wrote: "2027: Agents get identities. Dedicated accounts, credentials, budgets, permissions and audit trails become normal."

This part is prediction, not history.

But budgets need a unit, and tokens aren't it.

Cost per task, not cost per token.


Try Roster

I'm building Roster around this idea: AI employees with real responsibilities, tools, memory and schedules. Each one does recurring work you can count, which means it has a cost per task you can see.

If the same follow-ups, handoffs, and waiting loops keep eating your week, give them to an AI employee.

Try Roster →

Top comments (1)

Collapse
 
aifrontierpost profile image
AI Frontier Post •

Same token prices, $2.73 versus $1.50 per task — the Grok 4.7 effort-setting comparison makes "up to 30% less per task" read as a property of the benchmark task, not the model. Your step 7 runs the numbers on the flip side: a changing prefix more than doubles the cost, and no cheaper output price wins it back. So the practical hierarchy is prompt layout first, cache-read price second, model choice third — and your break-even framing already encodes that ordering.