Most LLM pricing pages answer one question.
What does a million tokens cost?
That's a useful number.
It's just not the number that shows up on the bill.
The bill is made of tasks. A task is a few turns, a few tool calls, a retry or two, and a prompt that gets resent every single time.
This month, the model launches started saying that out loud.
- On September 21, SpaceXAI released Grok 4.7, "served at the same price and speed as Grok 4.6." Its docs list $2 input and $6 output per 1M tokens.
- On September 22, Anthropic released Claude Opus 5.5. Its token prices are 20% lower than Opus 5. But Anthropic's headline number is different: "at default settings it will cost 40% less than Opus 5 on typical workloads."
- On September 28, Anthropic released Claude Sonnet 5.5 at the same $2 / $10 as Sonnet 5, and wrote: "In our testing, it costs up to 30% less per task than its predecessor."
- On September 29, at DevDay, OpenAI released GPT-6.1 Sol. OpenAI's pricing page lists it at $2 input and $10 output, a fifth of GPT-6 Astra's $10 / $50. The same day, GitHub added it to Copilot, saying it completed tasks "while using noticeably fewer tokens and steps."
Different companies.
Same pitch.
The price per token barely moved. The claim moved to the task.
That distinction matters, because it cuts both ways.
Artificial Analysis, a third party, compares Grok 4.7 (high) with Grok 4.6 (medium). Same token prices. It lists $2.73 versus $1.50 per task, because Grok 4.7 (high) used 66k output tokens per task and Grok 4.6 (medium) used 28k. Those are different effort settings, so it isn't a like-for-like model comparison. It is a clean example of the same price producing a different bill.
So let's stop comparing price lists and build a tiny cost-per-task calculator.
By the end, you will run one command:
npx tsx cost.ts
No API key.
No SDK.
Just list prices, a task model, and arithmetic you can check.
Table of Contents
- Project Setup
- Step 1: A Price Table With Sources
- Step 2: Price One Turn
- Step 3: Model a Task
- Step 4: Cache Hit Rate and Latency
- Step 5: Compare Models per Task
- Step 6: Find the Break Even Point
- Step 7: Prompt Layout vs Model Choice
- Step 8: Look at the Tail
- Where It Breaks Down
- The Bigger Idea
Project Setup
You will need Node.js 18 or newer.
mkdir cost-per-task
cd cost-per-task
npm init -y
npm install --save-dev typescript tsx @types/node
Save the following blocks, in order, as cost.ts.
Step 1: A Price Table With Sources
type Price = {
input: number; // $ per 1M uncached input tokens
cachedInput: number; // $ per 1M cache-hit input tokens
cacheWrite: number | null; // $ per 1M tokens written to cache (null = not listed)
output: number; // $ per 1M output tokens
source: string;
checked: string;
};
// Standard tier, short context. Checked by hand on Sep 30, 2026.
const PRICES: Record<string, Price> = {
"claude-sonnet-5-5": {
input: 2, cachedInput: 0.2, cacheWrite: 2.5, output: 10,
source: "https://platform.claude.com/docs/en/about-claude/pricing",
checked: "2026-09-30",
},
"claude-opus-5-5": {
input: 4, cachedInput: 0.2, cacheWrite: 5, output: 20,
source: "https://platform.claude.com/docs/en/about-claude/pricing",
checked: "2026-09-30",
},
"gpt-6.1-sol": {
input: 2, cachedInput: 0.1, cacheWrite: 2.5, output: 10,
source: "https://developers.openai.com/api/docs/pricing",
checked: "2026-09-30",
},
"grok-4.7": {
input: 2, cachedInput: 0.5, cacheWrite: null, output: 6,
source: "https://docs.x.ai/docs/models/grok-4.7",
checked: "2026-09-30",
},
};
Every row carries a URL and the date I checked it.
Prices change.
A price without a source is a rumor with a decimal point.
Three details are easy to miss:
- Cache reads are not priced the same everywhere. $0.10 on GPT-6.1 Sol, $0.20 on both Claude models, $0.50 on Grok 4.7.
- Anthropic and OpenAI list a separate cache write price. The Grok 4.7 docs don't list one, so the code bills those tokens at the normal input rate.
- These are short-context rates. OpenAI charges more above 272K input tokens and xAI above 200K. Our example stays well under both.
Step 2: Price One Turn
type Turn = {
freshInput: number; // input tokens billed at the full rate
cacheWrite: number; // input tokens stored in the cache this turn
cacheRead: number; // input tokens served from the cache
output: number; // output tokens, including reasoning tokens
latencyMs: number;
retry: boolean;
};
function turnCost(t: Turn, p: Price): number {
const writeRate = p.cacheWrite ?? p.input;
return (
t.freshInput * p.input +
t.cacheWrite * writeRate +
t.cacheRead * p.cachedInput +
t.output * p.output
) / 1_000_000;
}
const taskCost = (turns: Turn[], p: Price) =>
turns.reduce((sum, t) => sum + turnCost(t, p), 0);
A turn is one model call.
It has four kinds of tokens, and each one has its own rate.
In this calculator, output includes reasoning tokens. Check how your provider reports them. They're usually the part that grows when you turn effort up.
Step 3: Model a Task
This is the important step.
// EXAMPLE INPUTS. Not measurements of any model.
type TaskShape = {
systemTokens: number;
taskTokens: number;
toolTurns: number; // model calls a tool, gets a result back
toolResultTokens: number;
outputPerTurn: number;
retryTurns: number[]; // turn indexes that failed and ran again
cachePrefix: boolean; // is the start of the prompt stable?
firstTokenMs: number;
msPerOutputToken: number;
toolMs: number;
};
function simulateTask(s: TaskShape): Turn[] {
const turns: Turn[] = [];
let context = s.systemTokens + s.taskTokens;
let cached = 0;
for (let i = 0; i <= s.toolTurns; i++) {
const attempts = s.retryTurns.includes(i) ? 2 : 1;
for (let a = 0; a < attempts; a++) {
const read = s.cachePrefix ? cached : 0;
const write = s.cachePrefix ? context - cached : 0;
const isTool = i < s.toolTurns;
turns.push({
freshInput: s.cachePrefix ? 0 : context,
cacheWrite: write,
cacheRead: read,
output: s.outputPerTurn,
latencyMs:
s.firstTokenMs +
s.outputPerTurn * s.msPerOutputToken +
(isTool ? s.toolMs : 0),
retry: a > 0,
});
if (s.cachePrefix) cached = context;
}
context += s.outputPerTurn + s.toolResultTokens;
}
return turns;
}
Everything in TaskShape is an example input. I made these numbers up to show the math. They are not measurements of any model.
The key line is this one:
context += s.outputPerTurn + s.toolResultTokens;
Every tool call makes the next prompt longer.
The model re-reads the system prompt, the task, and every earlier result on every turn.
So input cost doesn't grow in a straight line with the number of turns. It grows faster.
A retry sends the same prompt again and pays for the output again.
With a stable prefix, most of that resent context comes from the cache. Without one, you pay full price for all of it, every turn.
Step 4: Cache Hit Rate and Latency
function cacheHitRate(turns: Turn[]): number {
let read = 0;
let total = 0;
for (const t of turns) {
read += t.cacheRead;
total += t.freshInput + t.cacheWrite + t.cacheRead;
}
return total === 0 ? 0 : read / total;
}
const timeToDoneMs = (turns: Turn[]) =>
turns.reduce((sum, t) => sum + t.latencyMs, 0);
function percentile(values: number[], p: number): number {
const sorted = [...values].sort((a, b) => a - b);
const index = Math.ceil((p / 100) * sorted.length) - 1;
return sorted[Math.max(0, index)];
}
Cache hit rate is cached input divided by all input.
timeToDoneMs adds up every turn, including tool time.
That's the latency a user actually feels.
Time to first token is one number inside it.
Step 5: Compare Models per Task
const exampleTask: TaskShape = {
systemTokens: 6_000,
taskTokens: 1_500,
toolTurns: 6,
toolResultTokens: 2_500,
outputPerTurn: 800,
retryTurns: [3],
cachePrefix: true,
firstTokenMs: 900,
msPerOutputToken: 12,
toolMs: 1_500,
};
const usd = (n: number) => `$${n.toFixed(4)}`;
const turns = simulateTask(exampleTask);
console.log("Step 5: same example task, every model");
console.log(` turns: ${turns.length} (${turns.filter((t) => t.retry).length} retry)`);
console.log(` cache hit rate: ${(cacheHitRate(turns) * 100).toFixed(1)}%`);
for (const [model, price] of Object.entries(PRICES)) {
console.log(` ${model.padEnd(18)} ${usd(taskCost(turns, price))} per task`);
}
Output:
Step 5: same example task, every model
turns: 8 (1 retry)
cache hit rate: 80.4%
claude-sonnet-5-5 $0.1546 per task
claude-opus-5-5 $0.2869 per task
gpt-6.1-sol $0.1434 per task
grok-4.7 $0.1489 per task
Same task. Same token counts. Only the prices change.
Sonnet 5.5 and GPT-6.1 Sol have identical input and output prices. GPT-6.1 Sol still comes out cheaper here, because its cache reads cost half as much.
Grok 4.7 has the cheapest output. It doesn't win either, because this task is mostly cached input.
80% of the input came from the cache.
On a long agent task, the cache price can matter more than the output price.
Step 6: Find the Break Even Point
// How many times the baseline's tokens can `rival` spend
// on this task and still cost the same?
function breakEven(baseline: Price, rival: Price, turns: Turn[]) {
return taskCost(turns, baseline) / taskCost(turns, rival);
}
console.log("\nStep 6: break-even token ratio vs claude-sonnet-5-5");
for (const [model, price] of Object.entries(PRICES)) {
const ratio = breakEven(PRICES["claude-sonnet-5-5"], price, turns);
console.log(` ${model.padEnd(18)} ${ratio.toFixed(2)}x the tokens`);
}
Output:
Step 6: break-even token ratio vs claude-sonnet-5-5
claude-sonnet-5-5 1.00x the tokens
claude-opus-5-5 0.54x the tokens
gpt-6.1-sol 1.08x the tokens
grok-4.7 1.04x the tokens
This is the question the per-task claims are really about.
On this task shape, GPT-6.1 Sol can use 1.08x the tokens Sonnet 5.5 uses and still tie. Opus 5.5 has to use 0.54x the tokens to tie.
So if one model needs 30% fewer tokens, that gap is much bigger than any price difference in this table.
It also means a per-token price list can't tell you which model is cheaper. Only your own token counts can.
Step 7: Prompt Layout vs Model Choice
const noCache = simulateTask({ ...exampleTask, cachePrefix: false });
const sonnet = PRICES["claude-sonnet-5-5"];
console.log("\nStep 7: prompt layout vs model choice (claude-sonnet-5-5)");
console.log(` stable prefix: ${usd(taskCost(turns, sonnet))}`);
console.log(` changing prefix: ${usd(taskCost(noCache, sonnet))}`);
console.log(` changing prefix on grok-4.7: ${usd(taskCost(noCache, PRICES["grok-4.7"]))}`);
Output:
Step 7: prompt layout vs model choice (claude-sonnet-5-5)
stable prefix: $0.1546
changing prefix: $0.3424
changing prefix on grok-4.7: $0.3168
"Changing prefix" means something near the top of the prompt changes every turn. A timestamp in the system prompt is the classic one.
That breaks the cache. Every turn pays full input price.
On this example, it more than doubles the cost.
And switching to the model with the cheapest output doesn't win it back.
Fix the prompt layout before you switch models.
Step 8: Look at the Tail
// EXAMPLE INPUTS: ten runs of the same task, varying only
// how many tool turns and retries each run needed.
const runs: [number, number[]][] = [
[3, []], [4, []], [4, []], [5, []], [5, [2]],
[6, []], [6, [3]], [7, []], [9, [4, 7]], [14, [5, 9, 12]],
];
console.log("\nStep 8: ten example runs on gpt-6.1-sol");
const costs: number[] = [];
const times: number[] = [];
for (const [toolTurns, retryTurns] of runs) {
const t = simulateTask({ ...exampleTask, toolTurns, retryTurns });
costs.push(taskCost(t, PRICES["gpt-6.1-sol"]));
times.push(timeToDoneMs(t) / 1000);
}
console.log(` cost p50 ${usd(percentile(costs, 50))} p95 ${usd(percentile(costs, 95))}`);
console.log(` time p50 ${percentile(times, 50).toFixed(1)}s p95 ${percentile(times, 95).toFixed(1)}s`);
Output:
Step 8: ten example runs on gpt-6.1-sol
cost p50 $0.1245 p95 $0.3296
time p50 82.5s p95 214.5s
Averages hide the expensive runs.
The p95 run costs more than 2.5x the median here, and takes more than 2.5x as long.
That run needed 14 tool calls and 3 retries.
Budgets should be set on the tail, not on the demo.
To run everything:
npx tsx cost.ts
I also wrote a small test, cost.test.ts, that checks these numbers against an independent hand calculation. All 7 checks pass.
Where It Breaks Down
This is a teaching calculator.
Here is what a real one needs.
The Token Counts Are Made Up
Everything in TaskShape is an example.
Real numbers come from your logs. The APIs return token usage with each response. Record it per task, not per request.
Vendor Per-Task Claims Use Vendor Tasks
"Up to 30% less per task" is Anthropic's measurement, on Anthropic's tests.
It might hold for your workload.
It's a hypothesis until you measure it.
Effort Settings Change Everything
The Artificial Analysis comparison above is the proof. Same price, different effort, almost twice the cost per task.
Treat effort as part of the model name when you compare.
Caching Rules Differ
Cache lifetimes, what counts as a hit, and whether caching is automatic or explicit all vary by provider. Anthropic, for example, prices 5-minute and 1-hour cache writes differently.
Our simulation assumes every stable prefix is a hit. Real hit rates are lower.
Latency Isn't Time to Done
Our latency is a formula with made-up constants.
Real latency has queues, rate limits, and tool calls that hang.
Prices Change
That's why every row has a checked date.
Put the price table in config, not in code, and re-check it when a model ships.
The Bigger Idea
When people compare models, they usually compare price lists.
I think that's the wrong unit.
A real cost model looks like this:
Task
↓
Turns (tool calls + retries)
↓
Tokens (fresh input, cache writes, cache reads, output)
↓
Prices (sourced, dated)
↓
Cost per task → p50, p95, time to done
The price table provides rates.
The task shape provides turns.
The cache provides discounts.
Retries provide surprises.
The logs provide evidence.
And the budget provides a limit.
Earlier today, in What Is an Agent Harness?, the harness had a step budget.
A step budget counts turns.
The next one should count dollars.
On September 19, in The Computer Is Becoming an API for AI, I wrote: "2027: Agents get identities. Dedicated accounts, credentials, budgets, permissions and audit trails become normal."
This part is prediction, not history.
But budgets need a unit, and tokens aren't it.
Cost per task, not cost per token.
Try Roster
I'm building Roster around this idea: AI employees with real responsibilities, tools, memory and schedules. Each one does recurring work you can count, which means it has a cost per task you can see.
If the same follow-ups, handoffs, and waiting loops keep eating your week, give them to an AI employee.



Top comments (1)
Same token prices, $2.73 versus $1.50 per task — the Grok 4.7 effort-setting comparison makes "up to 30% less per task" read as a property of the benchmark task, not the model. Your step 7 runs the numbers on the flip side: a changing prefix more than doubles the cost, and no cheaper output price wins it back. So the practical hierarchy is prompt layout first, cache-read price second, model choice third — and your break-even framing already encodes that ordering.