Model pricing is easy to compare. Completed work is harder.
Fireworks reports a customer comparison in which total token use moved from 49.3K to 29.9K, a 39% reduction, while the task score moved from 0.751 to 0.753. That is useful evidence for token efficiency. It is not enough evidence for choosing a model in a real tool-using system.
The missing unit is an accepted work product.
A model response can be generated without being usable. A coding agent can return a patch that fails the next test. A research agent can produce a fluent summary with weak source coverage. An operations agent can create a list that has no owner or deadline. The model bill captures tokens. The task also consumes retries, tool calls, context, wall time, and human correction.
I use four states:
- Generated output — something returned by the system.
- Candidate work — output that passed a format check.
- Accepted work — output that met the task contract.
- Reusable work — accepted output that another person can reopen.
For an evaluation, I keep the surrounding system fixed. The same person defines the task. The same prompt sets behavior. The same tools provide evidence. The same acceptance line judges both candidates. I change one variable: the model.
Then I run five real tasks: coding, research, writing, data, and recurring operations. I record tokens, tool calls, retries, wall time, human minutes, accepted status, reusable-artifact status, and the reason for rejection.
The calculation is deliberately plain:
model_cost = input_tokens * input_price + output_tokens * output_price
task_cost = model_cost + tool_fees + retry_cost + human_review_cost
accepted_work_cost = task_cost / accepted_work_products
The denominator must be visible. A ten-dollar run that produces five accepted outputs is not the same as a ten-dollar run that produces one draft and a long repair cycle.
The acceptance line must match the work. For code, I might require a reproduced bug, a narrow change, passing existing tests, one regression test, and no unrelated files. For research, I might require five primary sources, one claim per source, dates on current claims, named uncertainty, and a decision note. For operations, I might require a dated input, a reason for each exception, an owner, a deadline, and a record that can be reopened tomorrow.
This is also why I read the comparison through the Agent Stack™: Human, Behavior, Tools, Memory, Agents, Models. The model is one layer. Tools add outside work. Memory adds retrieval and repeated context. Agents add steps and retries. Human review defines acceptance. Comparing only the last layer hides the path that creates the bill.
The current sources are bounded. Fireworks' Ember-1 figures are vendor-reported and the model is presented as a Research Preview. OpenAI's evaluation guidance recommends task-specific tests and human judgment. Neither source proves the result for my workload. They make the next experiment clearer.
Before changing a model, run the work-unit trial. Measure the cost of something a person can accept, reopen, and use.
Canonical article: https://echonerve.com/stop-comparing-model-prices-measure-cost-per-completed-work/
Top comments (0)