DEV Community

Robert
Robert

Posted on Originally published at neuragrowth.co

Sonnet vs Haiku vs Opus on critique.flashcards: When To Reach For Which

If you are building a flashcard critique step and trying to pick between Sonnet, Haiku and Opus, the choice is not as simple as picking the cheapest model. Two of the three landed at the same cost per call, and the reason they did is different enough to matter.

We ran the critique.flashcards job 904 times across 90 days using three models: claude-sonnet-4-6, claude-haiku-4-5 and claude-opus-4-8. All three had zero failures on our side. The comparison below covers cost, speed and the token picture behind each number.

Haiku last ran on 2026-06-08. Sonnet and Opus both ran through 2026-07-30. One of those facts is the most useful in the piece.

What the cost figures actually show

Sonnet (anthropic/claude-sonnet-4-6) and Haiku (anthropic/claude-haiku-4-5) both came in at $0.00213 per call. Opus (anthropic/claude-opus-4-8) cost $0.00382 per call, which is 1.79 times the cost of the other two, or $0.00169 more per call.

Stopping at the per-call figure would miss the most important part. Sonnet and Haiku arrived at the same number by very different routes.

Sonnet billed us for 115 input tokens per call. Its total input was 2,796 tokens per call, with 92.9 percent coming from the prompt cache. We paid full rate on a small fraction of a large job.

Haiku billed us for 2,140 input tokens per call. Its total input was also 2,140 tokens, with zero percent cached. Every token was a cache miss, every time. Haiku and Sonnet cost the same per call, but Haiku was paying full rate on the whole prompt while Sonnet was paying almost nothing on most of it.

Opus billed us for 185 input tokens per call against a total of 3,653, with 92.7 percent cached. Its higher per-call cost comes from Opus rates on those billed tokens, not from a larger unbilled job.

Speed across all three models

Opus was fastest at a median of 1.8 seconds per call. Haiku ran at 1.9 seconds, which is 1.06 times slower and 0.1 seconds behind. Sonnet ran at 2.0 seconds, which is 1.11 times slower than Opus and 0.2 seconds behind.

The spread between fastest and slowest is 0.2 seconds across all three models.

Failure rates

All three models returned zero failures across their combined 904 calls. Sonnet ran 669 calls, Haiku 147, Opus 88. The facts name Sonnet as the tool with the most failures, and the figure behind that label is zero.

We cannot draw any conclusion about reliability from a record where no tool failed.

Haiku stopped on 2026-06-08

Haiku ran for six days, from 2026-06-03 to 2026-06-08, across 147 calls. Sonnet and Opus both ran through 2026-07-30.

The record does not say why we stopped using Haiku. What the record does show is that Haiku was the only model with no caching on this job. Whether that was a configuration issue or a property of the model on our prompts, the facts do not say.

Sonnet and Opus are the two models still in use.

What this comparison can and cannot tell you

These numbers come from one job, critique.flashcards, on our prompts, at our call volumes, over 90 days. The cache hit rates are a product of how we structure our prompts, not a property of the models in isolation.

A reader running a different job, with different prompt structure or different volumes, may see different cache behaviour and therefore a different cost picture. The per-call figures here are what we paid, not vendor list prices, and they should not be read as a benchmark.

Three models, 904 calls, one job. That is the scope of what we can honestly say.

Which two we would pick today and when we would pick differently

We are running Sonnet for the bulk of this job, 669 calls against Opus at 88. At $0.00213 per call with 92.9 percent of input cached, Sonnet is the lower-cost choice when prompt caching holds.

Opus costs 1.79 times as much per call. We are still using it, which means we are accepting that premium for a reason the record does not state explicitly.

The record does not say whether we would return to Haiku or under what conditions. What the record does say is that zero cache hits on a job where Sonnet cached 92.9 percent of its input is the clearest difference between the two.

Pick Sonnet for this job if your prompt caches well and $0.00213 per call is your target. Look at Opus if you can absorb a cost that is 1.79 times higher. Do not add Haiku back without first understanding why it had zero cache hits on the same job where Sonnet cached 92.9 percent of its input. The record does not answer that question, and neither do we.


Originally published at neuragrowth.co. NeuraGrowth is a one-person digital-products studio; this is the log of what its pipeline does and where it breaks.

Top comments (0)