DEV Community

Daniel Pertu
Daniel Pertu

Posted on

Four AI jobs, four model choices, and the one we picked for comments was not the most accurate

Nakodo has four AI jobs, and they are not alike:

  1. Read a brand's website and write a brief: what they sell, which markets, which search keywords. One call per campaign, long input, structured output.
  2. Label a sample of comments on a creator's videos. Hundreds of tiny classifications, and by far the most calls.
  3. Write the reasons a creator does or does not fit, for the brand to read.
  4. Draft the first outreach email.

The easy thing is to pick one model and use it everywhere. The output of all four is visible to a paying customer, so "good enough" is not a judgement I wanted to make by reading a handful of samples. scripts/benchmark-models.ts makes it instead, in three commands:

pnpm benchmark prepare   # crawl the test sites, pick channels, fetch comments, label them with the reference model
pnpm benchmark run       # run every candidate on every task, judge, write benchmark/results.md
pnpm benchmark cleanup   # delete the stored comment text
Enter fullscreen mode Exit fullscreen mode

The parts that make it a measurement

The test set is a file. prepare crawls five real brand websites, pulls fifteen channels with their recent videos and metrics out of a test campaign, and samples seven comments per video, into benchmark/data.json. Every candidate then reads that file. The API quota (about 30 YouTube units) is spent once, and a model that looks good is not just the model that happened to get the easier channels.

Labels are compared, not graded. Comment labels have a right answer, so a deliberately expensive reference model labels the sample once at medium effort and the candidates are scored as accuracy against it, plus a separate "language match" column, because a label in the wrong language is a different failure.

Written output is graded blind. A judge model sees the exact input the writer saw, then the candidate outputs shuffled by a seeded shuffle and relabelled A, B, C, D. It is told it does not know who wrote which, and that ties are allowed:

const order = shuffle(outputs, seed);
const ids = order.map((_, i) => String.fromCharCode(65 + i));
Enter fullscreen mode Exit fullscreen mode

The judge and the reference model are not candidates. They cost 20 to 100 times what the candidates cost per token, which is the point: they are only ever run on a test set of 205 comments and 15 channels.

Validity is its own axis. A model that writes beautifully and ignores the schema one time in five is not usable in a pipeline that runs unattended. So each task has a cheap structural check that is counted per call, separately from quality. The outreach one reads:

if (r.subject.trim() && words >= 50 && words <= 200 && !r.body.includes("—") && r.optOutLine.trim()) m.valid++;
Enter fullscreen mode Exit fullscreen mode

Yes, one of the validity conditions is that there is no em dash in the body. Our emails do not use them, the instruction says so, and a model that cannot follow that instruction is telling you something about the other instructions.

Cost is measured, not assumed. Real token counts per call, reasoning tokens included, cached input billed at the cached rate:

function cost(model: string, u: Usage): number {
  const [inp, cached, out] = PRICES[model];
  return ((u.inputTokens - u.cachedInputTokens) * inp + u.cachedInputTokens * cached + u.outputTokens * out) / 1e6;
}
Enter fullscreen mode Exit fullscreen mode

and reported as cost per 1,000 units of the thing the task actually produces: per brief, per comment, per creator, per draft.

Provider weather is not model quality. A 429 or a 5xx or a dropped connection says nothing about the model, so those are retried with a 5_000 * 3 ** i backoff up to five attempts before anything is counted. A failure that survives that counts as an invalid call, which is the honest accounting: it did not produce output.

The run is resumable. Every item writes a checkpoint of the metrics and the raw outputs so far, keyed by task and item. An interrupted run resumes instead of paying for the finished half again.

The comment text is deleted. It lives only in two gitignored files, no author names or ids are fetched at all, and cleanup removes them. A benchmark corpus is still other people's writing.

What came out

From the run on 1 October 2026, cost in USD per 1,000 units:

Task Best quality Picked Cost of the pick Cost of the cheapest candidate
Site research 8.80 gpt-6-luna $0.461 per 1,000 briefs $0.818
Comment labels 97.1% gpt-5.6-luna $0.033 per 1,000 comments $0.009
Creator rationale 8.87 gpt-6-luna $0.209 per 1,000 creators $0.163
Outreach drafts 9.53 gpt-6-luna $0.235 per 1,000 drafts $0.144

Two results were worth the whole exercise.

The cheapest model per token was the most expensive per brief. The nano-class model at $0.05 per million input tokens produced briefs at $0.818 per thousand, against $0.461 for a model with double the input price. It reasons and emits its way to a worse score at nearly twice the cost. Per-token list prices are not a ranking of anything you care about.

Three tasks went one way and comments went the other. On accuracy the two leading models were 96.4% and 97.1%, which is within noise on 205 comments. The gap that decided it was the validity column: the otherwise cheaper model returned a complete label array for only 80% of calls. The selection rule is written down in the report rather than left to my judgement on the day:

Among models with at least 90% valid output, the cheapest one within 3 points of the best accuracy (comments) or 0.5 rubric points of the best score (other tasks).

That is a 3.7 times higher per-comment cost to buy reliability on the job with the most calls, which is exactly the trade that needs a number rather than a vibe. Comments are also the one job that is already skipped when a better sample could not change the outcome, so the dearer model runs on fewer calls than its position in the list suggests.

Median latency ran from 25s to 57s across the candidates, and it did not decide anything. None of these calls sit in a request path. They sit in a Postgres job queue behind a worker, so latency buys throughput, not page speed, and the slowest candidate was also the worst.

The four jobs are described, in the language a customer reads rather than this one, on how it works: the brief is job one, the score and its reasons are job three, and the drafts are job four. How many AI drafts a plan gets per day is on pricing, and that number is a direct consequence of the per-draft cost in the table above.

The prompts, the rubric text and the scoring internals stay in the repo. The harness is the part worth copying: fixed input on disk, a reference answer where one exists, blind judging where one does not, validity counted separately from quality, real token costs including the reasoning you did not ask for, and a written rule that turns the table into a decision without you in the room.

Top comments (0)