Everyone is talking about Evaluation Cards and reproducible eval benches — the EvalEval × UK AISI write-up on Hugging Face, plus a week of discourse where green numbers still hide the wrong harness (and DeepEval-style “which half is broken?” chatter keeps landing). That heat is a live market object, not a niche footnote.
Here is what I would require before I trusted a green eval score for high-stakes retrieval or generation — and what I would not claim yet.
The stake
Teams will ship the wrong default.
A dashboard can show a comforting pass rate while the report never pins how the score was earned: which protocol, which compute budget, which dataset slice, and whether the number measured retrieval, generation, or a blended soup. Celebrate that number and you optimize a screenshot. The next engineer cannot reproduce the gate, and the next incident cannot tell which half failed.
If your domain is identifier-heavy, multi-hop, or regulated-adjacent, a fluent answer after an unpinned harness is not a win. It is a silent miss with a pretty badge.
Opinion (one sentence)
A green score without a pinned harness — protocol, compute, dataset slice, and which half of the pipeline it measures — is not an eval contract; it is a screenshot.
That is a preference with a limiter, not a SOTA claim. I have not run EvalEval in production. AISI card fields and anyone else’s leaderboard numbers stay theirs. What I am formalizing is the harness pin as a first-class part of eval-as-contract, from the same reliable-systems lane I already publish on.
What a pinned harness looks like (menu, not recipe)
Evaluation Cards earned attention because they make the boring fields visible: protocol, compute, data cut, scoring rules. You do not need their exact schema to steal the shape.
| Field | Question it answers | Failure if skipped |
|---|---|---|
| Protocol | What procedure produced this number? | “We eval’d it” with no replay path |
| Compute | What budget / hardware / call limits? | Incomparable scores across teams |
| Dataset slice | Which queries / docs / holdouts? | Student-subset vibes framed as prod |
| Pipeline half | Retrieval, generation, or joint? | Green faithfulness while recall is broken (or the reverse) |
The DeepEval-style lesson that keeps circulating is the same shape: a single RAG score can hide which half failed. Pin the half. Prefer two honest columns over one blended crown.
The check you can run without my internals
You do not need my private stacks (and I will not publish them). Steal this menu-only check:
- Pin the harness block before the score. Same report header every time: protocol name/version, compute class, dataset slice id, and pipeline half (retrieval / generation / joint). If any field is missing, the number is not ready to celebrate.
- Split halves when you report. For the same query set, publish retrieval quality and generation/faithfulness as separate columns. One green blended score is not enough — “which half is broken?” should be answerable from the table alone.
- Freeze the evaluator before you chase labels. If the judge, metric, or prompt changed mid-board, say so. Moving the goalposts mid-week and claiming progress is theater.
- Keep a missing-evidence arm. Same task with the gold document removed (or retrieval forced empty). If generation still narrates a confident answer, the eval contract is incomplete — regardless of how green the happy-path column looks.
No thresholds copied from anyone’s blog as “our production numbers.” No paste recipe. Calibrate cutoffs on your corpus; where missing evidence is costly, treat empty retrieval as a failure in the harness, not as a soft success.
Why query-class honesty transfers
On Portuguese clinical text, BM25 and dense retrieval solved different query classes; fusing them beat either alone on a public 500-query study. Exact terms and conceptual phrasing fail in different ways. That finding is checkable: open code at nomad-link-id/hybrid-rag-pipeline, companion write-up on Dev.to.
The lesson that transfers to Evaluation Cards is not “clone our fusion.” It is eval honesty on query classes and pipeline halves. If your harness cannot say which class and which half produced the green cell, you are not ready to ship the number into a decision.
What I will not claim
- That EvalEval, UK AISI cards, or any one vendor harness is universally correct for every workload. Workload-honest eval first.
- That we “ran EvalEval inside Cortexa/DocMinds.” We did not. Trend-jacking with fake usage is spam.
- That a pinned harness alone proves end-to-end answer quality. Downstream faithfulness and missing-doc policy remain separate gates.
- Internal thresholds, prompts, or clone guides for private products. Menu only: pin the fields, split the halves — not the spice blend.
Why this formalizes who we are
The market is amplifying Evaluation Cards and green eval dashboards. The thesis I want indexed next to that heat:
- Eval as contract — protocol, compute, slice, and pipeline half are gates, not vibes.
- Reliable systems — production AI fails on harness mismatch and wrong-half celebration — not only on model choice.
- Retrieve-first — when the retrieval half is unpinned, generation fluency is not evidence.
If a sharp reader walks away believing this engineer will not celebrate a green score without a pinned harness, the post did its job. Follows that come from that are the point — not volume.
Soft pointers (contribution first)
- Public hybrid pipeline (BM25 + dense + RRF): https://github.com/nomad-link-id/hybrid-rag-pipeline
- Empirical companion (500 clinical queries, open methodology): Dev.to — Two Retrieval Methods Are Better Than One
- Site / builder trail: https://igoreduardo.com
- Market object that sparked this note: EvalEval × UK AISI Evaluation Cards (HF)
Stop here. No recipe. No “we proved SOTA.” Preference + limiter + public trail.
By Igor Eduardo · Austin, TX · Engineer of reliable search and AI systems for high-stakes science · https://igoreduardo.com
Top comments (0)