Your LLM passes the demo every time. Then a customer sends a prompt you didn't test, and the answer is garbage. The gap between "it worked in the eval" and "it works in production" is almost always the same: you measured accuracy, but you never measured whether the model's confidence was trustworthy.
laya-evals closes that gap. It is the fifth tool in my open-source series on the laya decision engine, and the one that keeps the other four honest.
Judge accuracy: 0.8968 (Cohen's kappa 0.7938)
Calibration: ECE 0.0253, Brier 0.0779
Cost per 1,000 judgments: $0 (local inference)
CI gate: fails the build when accuracy or calibration regresses
What it does
Three things, in one pipeline:
1. Rubric-based LLM-as-judge. Score your eval sets with a local laya decision model instead of paying for GPT-4 as a judge. Choice questions, score questions, or no-ground-truth mode. 6.4 decisions per second on CPU.
2. Calibration audit. Accuracy tells you how often the model is right. Calibration tells you whether you can trust its confidence. laya-evals computes Expected Calibration Error, Brier score, and reliability bins, then tells you which confidence threshold to use for each question shape.
Why per shape? Because a 2-option question needs a different gate than a 20-option question. The SST-2 benchmark (2 options) needs a threshold of 0.865 for 95% accuracy. The MASSIVE intent benchmark (20 options) needs 0.994. One global min_confidence is not a policy, and laya-evals proves it with your own numbers.
3. CI regression gate. Wire it into GitHub Actions. Exit code 0 means pass, exit code 1 means a measured regression, exit code 2 means the run itself was invalid. Your eval suite becomes a build gate, not a suggestion.
The numbers
872 SST-2 sentences, 2-option rubric, local laya judge:
| Measurement | Result |
|---|---|
| Judge accuracy vs gold | 0.8968 |
| Cohen's kappa | 0.7938 |
| ECE (15 bins) | 0.0253 |
| Brier score | 0.0779 |
| Throughput | 6.4 decisions/s |
| API cost | $0 |
Cross-checked against Z.ai's GLM-5.3-flash as reference judge: 95.1% accuracy, 90.7% agreement, kappa 0.814. The local judge is cheaper and nearly as reliable.
I also re-measured four public laya benchmark claims with laya-evals and got values within 0.0003 of the published numbers. The reproduction pack is committed.
Try it
pip install laya-evals
from laya_evals import Judge, ece, brier_score, advise_thresholds
judge = Judge()
results = judge.run(eval_set, rubric)
print(ece(results.confidences, results.correct))
print(advise_thresholds(results))
The series
This is the fifth and last tool in the laya series, built on the laya decision engine by NandhaKishorM. The same engine that TypeSafe shipped as Jev and OpenAI shipped as the Decisions API on Luna this week. The difference: this one runs on your machine, for free, and the benchmarks are committed.
- laya-router: 54.9% cost reduction, $0 per routing decision
- laya-compactor: 70% fewer RAG tokens, same answer quality
- laya-triage: support triage from 51% to 90.5% after fine-tuning
- laya-phishield: explainable phishing detection, $0 per 1,000 emails
Every tool ships with its benchmarks, including the numbers that hurt.
Top comments (0)