DEV Community

Cover image for Judge cheap, audit confidence: a CI gate for LLM evals (open source)
gj0xv
gj0xv

Posted on

Judge cheap, audit confidence: a CI gate for LLM evals (open source)

Your LLM passes the demo every time. Then a customer sends a prompt you didn't test, and the answer is garbage. The gap between "it worked in the eval" and "it works in production" is almost always the same: you measured accuracy, but you never measured whether the model's confidence was trustworthy.

laya-evals closes that gap. It is the fifth tool in my open-source series on the laya decision engine, and the one that keeps the other four honest.

Judge accuracy: 0.8968 (Cohen's kappa 0.7938)
Calibration: ECE 0.0253, Brier 0.0779
Cost per 1,000 judgments: $0 (local inference)
CI gate: fails the build when accuracy or calibration regresses
Enter fullscreen mode Exit fullscreen mode

What it does

Three things, in one pipeline:

1. Rubric-based LLM-as-judge. Score your eval sets with a local laya decision model instead of paying for GPT-4 as a judge. Choice questions, score questions, or no-ground-truth mode. 6.4 decisions per second on CPU.

2. Calibration audit. Accuracy tells you how often the model is right. Calibration tells you whether you can trust its confidence. laya-evals computes Expected Calibration Error, Brier score, and reliability bins, then tells you which confidence threshold to use for each question shape.

Why per shape? Because a 2-option question needs a different gate than a 20-option question. The SST-2 benchmark (2 options) needs a threshold of 0.865 for 95% accuracy. The MASSIVE intent benchmark (20 options) needs 0.994. One global min_confidence is not a policy, and laya-evals proves it with your own numbers.

3. CI regression gate. Wire it into GitHub Actions. Exit code 0 means pass, exit code 1 means a measured regression, exit code 2 means the run itself was invalid. Your eval suite becomes a build gate, not a suggestion.

The numbers

872 SST-2 sentences, 2-option rubric, local laya judge:

Measurement Result
Judge accuracy vs gold 0.8968
Cohen's kappa 0.7938
ECE (15 bins) 0.0253
Brier score 0.0779
Throughput 6.4 decisions/s
API cost $0

Cross-checked against Z.ai's GLM-5.3-flash as reference judge: 95.1% accuracy, 90.7% agreement, kappa 0.814. The local judge is cheaper and nearly as reliable.

I also re-measured four public laya benchmark claims with laya-evals and got values within 0.0003 of the published numbers. The reproduction pack is committed.

Try it

pip install laya-evals
Enter fullscreen mode Exit fullscreen mode
from laya_evals import Judge, ece, brier_score, advise_thresholds

judge = Judge()
results = judge.run(eval_set, rubric)
print(ece(results.confidences, results.correct))
print(advise_thresholds(results))
Enter fullscreen mode Exit fullscreen mode

The series

This is the fifth and last tool in the laya series, built on the laya decision engine by NandhaKishorM. The same engine that TypeSafe shipped as Jev and OpenAI shipped as the Decisions API on Luna this week. The difference: this one runs on your machine, for free, and the benchmarks are committed.

  • laya-router: 54.9% cost reduction, $0 per routing decision
  • laya-compactor: 70% fewer RAG tokens, same answer quality
  • laya-triage: support triage from 51% to 90.5% after fine-tuning
  • laya-phishield: explainable phishing detection, $0 per 1,000 emails

Every tool ships with its benchmarks, including the numbers that hurt.

Top comments (0)