I built a deterministic LLM evaluation engine without an LLM judge
Most current LLM evaluation workflows eventually face a version of the same problem:
How do you evaluate the output of a model?
One common approach is to use another LLM as the evaluator.
That can be useful, but it also introduces another model, another source of variability, and another layer of judgment into the evaluation process.
I wanted to explore a different approach.
So I built DIBER Core.
What is DIBER?
DIBER is an open-source deterministic evaluation engine for LLM outputs.
It does not ask an LLM to judge another LLM.
Instead, it works from an explicit evaluation structure:
LLM run → Ground Truth → human-defined classifications → deterministic engine → audit report
The engine performs the normalization, weighting, comparison and statistical calculations.
The same structured inputs produce the same report.
It does not reduce everything to one score
A model can cover a large amount of information while still producing significant errors.
It can also provide a technically correct answer while omitting important points.
So DIBER keeps different dimensions separate.
It evaluates things such as:
- Coverage
- Quality
- Usable information
- Effective information
- Omissions
- Errors
- Mathematical errors
- Hallucinations
- Overclaims
- Inferences
- Extra unsupported claims
- Relation consistency
The goal is to make the behavior visible rather than hiding it behind one aggregate number.
Ground Truth hallucinations and extra claims
One distinction I wanted to preserve was between two different situations.
A hallucination can occur while handling an expected Ground Truth point.
But a model can also introduce a claim that was never part of the Ground Truth.
DIBER keeps these separate.
This makes it possible to inspect both the expected information structure and additional unsupported claims produced by the model.
Comparison is also part of the evaluation
Another problem appears when comparing two model runs.
If the result changes, what actually changed?
DIBER tracks experimental identity across:
- model
- prompt
- input
- context
and classifies comparisons as:
- replication
- single-variable change
- multidimensional
- indeterminate
A comparison is not automatically considered controlled simply because two runs can be placed next to each other.
Repeated runs
DIBER can also group repeated runs with the same known experimental configuration.
It calculates statistics including:
- sample count
- mean
- minimum
- maximum
- range
- standard deviation
- confidence interval
This makes it possible to examine not only the result of a run, but also the stability of a configuration across repeated runs.
It is designed to be used as software
DIBER Core is available as an npm package:
npm install @anonipro/diber-core
It includes a CLI:
npx diber evaluate \
--truth ./tests/gt.json \
--evals ./tests/runs.json \
--out ./report.json
and deterministic quality gates:
npx diber assert \
--max-hallucination 0.02 \
--min-quality 0.90 \
./report.json
It can also be integrated directly into Node.js applications and test suites, or used in the browser.
The visual layer
I also built a browser-based showcase around the Core.
It takes a synthetic experiment, runs the actual DIBER engine and visualizes the resulting metrics.
The point of the showcase is not to create another dashboard score.
It is to make the underlying evaluation structure visible:
What changed?
Which points changed?
Which failure modes changed?
Was the comparison controlled?
Was the result stable across repeated runs?
https://diber.anonipro.com/en/showcase
Open source
DIBER Core is released under the MIT License.
GitHub:
https://github.com/ANONIPRO/diber-core
npm:
https://www.npmjs.com/package/@anonipro/diber-core
I'm interested in technical feedback, especially around the evaluation methodology, metrics and comparison model.
If you work on LLM evaluation, model testing, QA, benchmarking, AI research or developer tooling, I'd be interested in seeing how you approach the same problem.
Try it, inspect the code, and tell me where the methodology breaks.
Top comments (0)