DEV Community

Cover image for I built a deterministic LLM evaluation engine without an LLM judge
Artur Woszczyk
Artur Woszczyk

Posted on

I built a deterministic LLM evaluation engine without an LLM judge

I built a deterministic LLM evaluation engine without an LLM judge

Most current LLM evaluation workflows eventually face a version of the same problem:

How do you evaluate the output of a model?

One common approach is to use another LLM as the evaluator.

That can be useful, but it also introduces another model, another source of variability, and another layer of judgment into the evaluation process.

I wanted to explore a different approach.

So I built DIBER Core.

What is DIBER?

DIBER is an open-source deterministic evaluation engine for LLM outputs.

It does not ask an LLM to judge another LLM.

Instead, it works from an explicit evaluation structure:

LLM run → Ground Truth → human-defined classifications → deterministic engine → audit report

The engine performs the normalization, weighting, comparison and statistical calculations.

The same structured inputs produce the same report.

It does not reduce everything to one score

A model can cover a large amount of information while still producing significant errors.

It can also provide a technically correct answer while omitting important points.

So DIBER keeps different dimensions separate.

It evaluates things such as:

  • Coverage
  • Quality
  • Usable information
  • Effective information
  • Omissions
  • Errors
  • Mathematical errors
  • Hallucinations
  • Overclaims
  • Inferences
  • Extra unsupported claims
  • Relation consistency

The goal is to make the behavior visible rather than hiding it behind one aggregate number.

Ground Truth hallucinations and extra claims

One distinction I wanted to preserve was between two different situations.

A hallucination can occur while handling an expected Ground Truth point.

But a model can also introduce a claim that was never part of the Ground Truth.

DIBER keeps these separate.

This makes it possible to inspect both the expected information structure and additional unsupported claims produced by the model.

Comparison is also part of the evaluation

Another problem appears when comparing two model runs.

If the result changes, what actually changed?

DIBER tracks experimental identity across:

  • model
  • prompt
  • input
  • context

and classifies comparisons as:

  • replication
  • single-variable change
  • multidimensional
  • indeterminate

A comparison is not automatically considered controlled simply because two runs can be placed next to each other.

Repeated runs

DIBER can also group repeated runs with the same known experimental configuration.

It calculates statistics including:

  • sample count
  • mean
  • minimum
  • maximum
  • range
  • standard deviation
  • confidence interval

This makes it possible to examine not only the result of a run, but also the stability of a configuration across repeated runs.

It is designed to be used as software

DIBER Core is available as an npm package:

npm install @anonipro/diber-core
Enter fullscreen mode Exit fullscreen mode

It includes a CLI:

npx diber evaluate \
  --truth ./tests/gt.json \
  --evals ./tests/runs.json \
  --out ./report.json
Enter fullscreen mode Exit fullscreen mode

and deterministic quality gates:

npx diber assert \
  --max-hallucination 0.02 \
  --min-quality 0.90 \
  ./report.json
Enter fullscreen mode Exit fullscreen mode

It can also be integrated directly into Node.js applications and test suites, or used in the browser.

The visual layer

I also built a browser-based showcase around the Core.

It takes a synthetic experiment, runs the actual DIBER engine and visualizes the resulting metrics.

The point of the showcase is not to create another dashboard score.

It is to make the underlying evaluation structure visible:

What changed?

Which points changed?

Which failure modes changed?

Was the comparison controlled?

Was the result stable across repeated runs?

https://diber.anonipro.com/en/showcase

Open source

DIBER Core is released under the MIT License.

GitHub:

https://github.com/ANONIPRO/diber-core

npm:

https://www.npmjs.com/package/@anonipro/diber-core

I'm interested in technical feedback, especially around the evaluation methodology, metrics and comparison model.

If you work on LLM evaluation, model testing, QA, benchmarking, AI research or developer tooling, I'd be interested in seeing how you approach the same problem.

Try it, inspect the code, and tell me where the methodology breaks.

Top comments (0)