How EvalPort's Grader System Works
When designing EvalPort, the grader system was the hardest part to get right. Every eval framework has its own way of scoring LLM outputs — DeepEval uses metric classes, Promptfoo uses assertion objects, Inspect AI uses solver functions. We needed a system expressive enough to cover 90%+ of real-world eval needs, but simple enough that any framework could implement it.
The result: 11 grader types that carry their own semantics. A grader isn't just a name — it specifies its parameters, its model, its threshold. An eval suite is self-describing.
The 11 Grader Types
exact_match — Compare output to expected output, optionally ignoring case.
contains — Check if the output contains a substring.
regex — Match against a regular expression.
semantic_similarity — Embed output and expected output, compare cosine similarity against a threshold.
llm_judge — Use an LLM to evaluate the output against a prompt template. The most powerful grader.
json_schema — Validate that the output is valid JSON matching a JSON Schema.
json_path — Extract a value from JSON output using a JSONPath expression, then compare it.
code — Run a function to evaluate the output.
human — Defer to human review.
model_graded — Compare the output to a reference answer using a model.
custom — Escape hatch for graders not covered by built-in types.
How Graders Connect to Test Cases
A test case references graders by ID. Multiple graders can evaluate the same test case. The ResultSet records each grader's score separately.
Why This Design Works
Self-describing: An eval suite carries everything a framework needs to execute it.
Framework-agnostic: Any framework can implement any subset of grader types.
Extensible: The custom type lets frameworks bring their own graders.
Comparable: Results from different frameworks use the same grader IDs.
Try It
pip install evalport-sdk
npm install evalport-sdk
Spec: https://github.com/adhabnr-ux/evalport/blob/main/spec/SPEC.md
Top comments (0)