I build QuizPilot, an open-source Chrome extension that reads practice questions on a web page and suggests answers. Since TypeSafe released Jev, our code sends every multiple-choice and true/false question to it. This post covers how we map quiz questions onto Jev, and what a small benchmark showed.
The short version
- Accuracy: Jev answered 85–86 of 90 questions correctly across two runs. A lightweight LLM got 89, a frontier reasoning LLM got all 90.
- Speed: one Jev request answered 60 questions in 0.17–0.44 s. The lightweight LLM took 2.6 s, the reasoning LLM 8.5 s.
- Where it slips: every true/false and multi-select answer was right. All of its mistakes were single-choice questions that need several steps of arithmetic.
- Confidence is useful: 7 of its 9 wrong answers came back with a confidence below 0.6, against an average of 0.93 for its right answers.
Why a decision model for quiz questions
A multiple-choice question already lists every possible answer. An LLM still writes its reply as text, and we parse the letter out. Jev takes a state and typed questions (pick one of these options, or yes/no), and returns the choice with a probability. No prose to parse, no answer outside the options.
How a quiz question becomes a Jev question
- Single choice → one Jev categorical question; the choices are the options.
-
True/false → one yes/no (
noul) question: "Is this statement true?" - Multiple select → one yes/no question per option ("Is option C one of the correct answers?"). Every option at 0.5 or above is ticked; if none reaches 0.5, the likeliest one is.
- Fill-in, short answer, picture questions → not sent to Jev; it doesn't write text or read images.
A trimmed request for one true/false question:
{
"model": "typesafe/jev-1.13",
"state": "Questions from a quiz page. Each question is independent. Judge by factual correctness, not by wording or position of the options.",
"questions": {
"q0": {
"type": "noul",
"instructions": {
"statement": "The harmonic series 1 + 1/2 + 1/3 + … converges.",
"task": "Is this statement true?",
"criteria": { "true": "The statement is correct", "false": "The statement is incorrect" }
}
}
}
}
A yes/no answer is a single probability, so we turn it into a confidence by its distance from 0.5: |2p − 1|. A multiple-select question's confidence is the lowest confidence among its options. A whole page goes out in one request, split only when it nears Jev's input limit.
The adapter is open source: github.com/jelly-ham/quizpilot-extension.
The benchmark
90 questions with known answers, written by me, none from a real exam:
- Basic set (60): 20 single choice, 20 multi-select, 20 true/false; half easy, half harder.
- Hard set (30): mostly multi-step arithmetic — trailing zeros of 100!, 2^100 mod 7, inclusion–exclusion counting.
Comparison models: a lightweight general LLM with thinking off, and a frontier reasoning LLM at medium effort. Jev and the lightweight LLM ran each set twice.
| Model | Basic (60) | Hard (30) | Total (90) | Latency |
|---|---|---|---|
| Jev | 59, 59 | 26, 27 | 85–86 | 0.17–0.44 s |
| Lightweight LLM | 60 | 29, 29 | 2.6 s |
| Reasoning LLM | 60 | 30 | 8.5 s |
Jev by question type, both runs combined: true/false 40/40, multi-select 50/50, single choice 61/70.
Where Jev gets it wrong
All nine mistakes were single-choice questions whose answer has to be computed, not recognised: trailing zeros of 100!, 2^100 mod 7, the sum of primes from 1 to 100, the number of real roots of x³ − 3x = 1, and (once) the interior angles of a 12-gon. Fact-recall questions, including harder ones, were all right.
Confidence tells you when to re-check
Across 180 Jev answers, 9 were wrong. Seven of the nine came back with confidence between 0.26 and 0.49; the two exceptions were the same question (trailing zeros of 100!) at 0.69 and 0.71. Right answers averaged 0.93.
Re-checking everything below 0.6 would have touched 14 of 180 answers (8%). Using the lightweight LLM's answers from the same benchmark as the fallback, the combination scores 60/60 on the basic set and 28/30 on the hard set. (Combined from separate runs, not a live run with automatic review.)
Limits
- 90 questions we wrote ourselves show large gaps, not small ones.
- Text only; picture questions aren't covered.
- Jev 1.13 via OpenRouter, October 6, 2026.
Full post with more detail: quizpilot.link/en/blog/jev. Happy to answer questions about the adapter or the question set.
Top comments (0)