DEV Community

jack
jack

Posted on Originally published at quizpilot.link

How we use Jev to answer quiz questions (a 90-question benchmark)

I build QuizPilot, an open-source Chrome extension that reads practice questions on a web page and suggests answers. Since TypeSafe released Jev, our code sends every multiple-choice and true/false question to it. This post covers how we map quiz questions onto Jev, and what a small benchmark showed.

The short version

  • Accuracy: Jev answered 85–86 of 90 questions correctly across two runs. A lightweight LLM got 89, a frontier reasoning LLM got all 90.
  • Speed: one Jev request answered 60 questions in 0.17–0.44 s. The lightweight LLM took 2.6 s, the reasoning LLM 8.5 s.
  • Where it slips: every true/false and multi-select answer was right. All of its mistakes were single-choice questions that need several steps of arithmetic.
  • Confidence is useful: 7 of its 9 wrong answers came back with a confidence below 0.6, against an average of 0.93 for its right answers.

Why a decision model for quiz questions

A multiple-choice question already lists every possible answer. An LLM still writes its reply as text, and we parse the letter out. Jev takes a state and typed questions (pick one of these options, or yes/no), and returns the choice with a probability. No prose to parse, no answer outside the options.

How a quiz question becomes a Jev question

  • Single choice → one Jev categorical question; the choices are the options.
  • True/false → one yes/no (noul) question: "Is this statement true?"
  • Multiple select → one yes/no question per option ("Is option C one of the correct answers?"). Every option at 0.5 or above is ticked; if none reaches 0.5, the likeliest one is.
  • Fill-in, short answer, picture questions → not sent to Jev; it doesn't write text or read images.

A trimmed request for one true/false question:

{
  "model": "typesafe/jev-1.13",
  "state": "Questions from a quiz page. Each question is independent. Judge by factual correctness, not by wording or position of the options.",
  "questions": {
    "q0": {
      "type": "noul",
      "instructions": {
        "statement": "The harmonic series 1 + 1/2 + 1/3 + … converges.",
        "task": "Is this statement true?",
        "criteria": { "true": "The statement is correct", "false": "The statement is incorrect" }
      }
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

A yes/no answer is a single probability, so we turn it into a confidence by its distance from 0.5: |2p − 1|. A multiple-select question's confidence is the lowest confidence among its options. A whole page goes out in one request, split only when it nears Jev's input limit.

The adapter is open source: github.com/jelly-ham/quizpilot-extension.

The benchmark

90 questions with known answers, written by me, none from a real exam:

  • Basic set (60): 20 single choice, 20 multi-select, 20 true/false; half easy, half harder.
  • Hard set (30): mostly multi-step arithmetic — trailing zeros of 100!, 2^100 mod 7, inclusion–exclusion counting.

Comparison models: a lightweight general LLM with thinking off, and a frontier reasoning LLM at medium effort. Jev and the lightweight LLM ran each set twice.

| Model | Basic (60) | Hard (30) | Total (90) | Latency |
|---|---|---|---|
| Jev | 59, 59 | 26, 27 | 85–86 | 0.17–0.44 s |
| Lightweight LLM | 60 | 29, 29 | 2.6 s |
| Reasoning LLM | 60 | 30 | 8.5 s |

Jev by question type, both runs combined: true/false 40/40, multi-select 50/50, single choice 61/70.

Where Jev gets it wrong

All nine mistakes were single-choice questions whose answer has to be computed, not recognised: trailing zeros of 100!, 2^100 mod 7, the sum of primes from 1 to 100, the number of real roots of x³ − 3x = 1, and (once) the interior angles of a 12-gon. Fact-recall questions, including harder ones, were all right.

Confidence tells you when to re-check

Across 180 Jev answers, 9 were wrong. Seven of the nine came back with confidence between 0.26 and 0.49; the two exceptions were the same question (trailing zeros of 100!) at 0.69 and 0.71. Right answers averaged 0.93.

Re-checking everything below 0.6 would have touched 14 of 180 answers (8%). Using the lightweight LLM's answers from the same benchmark as the fallback, the combination scores 60/60 on the basic set and 28/30 on the hard set. (Combined from separate runs, not a live run with automatic review.)

Limits

  • 90 questions we wrote ourselves show large gaps, not small ones.
  • Text only; picture questions aren't covered.
  • Jev 1.13 via OpenRouter, October 6, 2026.

Full post with more detail: quizpilot.link/en/blog/jev. Happy to answer questions about the adapter or the question set.

Top comments (0)