<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: jack</title>
    <description>The latest articles on DEV Community by jack (@jack9999).</description>
    <link>https://dev.to/jack9999</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4166603%2Fec063292-6035-4c44-aeb5-1d863072bc56.png</url>
      <title>DEV Community: jack</title>
      <link>https://dev.to/jack9999</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/jack9999"/>
    <language>en</language>
    <item>
      <title>How we use Jev to answer quiz questions (a 90-question benchmark)</title>
      <dc:creator>jack</dc:creator>
      <pubDate>Tue, 06 Oct 2026 13:59:50 +0000</pubDate>
      <link>https://dev.to/jack9999/how-we-use-jev-to-answer-quiz-questions-a-90-question-benchmark-7ho</link>
      <guid>https://dev.to/jack9999/how-we-use-jev-to-answer-quiz-questions-a-90-question-benchmark-7ho</guid>
      <description>&lt;p&gt;I build QuizPilot, an open-source Chrome extension that reads practice questions on a web page and suggests answers. Since TypeSafe released Jev, our code sends every multiple-choice and true/false question to it. This post covers how we map quiz questions onto Jev, and what a small benchmark showed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The short version
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Accuracy:&lt;/strong&gt; Jev answered 85–86 of 90 questions correctly across two runs. A lightweight LLM got 89, a frontier reasoning LLM got all 90.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Speed:&lt;/strong&gt; one Jev request answered 60 questions in 0.17–0.44 s. The lightweight LLM took 2.6 s, the reasoning LLM 8.5 s.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Where it slips:&lt;/strong&gt; every true/false and multi-select answer was right. All of its mistakes were single-choice questions that need several steps of arithmetic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Confidence is useful:&lt;/strong&gt; 7 of its 9 wrong answers came back with a confidence below 0.6, against an average of 0.93 for its right answers.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why a decision model for quiz questions
&lt;/h2&gt;

&lt;p&gt;A multiple-choice question already lists every possible answer. An LLM still writes its reply as text, and we parse the letter out. Jev takes a state and typed questions (pick one of these options, or yes/no), and returns the choice with a probability. No prose to parse, no answer outside the options.&lt;/p&gt;

&lt;h2&gt;
  
  
  How a quiz question becomes a Jev question
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Single choice&lt;/strong&gt; → one Jev categorical question; the choices are the options.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;True/false&lt;/strong&gt; → one yes/no (&lt;code&gt;noul&lt;/code&gt;) question: "Is this statement true?"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multiple select&lt;/strong&gt; → one yes/no question per option ("Is option C one of the correct answers?"). Every option at 0.5 or above is ticked; if none reaches 0.5, the likeliest one is.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fill-in, short answer, picture questions&lt;/strong&gt; → not sent to Jev; it doesn't write text or read images.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A trimmed request for one true/false question:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"typesafe/jev-1.13"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"state"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Questions from a quiz page. Each question is independent. Judge by factual correctness, not by wording or position of the options."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"questions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"q0"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"noul"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"instructions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"statement"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"The harmonic series 1 + 1/2 + 1/3 + … converges."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"task"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Is this statement true?"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"criteria"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"true"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"The statement is correct"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"false"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"The statement is incorrect"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A yes/no answer is a single probability, so we turn it into a confidence by its distance from 0.5: &lt;code&gt;|2p − 1|&lt;/code&gt;. A multiple-select question's confidence is the lowest confidence among its options. A whole page goes out in one request, split only when it nears Jev's input limit.&lt;/p&gt;

&lt;p&gt;The adapter is open source: &lt;a href="https://github.com/jelly-ham/quizpilot-extension" rel="noopener noreferrer"&gt;github.com/jelly-ham/quizpilot-extension&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The benchmark
&lt;/h2&gt;

&lt;p&gt;90 questions with known answers, written by me, none from a real exam:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Basic set (60):&lt;/strong&gt; 20 single choice, 20 multi-select, 20 true/false; half easy, half harder.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hard set (30):&lt;/strong&gt; mostly multi-step arithmetic — trailing zeros of 100!, 2^100 mod 7, inclusion–exclusion counting.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Comparison models: a lightweight general LLM with thinking off, and a frontier reasoning LLM at medium effort. Jev and the lightweight LLM ran each set twice.&lt;/p&gt;

&lt;p&gt;| Model | Basic (60) | Hard (30) | Total (90) | Latency |&lt;br&gt;
|---|---|---|---|&lt;br&gt;
| Jev | 59, 59 | 26, 27 | 85–86 | 0.17–0.44 s |&lt;br&gt;
| Lightweight LLM | 60 | 29, 29 | 2.6 s |&lt;br&gt;
| Reasoning LLM | 60 | 30 | 8.5 s |&lt;/p&gt;

&lt;p&gt;Jev by question type, both runs combined: true/false 40/40, multi-select 50/50, single choice 61/70.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Jev gets it wrong
&lt;/h2&gt;

&lt;p&gt;All nine mistakes were single-choice questions whose answer has to be computed, not recognised: trailing zeros of 100!, 2^100 mod 7, the sum of primes from 1 to 100, the number of real roots of x³ − 3x = 1, and (once) the interior angles of a 12-gon. Fact-recall questions, including harder ones, were all right.&lt;/p&gt;

&lt;h2&gt;
  
  
  Confidence tells you when to re-check
&lt;/h2&gt;

&lt;p&gt;Across 180 Jev answers, 9 were wrong. Seven of the nine came back with confidence between 0.26 and 0.49; the two exceptions were the same question (trailing zeros of 100!) at 0.69 and 0.71. Right answers averaged 0.93.&lt;/p&gt;

&lt;p&gt;Re-checking everything below 0.6 would have touched 14 of 180 answers (8%). Using the lightweight LLM's answers from the same benchmark as the fallback, the combination scores 60/60 on the basic set and 28/30 on the hard set. (Combined from separate runs, not a live run with automatic review.)&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;90 questions we wrote ourselves show large gaps, not small ones.&lt;/li&gt;
&lt;li&gt;Text only; picture questions aren't covered.&lt;/li&gt;
&lt;li&gt;Jev 1.13 via OpenRouter, October 6, 2026.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Full post with more detail: &lt;a href="https://quizpilot.link/en/blog/jev/" rel="noopener noreferrer"&gt;quizpilot.link/en/blog/jev&lt;/a&gt;. Happy to answer questions about the adapter or the question set.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>benchmark</category>
      <category>chrome</category>
    </item>
  </channel>
</rss>
