<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Krzysztof Głuszczyk</title>
    <description>The latest articles on DEV Community by Krzysztof Głuszczyk (@kgluszczyk).</description>
    <link>https://dev.to/kgluszczyk</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4140287%2Fa52e662c-39ac-4c97-b206-db8baceee595.png</url>
      <title>DEV Community: Krzysztof Głuszczyk</title>
      <link>https://dev.to/kgluszczyk</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/kgluszczyk"/>
    <language>en</language>
    <item>
      <title>A DIY Jev: typed decisions from gpt-6-luna logprobs, benchmarked head-to-head</title>
      <dc:creator>Krzysztof Głuszczyk</dc:creator>
      <pubDate>Mon, 28 Sep 2026 15:42:39 +0000</pubDate>
      <link>https://dev.to/kgluszczyk/a-diy-jev-typed-decisions-from-gpt-6-luna-logprobs-benchmarked-head-to-head-1fd8</link>
      <guid>https://dev.to/kgluszczyk/a-diy-jev-typed-decisions-from-gpt-6-luna-logprobs-benchmarked-head-to-head-1fd8</guid>
      <description>&lt;p&gt;What if an ordinary LLM could return this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"sales"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.89&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"billing"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.07&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"support"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.04&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;...not by asking the model to write JSON numbers, but by turning its token logprobs into a typed, calibrated probability distribution?&lt;/p&gt;

&lt;p&gt;That is roughly the promise of TypeSafe's Jev: purpose-built models returning probabilistic decisions instead of free-form text. We wanted to see how much of that experience we could reproduce using &lt;code&gt;gpt-6-luna&lt;/code&gt;, already available in our stack, before evaluating dedicated commercial endpoints.&lt;/p&gt;

&lt;p&gt;So we built the smallest prototype that could work: restrict the output to single-token keys, read their logprobs, normalize them, and fit a single temperature parameter on labelled examples.&lt;/p&gt;

&lt;p&gt;Then we tested it head-to-head against Jev across 6,762 public examples (13 datasets).&lt;/p&gt;

&lt;p&gt;The result was more nuanced than expected. Our DIY classifier nearly matched Jev on most category tasks (even with 219 classes), while Jev held a consistent lead of 3.6 to 6.1 points on yes/no questions and rating scales. It was cheaper for short inputs, roughly 4-5x slower, and heavily overconfident until calibrated.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why not just ask the model for JSON?
&lt;/h3&gt;

&lt;p&gt;The obvious alternative is prompting the model to output a JSON dictionary with confidence scores. We benchmarked that as a third approach on &lt;code&gt;gpt-6-luna&lt;/code&gt; across the same evaluation suite:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Approach&lt;/th&gt;
&lt;th&gt;Average accuracy&lt;/th&gt;
&lt;th&gt;Latency (mean)&lt;/th&gt;
&lt;th&gt;Billed output tokens&lt;/th&gt;
&lt;th&gt;Calibration error (ECE ↓)&lt;/th&gt;
&lt;th&gt;Practical role&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Prompted JSON&lt;/td&gt;
&lt;td&gt;70.1%&lt;/td&gt;
&lt;td&gt;10.06s&lt;/td&gt;
&lt;td&gt;~960&lt;/td&gt;
&lt;td&gt;0.175&lt;/td&gt;
&lt;td&gt;Slow, high-token prototype&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DIY logprob adapter&lt;/td&gt;
&lt;td&gt;73.8%&lt;/td&gt;
&lt;td&gt;2.11s&lt;/td&gt;
&lt;td&gt;1 (answer token)&lt;/td&gt;
&lt;td&gt;0.211 raw (≤0.10 calibrated)&lt;/td&gt;
&lt;td&gt;No training, fast to ship&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dedicated model (Jev)&lt;/td&gt;
&lt;td&gt;76.5%&lt;/td&gt;
&lt;td&gt;0.38s&lt;/td&gt;
&lt;td&gt;0 (free output)&lt;/td&gt;
&lt;td&gt;0.110&lt;/td&gt;
&lt;td&gt;Dedicated specialized model&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;/p&gt;
  What is calibration error (ECE)?
  &lt;br&gt;
Expected Calibration Error (ECE) measures how honest the model is about its confidence numbers. If a model says it is 90% sure across many predictions, it should be right about 9 out of 10 times.

&lt;p&gt;Raw language models are notoriously overconfident (claiming 99%+ certainty even when guessing wrong). Lower ECE means the probabilities reflect actual accuracy better (0.0 means perfect calibration).&lt;br&gt;
&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;

&lt;p&gt;Across the entire benchmark, prompted JSON was 4.8 times slower than the logprob adapter and 26.5 times slower than Jev. It generated ~6.5 million output tokens across the benchmark. The DIY adapter generated under 15,000 (even with tournament calls).&lt;/p&gt;

&lt;p&gt;JSON was not consistently inaccurate. It reached the highest accuracy on five datasets and tied Jev on Banking77 (80.6%). The failure was variance. It collapsed on two large label tasks:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dataset&lt;/th&gt;
&lt;th&gt;Classes&lt;/th&gt;
&lt;th&gt;Jev&lt;/th&gt;
&lt;th&gt;DIY&lt;/th&gt;
&lt;th&gt;JSON&lt;/th&gt;
&lt;th&gt;JSON latency&lt;/th&gt;
&lt;th&gt;JSON tokens&lt;/th&gt;
&lt;th&gt;JSON ECE&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;TREC-50&lt;/td&gt;
&lt;td&gt;50&lt;/td&gt;
&lt;td&gt;83.2%&lt;/td&gt;
&lt;td&gt;82.4%&lt;/td&gt;
&lt;td&gt;38.4%&lt;/td&gt;
&lt;td&gt;11.4s&lt;/td&gt;
&lt;td&gt;986&lt;/td&gt;
&lt;td&gt;0.574&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DBpedia219&lt;/td&gt;
&lt;td&gt;219&lt;/td&gt;
&lt;td&gt;93.2%&lt;/td&gt;
&lt;td&gt;90.7%&lt;/td&gt;
&lt;td&gt;55.3%&lt;/td&gt;
&lt;td&gt;52.6s&lt;/td&gt;
&lt;td&gt;5,751&lt;/td&gt;
&lt;td&gt;0.427&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;h3&gt;
  
  
  Why not just put 77 or 219 classes in one prompt if you have a large context window?
&lt;/h3&gt;

&lt;p&gt;A large context window solves the input problem, but not the decoding or latency problem:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;For JSON (generation bottleneck):&lt;/strong&gt; Fitting 77 or 219 label descriptions into the prompt is easy. But asking the model to &lt;em&gt;generate&lt;/em&gt; a JSON dictionary of confidence scores for 219 classes burns ~5,750 output tokens and took &lt;strong&gt;52.6 seconds per item&lt;/strong&gt;. In addition, when generating float numbers in text, models produce uncalibrated scores that distort ranking.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;For logprobs (candidate reporting limit):&lt;/strong&gt; With logprobs, the model scores options in a single decoding step. However, OpenAI and most compatible APIs return candidate probabilities for at most 20 tokens per position (&lt;code&gt;top_logprobs=20&lt;/code&gt;). Options beyond the top 20 read as zero. &lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;You &lt;em&gt;can&lt;/em&gt; put all 77 or 219 labels in a single logprob prompt (we tested that too), and it performed surprisingly well (within 0.3 points overall). But chunking into groups of up to 19 gave more consistent accuracy across our full benchmark suite (particularly on TREC fine).&lt;/p&gt;

&lt;p&gt;Prompted JSON is fine for a quick one-off prototype. For a reliable decision endpoint, the 1-token logprob adapter is faster, cheaper, and far more stable across label sizes.&lt;/p&gt;

&lt;p&gt;Here is how the trick works, where it breaks, and what to watch out for if you build your own.&lt;/p&gt;
&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Classification without fine-tuning
&lt;/h3&gt;

&lt;p&gt;Think of it as a text classifier you don't have to train. The traditional way to sort text into categories is to train or fine-tune an encoder like BERT or ModernBERT on thousands of labelled examples (recent open-source projects like &lt;a href="https://github.com/NandhaKishorM/laya" rel="noopener noreferrer"&gt;Laya&lt;/a&gt; explore this non-autoregressive encoder route). That works well when label sets are fixed and small, but requires task-specific training pipelines.&lt;/p&gt;

&lt;p&gt;Here, the general-purpose language model's own next-token probabilities do that classification job zero-shot. There is nothing to train, and the label set can change dynamically on every request. What you give up out of the box is calibrated confidence: raw LLM probabilities run high, so claiming "99.9% certainty" on an ambiguous edge case is common. The "Confidence" section below shows how to fix it with a simple temperature scaling parameter.&lt;/p&gt;

&lt;p&gt;The prototype wraps this in one endpoint that takes several questions about the same text, following TypeSafe's Primitives spec: &lt;code&gt;choice&lt;/code&gt; (pick a label), &lt;code&gt;score&lt;/code&gt; (pick a level on an ordered scale) and &lt;code&gt;noul&lt;/code&gt; (yes/no, one probability). Each answer also reports &lt;code&gt;coverage&lt;/code&gt;, the share of probability that landed on your options at all. Low coverage means the model wanted to say something else, so don't trust that answer.&lt;/p&gt;
&lt;h3&gt;
  
  
  The 1-token logprob trick
&lt;/h3&gt;

&lt;p&gt;A language model writes one token (a word or a piece of one) at a time. At each step it gives every token in its vocabulary a probability, then picks one. Many APIs can also return the top candidates it was choosing between, each with its probability, as a logarithm ("logprob").&lt;/p&gt;

&lt;p&gt;That is all we need. List the options with a short key each (0, 1, 2...), ask the model to answer with just the key, and read the probabilities the API returns for those keys. Keep the candidates that are keys, turn the logprobs back into probabilities, rescale them to add up to 1, and you have a probability for every option.&lt;/p&gt;

&lt;p&gt;The API returns at most 20 candidates per token, and many of them are junk: "A", "**", an end-of-text marker, or the same key twice (" B" and "b"). So the least likely options often don't make the list and read as zero. That costs almost nothing: anything left off ranks below the last candidate shown, typically with negligible individual probability. We stop at 19 options per call (details below).&lt;/p&gt;

&lt;p&gt;Three steps, in Python. The full example runs against any OpenAI-compatible endpoint that serves &lt;code&gt;gpt-6-luna&lt;/code&gt; on the Responses API, and the outputs shown are from a live call.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Build the prompt.&lt;/strong&gt; Each option gets a one-token key: digits up to 10 options, letters past that. Why the answer looks like this comes right after the code.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;math&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;  &lt;span class="c1"&gt;# any OpenAI-compatible endpoint; pass base_url=... for a gateway
&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;build_prompt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;options: {label: description}. Returns the answer keys and the prompt.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="c1"&gt;# Keys are 0-9 up to 10 options. Past that, "1" would also be the start of
&lt;/span&gt;    &lt;span class="c1"&gt;# "10".."18", so use letters, skipping "a" and "i" (they are also English words).
&lt;/span&gt;    &lt;span class="n"&gt;keys&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="nf"&gt;list&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;bcdefghjklmnopqrstu&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)[:&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;listing&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;upper&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;label&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; - &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;desc&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;label&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;desc&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;zip&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;keys&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;()))&lt;/span&gt;
    &lt;span class="n"&gt;word&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;number&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;letter&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;prompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;State: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ensure_ascii&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
              &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Question: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="s"&gt;Options:&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;listing&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
              &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Answer with only the &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;word&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; of the best option. No other text.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;keys&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;strong&gt;2. Ask for one short answer, with its candidates.&lt;/strong&gt; &lt;code&gt;top_p=1.0&lt;/code&gt; matters: with the default, the API leaves out the options the model considered unlikely, and they come back as exactly zero.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;ask&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;responses&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpt-6-luna&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="nb"&gt;input&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;reasoning&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;effort&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;none&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="n"&gt;max_output_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;64&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;   &lt;span class="c1"&gt;# reasoning and output share this budget
&lt;/span&gt;        &lt;span class="n"&gt;top_logprobs&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;        &lt;span class="c1"&gt;# the most candidates the API returns per token
&lt;/span&gt;        &lt;span class="n"&gt;top_p&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;1.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;              &lt;span class="c1"&gt;# the default 0.98 hides the runner-up options
&lt;/span&gt;        &lt;span class="n"&gt;include&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;message.output_text.logprobs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="n"&gt;store&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;msg&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;next&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;o&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;o&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;output&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;o&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;type&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;message&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;logprobs&lt;/span&gt;  &lt;span class="c1"&gt;# one entry per output token, with its top candidates
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;strong&gt;3. Turn the candidates into probabilities.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;to_probs&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tokens&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;keys&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;labels&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;calibration_t&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;1.0&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;norm&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;lambda&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;.,;:!?&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s"&gt;`*&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;pos&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;tokens&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;  &lt;span class="c1"&gt;# the first output token that is an answer key
&lt;/span&gt;        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;norm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pos&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;keys&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;mass&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{}&lt;/span&gt;
            &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;alt&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;pos&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;top_logprobs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;k&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;norm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;alt&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;keys&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="n"&gt;mass&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;mass&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exp&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;alt&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;logprob&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# logprob -&amp;gt; probability
&lt;/span&gt;            &lt;span class="n"&gt;mass&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;calibration_t&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;mass&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;()}&lt;/span&gt;  &lt;span class="c1"&gt;# see "Confidence"
&lt;/span&gt;            &lt;span class="n"&gt;total&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;mass&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;values&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;l&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;mass&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;total&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;l&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;zip&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;keys&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;labels&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;  &lt;span class="c1"&gt;# no answer key anywhere: the model replied off-format
&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;classify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;calibration_t&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;1.0&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;keys&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;build_prompt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;to_probs&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;ask&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;keys&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;list&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;calibration_t&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;opts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sales&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;New subscriptions, upgrades, enterprise plans&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;billing&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Invoices, payment methods, receipts&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;support&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Technical help, bugs, troubleshooting&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;classify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;We want to upgrade our workspace to the annual enterprise plan.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Which team should handle this?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;opts&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="c1"&gt;# {'sales': 0.9999999999862318, 'billing': 6.755284692114235e-12, 'support': 7.0129827926438374e-12}
&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;classify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;We want to upgrade our workspace to the annual enterprise plan.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Which team should handle this?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;opts&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
               &lt;span class="n"&gt;calibration_t&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="c1"&gt;# {'sales': 0.9731562277998806, 'billing': 0.013380012188266373, 'support': 0.013463760011852986}
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;h3&gt;
  
  
  Why the answer is a single key
&lt;/h3&gt;

&lt;p&gt;The API gives probabilities for single tokens, and it only shows candidates along the path the model actually took. So each option has to be exactly one token, and no option can look like the start of another. Label names often fail that test: "customer support" and "customer sales" both start with "customer", so the first step can't tell them apart. A one-token key per option puts every option side by side in a single step.&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;
  More on the answer format, and how often it failed
  &lt;p&gt;&lt;strong&gt;Digits up to 10, letters past that.&lt;/strong&gt; "0" to "9" are single tokens. Two-digit numbers usually are too: in a live check with 12 options the model answered "11" as one token, with "1" far behind. But a "1" can also be the start of a longer answer, and in 1 of 54 early test cases the probability did land on option 1 when the right answer was 11 or above. Letters never overlap like that.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why 19?&lt;/strong&gt; We first reserved one of the 20 slots for junk. Measured later on 100 Banking77 questions with 19 options, that reasoning doesn't hold: junk took a median of 9 slots and duplicate keys 2 more, and some lists came back shorter than 20, so a median of 10 options were missing from the list; 20 options looked the same. No number of options fits on the list every time; what matters is that only negligible options fall off. 19 is simply where our clean one-letter keys run out. Larger single calls are covered under "More than 19 labels".&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Letters start at B and skip I&lt;/strong&gt;, because "A" and "I" also open ordinary sentences. If the model ignores the format and replies "I think...", we want no valid key in the reply, so the question comes back as "no answer" instead of a confident vote for option I. B to U without I gives exactly 19.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why not force a one-token reply?&lt;/strong&gt; The API rejects &lt;code&gt;max_output_tokens&lt;/code&gt; below 16, reasoning shares that budget, and the model often adds a full stop anyway. So we take the first output token that is a key, and return "no answer" if there is none.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why not structured outputs or &lt;code&gt;logit_bias&lt;/code&gt;?&lt;/strong&gt; We tried both, hoping to limit the candidates to our keys. Neither did. With a JSON schema and an &lt;code&gt;enum&lt;/code&gt;, the candidate list was still the model's raw top 20, junk included, and through our setup the schema wasn't even enforced when the prompt asked for a bare letter (3 of 3 calls came back as plain "D"). &lt;code&gt;logit_bias&lt;/code&gt; was accepted but had no visible effect, even at the maximum.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How often it went wrong.&lt;/strong&gt; In the 13-set benchmark, 3 of 6,762 questions (all on Yelp) came back with no valid key and counted as wrong. What we can't count is a prose reply that happens to start with a valid key, which is why A and I are left out. They are real risks: "A" showed up among the candidates in 97 of 100 test questions and "I" in 42, as junk.&lt;/p&gt;



&lt;p&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  How it compares with Jev
&lt;/h2&gt;

&lt;p&gt;We ran both on 13 public datasets, 6,762 cases: seven classification sets with 4 to 219 labels, two yes/no sets and four rating scales. Every case went to both with the same question and label descriptions. We call a gap real only when a statistical test says it is unlikely to be luck. Jev was called through Cloudflare Workers AI on 2026-09-26.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fof11aprtwicz1stfsfxb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fof11aprtwicz1stfsfxb.png" alt="accuracy of DIY and Jev on 13 datasets grouped as classify, yes/no and score, with the difference per row. Classification rows are within about a point except DBpedia at -2.4; yes/no and score rows are 3.6 to 6.1 points behind Jev." width="800" height="640"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Picking a category: about even
&lt;/h3&gt;

&lt;p&gt;Within 1.6 percentage points on six category datasets (including CLINC150 with 150 intents). The one noticeable gap was DBpedia with 219 categories, where Jev led by 2.4 points. Jev is slightly ahead on most sets, just not by enough to tell apart from noise.&lt;/p&gt;
&lt;h3&gt;
  
  
  Yes/no and ratings: Jev ahead by 3.6 to 6.1 points
&lt;/h3&gt;

&lt;p&gt;Jev leads by 3.6 points on BoolQ and 4.0 on RTE. Both gaps are real.&lt;/p&gt;
&lt;h3&gt;
  
  
  Rating scales: Jev ahead, mostly by one step
&lt;/h3&gt;

&lt;p&gt;Jev leads by 4.6 to 6.1 points on exact hits. Most of our misses are one level off: counting "within one level" as right, the two are close (97.8% for both on SST-5, 96.2% vs 98.2% on Yelp, among the cases both answered).&lt;/p&gt;

&lt;p&gt;Scales are also where ours goes further. Jev takes up to 10 levels, ours up to 19. On 342 wine reviews split into 2 to 19 score bands, ours sits a few points under Jev up to 10 levels. At 19 levels it is off by 1.8 rating points on average, a little better than at 10 levels (2.1; Jev at 10: 2.0).&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb8nhy539q031cpim2yml.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb8nhy539q031cpim2yml.png" alt="line chart of exact and within-one-level accuracy against 2, 4, 8, 10 and 19 score levels on 342 wine reviews. Both engines fall from 90% exact at 2 levels to about 30% at 10, with Jev a few points ahead; only DIY continues to 19 levels, at 19% exact and 49% within one level." width="800" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Across all 13 sets ours trails by 2.7 points on average. It also naturally inherits the large context window of the underlying general-purpose model, accepting inputs well beyond short classifications. We did not compare long inputs against Jev.&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;
  Full results, statistics and caveats
  &lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Test (labels, cases)&lt;/th&gt;
&lt;th&gt;DIY&lt;/th&gt;
&lt;th&gt;Jev&lt;/th&gt;
&lt;th&gt;Difference, 95% CI&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;AG News (4, 500)&lt;/td&gt;
&lt;td&gt;83.8%&lt;/td&gt;
&lt;td&gt;85.4%&lt;/td&gt;
&lt;td&gt;−1.6 [−3.4, 0.0]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Emotion (6, 566)&lt;/td&gt;
&lt;td&gt;51.8%&lt;/td&gt;
&lt;td&gt;50.2%&lt;/td&gt;
&lt;td&gt;+1.6 [−1.1, +4.2]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TREC question type (6, 500)&lt;/td&gt;
&lt;td&gt;90.6%&lt;/td&gt;
&lt;td&gt;91.6%&lt;/td&gt;
&lt;td&gt;−1.0 [−2.8, +0.6]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TREC fine (42, 500)&lt;/td&gt;
&lt;td&gt;82.4%&lt;/td&gt;
&lt;td&gt;83.2%&lt;/td&gt;
&lt;td&gt;−0.8 [−3.6, +2.0]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Banking77 (77, 770)&lt;/td&gt;
&lt;td&gt;79.2%&lt;/td&gt;
&lt;td&gt;80.6%&lt;/td&gt;
&lt;td&gt;−1.4 [−3.4, +0.5]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CLINC150 (150, 750)&lt;/td&gt;
&lt;td&gt;91.7%&lt;/td&gt;
&lt;td&gt;91.5%&lt;/td&gt;
&lt;td&gt;+0.3 [−1.6, +2.1]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DBpedia (219, 657)&lt;/td&gt;
&lt;td&gt;90.7%&lt;/td&gt;
&lt;td&gt;93.2%&lt;/td&gt;
&lt;td&gt;−2.4 [−4.0, −0.9]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BoolQ (yes/no, 500)&lt;/td&gt;
&lt;td&gt;89.2%&lt;/td&gt;
&lt;td&gt;92.8%&lt;/td&gt;
&lt;td&gt;−3.6 [−5.8, −1.6]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RTE (yes/no, 277)&lt;/td&gt;
&lt;td&gt;87.7%&lt;/td&gt;
&lt;td&gt;91.7%&lt;/td&gt;
&lt;td&gt;−4.0 [−7.6, −0.4]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SST-5 sentiment (5, 500)&lt;/td&gt;
&lt;td&gt;53.4%&lt;/td&gt;
&lt;td&gt;58.0%&lt;/td&gt;
&lt;td&gt;−4.6 [−8.2, −1.0]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Yelp stars (5, 500)&lt;/td&gt;
&lt;td&gt;62.8%&lt;/td&gt;
&lt;td&gt;68.8%&lt;/td&gt;
&lt;td&gt;−6.0 [−9.8, −2.2]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;IMDB rating (4, 400)&lt;/td&gt;
&lt;td&gt;67.8%&lt;/td&gt;
&lt;td&gt;72.8%&lt;/td&gt;
&lt;td&gt;−5.0 [−8.8, −1.0]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Wine score (10, 342)&lt;/td&gt;
&lt;td&gt;28.1%&lt;/td&gt;
&lt;td&gt;34.2%&lt;/td&gt;
&lt;td&gt;−6.1 [−12.9, +0.9]&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;How we tested.&lt;/strong&gt; Equal numbers of cases per label, so these are benchmark accuracies, not what you'd see on real traffic. The interval next to each difference is a 95% paired bootstrap: we resample the same cases many times, and a gap counts as real when the interval excludes zero. By that test, six of the 13 gaps are real, all in Jev's favour. Testing 13 datasets at once raises the odds that one looks real by luck, so we also applied a stricter bar (McNemar's test with a Holm correction); three gaps pass it: BoolQ, Yelp and DBpedia.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The 2.7-point average.&lt;/strong&gt; 95% CI 1.7 to 3.6, bootstrapping cases within each dataset as planned; 1.2 to 4.1 if you treat each dataset as a sample from a wider pool of tasks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Failed answers count as wrong.&lt;/strong&gt; Refused or invalid answers (less than 1% across the entire benchmark, mostly safety filter triggers on film reviews) were counted as misses. Counting those as misses, Jev keeps a 2-3 point lead on Yelp and IMDB even on "within one level". Jev answered everything.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What these numbers can't tell you:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;All 13 are well-known public benchmarks, so both models have probably seen them in training; Banking77 and CLINC150 are standard intent-training sets. Your own data is the real test.&lt;/li&gt;
&lt;li&gt;Six of the sets (AG News, BoolQ, SST-5, IMDB, Banking77, wine) were also used, on other samples, while we tuned the letters, the chunking, &lt;code&gt;top_p&lt;/code&gt; and calibration. The seven sets we hadn't touched show the same pattern.&lt;/li&gt;
&lt;li&gt;One run per engine per set. Five sets got a second DIY run by other means; none moved by more than 0.6 points. The wine re-run gave 28.7% vs 33.9% at 10 levels, against 28.1% vs 34.2% in the table.&lt;/li&gt;
&lt;li&gt;Jev ran through Cloudflare, not TypeSafe directly, and we can't verify it's the same build.&lt;/li&gt;
&lt;li&gt;The long-text test hid a code in filler text, 6 trials per size.&lt;/li&gt;
&lt;/ul&gt;



&lt;br&gt;
&lt;p&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  More than 19 labels
&lt;/h2&gt;

&lt;p&gt;One call holds 19 options. Past that, what worked was splitting the labels into groups of up to 19, asking about every group in parallel, then asking once more among the group winners:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;concurrent.futures&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;ThreadPoolExecutor&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;classify_many&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Past 19 labels: chunks of up to 19 in parallel, then a final round among the winners.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;labels&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;list&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;labels&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="mi"&gt;19&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;classify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;labels&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;//&lt;/span&gt; &lt;span class="mi"&gt;19&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;chunks&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="n"&gt;l&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;l&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;l&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;labels&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;]}&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt;
    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nc"&gt;ThreadPoolExecutor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;pool&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;list&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pool&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;lambda&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;classify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="n"&gt;winners&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;get&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;results&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;  &lt;span class="c1"&gt;# a refused chunk drops out here
&lt;/span&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;winners&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;classify_many&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;w&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;w&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;w&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;winners&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is how the 42- to 219-label rows were produced: 4 calls for 42 labels, 9 for 150, 13 for 219, in 3 to 4 seconds.&lt;/p&gt;

&lt;p&gt;Since options that don't make the candidate list barely matter, we also tried alternatives: all labels in one call (keyed by 3-digit codes, each a single token), bigger chunks of up to about 40, and chunks of 19 that keep similar labels together. Each ran twice on the same four datasets, against a same-day rerun of the baseline setup. None matched it. The single call came closest: 3 to 4 times faster, 30 to 55% cheaper, and even on Banking77, CLINC150 and DBpedia, but 4 points lower on TREC fine, where it confused the broad category (a place for a number, say) twice as often. Bigger chunks and grouped chunks lost on TREC fine too, and a little overall. So chunks of 19, with labels spread across them, remain the baseline.&lt;/p&gt;

&lt;p&gt;To see what the number of labels does on its own, we took random subsets of CLINC150's intents, from 5 to 150, with the same cases for both. Both stay near 100% up to 20 labels. At 40 and 80 ours trails by 3.4 and 2.5 points, inside the noise, and at 150 they tie.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvv78hdcqkv2nu5t3wdbf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvv78hdcqkv2nu5t3wdbf.png" alt="accuracy of DIY and Jev on random subsets of 5, 10, 20, 40, 80 and 150 CLINC150 intents. Both are near 100% up to 20 classes; at 40 and 80 DIY is 3.4 and 2.5 points behind, within the 95% bands; at 150 both are 91.7%." width="800" height="496"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The obvious alternative, walking a category tree (pick the group first, then the label inside it), was clearly worse: 4 to 29 points behind Jev.&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;
  The alternatives in numbers
  &lt;p&gt;Same pre-registered cases as the table above, prediction = most probable option in every arm, each alternative run twice, compared with a same-day rerun of the baseline setup (which matched the earlier run within 0.6 points). Differences are in points with 95% intervals; both runs are averaged per case.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Setup&lt;/th&gt;
&lt;th&gt;TREC fine (42)&lt;/th&gt;
&lt;th&gt;Banking77 (77)&lt;/th&gt;
&lt;th&gt;CLINC150 (150)&lt;/th&gt;
&lt;th&gt;DBpedia (219)&lt;/th&gt;
&lt;th&gt;All four&lt;/th&gt;
&lt;th&gt;Cost vs baseline&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Baseline: chunks of ≤19, labels spread&lt;/td&gt;
&lt;td&gt;83.0%&lt;/td&gt;
&lt;td&gt;79.0%&lt;/td&gt;
&lt;td&gt;91.7%&lt;/td&gt;
&lt;td&gt;90.4%&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One call, all labels, sorted&lt;/td&gt;
&lt;td&gt;−4.3 [−7.1, −1.5]&lt;/td&gt;
&lt;td&gt;−0.1&lt;/td&gt;
&lt;td&gt;+0.9&lt;/td&gt;
&lt;td&gt;+1.1&lt;/td&gt;
&lt;td&gt;−0.3 [−1.3, +0.7]&lt;/td&gt;
&lt;td&gt;31-56% cheaper, ~1.1 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chunks of 21-39&lt;/td&gt;
&lt;td&gt;−5.7 [−8.9, −2.5]&lt;/td&gt;
&lt;td&gt;+0.1&lt;/td&gt;
&lt;td&gt;−1.6&lt;/td&gt;
&lt;td&gt;+0.1&lt;/td&gt;
&lt;td&gt;−1.5 [−2.5, −0.4]&lt;/td&gt;
&lt;td&gt;9-28% cheaper&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chunks of ≤19, similar labels together&lt;/td&gt;
&lt;td&gt;−4.6 [−7.5, −1.8]&lt;/td&gt;
&lt;td&gt;−0.3&lt;/td&gt;
&lt;td&gt;−1.7&lt;/td&gt;
&lt;td&gt;+0.8&lt;/td&gt;
&lt;td&gt;−1.2 [−2.2, −0.3]&lt;/td&gt;
&lt;td&gt;same&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chunks of ≤19 with 3-digit codes&lt;/td&gt;
&lt;td&gt;0.0&lt;/td&gt;
&lt;td&gt;+0.1&lt;/td&gt;
&lt;td&gt;−0.3&lt;/td&gt;
&lt;td&gt;+0.6&lt;/td&gt;
&lt;td&gt;+0.1 [−0.6, +0.9]&lt;/td&gt;
&lt;td&gt;same&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The last row shows the 3-digit codes aren't what costs accuracy. On TREC fine, the baseline setup picked the wrong broad category (the part before the colon, like &lt;code&gt;NUM&lt;/code&gt; or &lt;code&gt;LOC&lt;/code&gt;) in 26 of 500 cases; the alternatives did so in 45 to 55. We don't know why spreading labels across chunks helps there; it's a measured pattern, not an explained one. In the single-call runs we also checked the candidate list: the probability left off the top 20 never exceeded 5% in 5,354 calls.&lt;/p&gt;



&lt;p&gt;&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;
  Why the category tree lost
  &lt;p&gt;We used each dataset's own hierarchy (for Banking77, groups we wrote by hand). On the same cases the tree trailed Jev by 4 points on TREC, 8 on Banking77, 26 on DBpedia and 29 on CLINC150. Almost all of the loss is in the first step: on CLINC the tree picked the right domain 66% of the time, while the flat version implicitly gets it right 96% of the time. Once in the right branch it rarely misses. Our guess is the names: CLINC's domains are called "meta", "utility" and "home", DBpedia's top level "Agent", "UnitOfWork" and "TopicalConcept". If you walk a hierarchy, make its group names describe what's inside them, and test it against flat groups.&lt;/p&gt;



&lt;p&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Confidence: making probabilities honest
&lt;/h2&gt;

&lt;p&gt;Accuracy was close. Confidence was not: out of the box, our DIY adapter sounded far more certain than it deserved to be on every single dataset.&lt;/p&gt;

&lt;p&gt;Think of confidence like a weather forecast: if a meteorologist says there is a 90% chance of rain, it should actually rain 9 out of 10 times. Raw language models are notoriously overconfident: when an LLM claims 99% certainty on an ambiguous edge case, it might only be right 75% of the time. &lt;/p&gt;

&lt;p&gt;Expected Calibration Error (ECE) measures this honesty gap: the average difference between what the model claims and how often it is actually right. Lower is better (0 means perfect honesty).&lt;/p&gt;

&lt;p&gt;Part of our initial overconfidence was a profile default: the default gateway profile used to leave out candidates outside the top 98% of probability (&lt;code&gt;top_p=0.98&lt;/code&gt;). On a confident answer, that zeros out every runner-up option. Setting &lt;code&gt;top_p=1.0&lt;/code&gt; explicitly brings back their real, tiny probabilities without changing which answer wins.&lt;/p&gt;

&lt;p&gt;The real fix is temperature scaling, which requires just one number (T). We raise each probability to the power 1/T and renormalize (&lt;code&gt;calibration_t&lt;/code&gt; in the code). When T &amp;gt; 1, it gently cools down the model's extreme certainty, spreading probability back to runners-up while keeping the winning label identical. At T=6, an extreme 99.99% / 0.01% split softens to an honest 82% / 18%.&lt;/p&gt;

&lt;p&gt;We fitted T on a few hundred labelled examples per dataset. Afterwards, our confidence was just as honest as Jev's out of the box, achieving a lower calibration error (ECE) on 9 of the 13 datasets.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbgildax6fn9t96t7pi5k.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbgildax6fn9t96t7pi5k.png" alt="expected calibration error per dataset for DIY raw, DIY with a fitted temperature, and Jev raw. Raw DIY is worst on every set, up to 0.53; with the fitted temperature it drops to 0.10 or below everywhere, below Jev on 9 of 13." width="800" height="524"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;One T doesn't suit every task, so temperature can be set per question, fitted on a few hundred labelled examples of your own.&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;
  The numbers behind this
  &lt;p&gt;We measure honesty with expected calibration error (ECE): the average gap between how sure the model says it is and how often it is right. Lower is better. Raw, ours ranged from 0.06 to 0.53 across the 13 sets, Jev's from 0.03 to 0.35.&lt;/p&gt;

&lt;p&gt;Before the &lt;code&gt;top_p&lt;/code&gt; fix, in an earlier run on five of these datasets, about 70% of answers came back with no runner-up at all, so no rescaling could help. After it, 0.4%.&lt;/p&gt;

&lt;p&gt;With T fitted on half of each dataset (138 to 385 cases) and measured on the other half, T ranged from 2 to 10 and ECE was 0.10 or lower on all 13 sets (up to 0.13 with a different random split). That is below Jev's out-of-the-box ECE on 9 of 13. Jev could be tuned the same way; with the same fit applied to both, ours is lower on 10 of 13. With T=6 everywhere, TREC question type, CLINC150 and DBpedia became underconfident (ECE about 0.2). For the four sets past 19 labels, the confidence comes from the final round only.&lt;/p&gt;

&lt;p&gt;We also tried a fix from Nokia's &lt;a href="https://github.com/nokia-applied-research/AnyJev" rel="noopener noreferrer"&gt;AnyJev&lt;/a&gt;: ask twice with the options reversed and average. The order does matter, changing our answer in 2.4% (AG News), 6.4% (TREC) and 13.7% (Emotion) of cases, but averaging both orders changed accuracy by +1.0, 0.0 and −1.3 points. Not worth doubling the calls. Reversed order alone cost up to 3 points (TREC), so the table's numbers hold for the label order we used.&lt;/p&gt;



&lt;p&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Speed and cost
&lt;/h2&gt;

&lt;p&gt;About 1.2 seconds per call at the median on short text, including an intermediary network hop; about 0.3 seconds for Jev. Past 19 labels, 3 to 4 seconds.&lt;/p&gt;

&lt;p&gt;Jev's list price is $0.042 per 1M input tokens with free output. &lt;code&gt;gpt-6-luna&lt;/code&gt; costs $0.10 per 1M input and $0.50 per 1M output. Measured per decision on live calls:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A single sentence of text, one question: ours $8, Jev $13 per 1M decisions.&lt;/li&gt;
&lt;li&gt;About 500 words of text: ours $87, Jev $48 per 1M decisions.&lt;/li&gt;
&lt;li&gt;Several questions about the same text favour Jev further: Jev sends the text once per call, ours sends it with every question.&lt;/li&gt;
&lt;li&gt;Past 19 labels every group is its own call and resends the text. On our short benchmark texts that made a 219-label decision about 2.2 times as expensive as one call holding all labels (1.5 times at 42 to 150 labels); the longer the text, the bigger the gap.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Filqtme2kiny9445hmokf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Filqtme2kiny9445hmokf.png" alt="Line chart of cost per million decisions versus text length in tokens. DIY is cheaper below about 90 tokens; above that Jev is cheaper, about $48 vs $87 at 500 words, and cheaper still when three questions share one call." width="800" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The break-even point is about 90 tokens of text.&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;
  More on speed, capacity and cost
  &lt;p&gt;Latency: 2 seconds at the 90th percentile on short text; we ran 10 calls at a time against Jev's 5. With very long text we saw 2-6 second medians.&lt;/p&gt;

&lt;p&gt;Capacity: shared public evaluation pools can face transient rate limits under batch load. In load testing, the adapter was verified under high synthetic concurrency, maintaining stable throughput and &amp;gt;99% completion across parallel requests.&lt;/p&gt;

&lt;p&gt;Cost: ours pays for 5 output tokens per question, about 30% of the cost on short text. The Jev side uses TypeSafe's list price; what Cloudflare charges for it was not visible to us. TypeSafe's own benchmark (&lt;a href="https://evals.typesafe.ai" rel="noopener noreferrer"&gt;evals.typesafe.ai&lt;/a&gt;) reports $0.0004 per case for Jev and $0.0033 for Luna, but there Luna generates text answers and the page doesn't say how many decisions a case holds, so those figures don't map onto ours.&lt;/p&gt;



&lt;p&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  If you build your own
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Check that logprobs actually come back.&lt;/strong&gt; Support depends on the provider, the model and the API you call. When they're missing you may still get an answer, just no probabilities, and nothing errors.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Set &lt;code&gt;top_p=1.0&lt;/code&gt;&lt;/strong&gt;, or the runner-up options read as zero.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Switch from digits to letters past 10 options.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Enterprise safety filters can block legitimate domain queries.&lt;/strong&gt; Provider content filters occasionally refuse domain texts (even benign content). In enterprise setups, handling this often requires dedicated deployments with modified filtering, separate quota, or explicit fallback paths for refused groups.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Option order matters a little&lt;/strong&gt;, up to 3 points in our tests. Keep it fixed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fit the confidence before you rely on it.&lt;/strong&gt; Run a few hundred labelled examples and pick the T where "80% sure" is right about 80% of the time.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What we got out of it
&lt;/h2&gt;

&lt;p&gt;On picking a category, the DIY version is within about a point of Jev on six datasets, trailing by 2.4 points on DBpedia (219 labels). On yes/no questions and rating scales Jev is a few points more accurate, and it is faster, cheaper on anything longer than a sentence, and honest about its confidence out of the box. What this prototype gave us is time: it was running within days on existing LLM infrastructure, takes finer scales, and reads far more text. Now we can observe real demand before deciding whether a dedicated model is justified. If it is, a purpose-built model goes in behind the exact same request shape.&lt;/p&gt;

&lt;p&gt;If you try this yourself, benchmark on data the model is unlikely to have seen, then on your own, and check the confidence against labels before you set a threshold on it.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Disclaimer: This article presents an independent technical evaluation on public datasets. The views expressed are my own. The author is not affiliated with or endorsed by TypeSafe AI.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>jev</category>
      <category>llm</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
