DEV Community

Cover image for A DIY Jev: typed decisions from gpt-6-luna logprobs, benchmarked head-to-head
Krzysztof Głuszczyk
Krzysztof Głuszczyk

Posted on AI-assisted

A DIY Jev: typed decisions from gpt-6-luna logprobs, benchmarked head-to-head

What if an ordinary LLM could return this:

{"refund": 0.72, "fraud": 0.21, "other": 0.07}
Enter fullscreen mode Exit fullscreen mode

...not by asking the model to write JSON numbers, but by turning its token logprobs into a typed, calibrated probability distribution?

That is roughly the promise of TypeSafe's Jev: purpose-built models returning probabilistic decisions instead of free-form text. We wanted to see how much of that experience we could reproduce using gpt-6-luna, already available in our stack, before evaluating dedicated commercial endpoints.

So we built the smallest prototype that could work: restrict the output to single-token keys, read their logprobs, normalize them, and fit a single temperature parameter on labelled examples.

Then we tested it head-to-head against Jev across 6,762 public examples (13 datasets).

The result was more nuanced than expected. Our DIY classifier nearly matched Jev on most category tasks (even with 219 classes), but lost decisively on yes/no questions and rating scales. It was cheaper for short inputs, roughly 4x slower, and heavily overconfident until calibrated.

Why not just ask the model for JSON?

The obvious alternative is prompting the model to output a JSON dictionary with confidence scores. We benchmarked that as a third approach on gpt-6-luna across the exact same 6,762 cases:

Approach Accuracy Latency Tokens / call ECE Practical role
Prompted JSON 70.1% 10.06s 961 0.175 Slow, expensive prototype
DIY logprob adapter 73.8% 2.11s 1 0.211 Low-cost prototype
Dedicated model (Jev) 76.5% 0.38s 0 0.110 Dedicated specialized model

Across all 6,762 cases, prompted JSON was 4.8 times slower than the logprob adapter and 26.5 times slower than Jev. It generated ~6.5 million output tokens across the benchmark. The DIY adapter generated 6,762.

JSON was not consistently inaccurate. It reached the highest accuracy on five datasets and tied Jev on Banking77 (80.6%). The failure was variance. It collapsed on two large label tasks:

Dataset Classes Jev DIY JSON JSON latency JSON tokens JSON ECE
TREC-50 50 83.2% 82.4% 38.4% 11.4s 986 0.574
DBpedia219 219 93.2% 90.7% 55.3% 52.6s 5,751 0.427

Two practical reasons make the JSON approach fail at scale:

  1. Token explosion and latency: Writing JSON keys and float numbers across chunked tournament rounds burns up to 5,751 output tokens per triage call. At 52.6 seconds per item, asking for JSON turns a routing classifier into an unusable bottleneck. TypeSafe's own benchmark (evals.typesafe.ai) evaluated Luna generating text answers, where it ended up roughly 8x more expensive ($0.0033 vs $0.0004 per decision) and significantly slower. With logprobs, the model generates one token, reading all candidates from that single decoding step.
  2. Confidence distortion under tournament chunking: When a label set exceeds 19 options and must be evaluated in tournament chunks, verbalized text confidences break down. In each local chunk, the model assigns high confidences (0.85+) to mediocre local candidates. In subsequent rounds, these inflated local scores eliminate the true global winner, driving ECE to 0.574 on TREC-50 and 0.427 on DBpedia. Reading token logprobs gives direct access to the model's next-token probabilities, avoiding local verbalized score inflation.

Prompted JSON is fine for a quick one-off prototype. For a reliable decision endpoint, the 1-token logprob adapter is faster, cheaper, and far more stable across label sizes.

Here is how the trick works, where it breaks, and what to watch out for if you build your own.

How it works

A language model writes one token (a word or a piece of one) at a time. At each step it gives every token in its vocabulary a probability, then picks one. Many APIs can also return the top candidates it was choosing between, each with its probability, as a logarithm ("logprob").

That is all we need. List the options with a short key each (0, 1, 2...), ask the model to answer with just the key, and read the probabilities the API returns for those keys. Keep the candidates that are keys, turn the logprobs back into probabilities, rescale them to add up to 1, and you have a probability for every option.

The API returns at most 20 candidates per token, and many of them are junk: "A", "**", an end-of-text marker, or the same key twice (" B" and "b"). So the least likely options often don't make the list and read as zero. That costs almost nothing: anything left off ranks below the last candidate shown, typically with negligible individual probability. We stop at 19 options per call (details below).

Three steps, in Python. The full example runs against any OpenAI-compatible endpoint that serves gpt-6-luna on the Responses API, and the outputs shown are from a live call.

1. Build the prompt. Each option gets a one-token key: digits up to 10 options, letters past that. Why the answer looks like this comes right after the code.

import json, math
from openai import OpenAI

client = OpenAI()  # any OpenAI-compatible endpoint; pass base_url=... for a gateway

def build_prompt(text, question, options):
    """options: {label: description}. Returns the answer keys and the prompt."""
    n = len(options)
    # Keys are 0-9 up to 10 options. Past that, "1" would also be the start of
    # "10".."18", so use letters, skipping "a" and "i" (they are also English words).
    keys = [str(i) for i in range(n)] if n <= 10 else list("bcdefghjklmnopqrstu")[:n]
    listing = "\n".join(f"{k.upper()}: {label} - {desc}"
                        for k, (label, desc) in zip(keys, options.items()))
    word = "number" if n <= 10 else "letter"
    prompt = (f"State: {json.dumps({'text': text}, ensure_ascii=False)}\n\n"
              f"Question: {question}\n\nOptions:\n{listing}\n\n"
              f"Answer with only the {word} of the best option. No other text.")
    return keys, prompt
Enter fullscreen mode Exit fullscreen mode

2. Ask for one short answer, with its candidates. top_p=1.0 matters: with the default, the API leaves out the options the model considered unlikely, and they come back as exactly zero.

def ask(prompt):
    resp = client.responses.create(
        model="gpt-6-luna",
        input=prompt,
        reasoning={"effort": "none"},
        max_output_tokens=64,   # reasoning and output share this budget
        top_logprobs=20,        # the most candidates the API returns per token
        top_p=1.0,              # the default 0.98 hides the runner-up options
        include=["message.output_text.logprobs"],
        store=False,
    )
    msg = next(o for o in resp.output if o.type == "message")
    return msg.content[0].logprobs  # one entry per output token, with its top candidates
Enter fullscreen mode Exit fullscreen mode

3. Turn the candidates into probabilities.

def to_probs(tokens, keys, labels, calibration_t=1.0):
    norm = lambda t: t.strip().strip(".,;:!?'\"`*").lower()
    for pos in tokens:  # the first output token that is an answer key
        if norm(pos.token) in keys:
            mass = {}
            for alt in pos.top_logprobs:
                k = norm(alt.token)
                if k in keys:
                    mass[k] = mass.get(k, 0.0) + math.exp(alt.logprob)  # logprob -> probability
            mass = {k: v ** (1 / calibration_t) for k, v in mass.items()}  # see "Confidence"
            total = sum(mass.values())
            return {l: mass.get(k, 0.0) / total for k, l in zip(keys, labels)}
    return None  # no answer key anywhere: the model replied off-format

def classify(text, question, options, calibration_t=1.0):
    keys, prompt = build_prompt(text, question, options)
    return to_probs(ask(prompt), keys, list(options), calibration_t)

opts = {"billing": "Charges, payment, refunds", "support": "General help", "sales": "New bookings"}
print(classify("I was charged twice for the same booking.", "Which team should handle this?", opts))
# {'billing': 0.9999999999862318, 'support': 6.755284692114235e-12, 'sales': 7.0129827926438374e-12}
print(classify("I was charged twice for the same booking.", "Which team should handle this?", opts,
               calibration_t=6))
# {'billing': 0.9731562277998806, 'support': 0.013380012188266373, 'sales': 0.013463760011852986}
Enter fullscreen mode Exit fullscreen mode

Why the answer is a single key

The API gives probabilities for single tokens, and it only shows candidates along the path the model actually took. So each option has to be exactly one token, and no option can look like the start of another. Label names often fail that test: "customer support" and "customer sales" both start with "customer", so the first step can't tell them apart. A one-token key per option puts every option side by side in a single step.

More on the answer format, and how often it failed

Digits up to 10, letters past that. "0" to "9" are single tokens. Two-digit numbers usually are too: in a live check with 12 options the model answered "11" as one token, with "1" far behind. But a "1" can also be the start of a longer answer, and in 1 of 54 early test cases the probability did land on option 1 when the right answer was 11 or above. Letters never overlap like that.

Why 19? We first reserved one of the 20 slots for junk. Measured later on 100 Banking77 questions with 19 options, that reasoning doesn't hold: junk took a median of 9 slots and duplicate keys 2 more, and some lists came back shorter than 20, so a median of 10 options were missing from the list; 20 options looked the same. No number of options fits on the list every time; what matters is that only negligible options fall off. 19 is simply where our clean one-letter keys run out. Larger single calls are covered under "More than 19 labels".

Letters start at B and skip I, because "A" and "I" also open ordinary sentences. If the model ignores the format and replies "I think...", we want no valid key in the reply, so the question comes back as "no answer" instead of a confident vote for option I. B to U without I gives exactly 19.

Why not force a one-token reply? The API rejects max_output_tokens below 16, reasoning shares that budget, and the model often adds a full stop anyway. So we take the first output token that is a key, and return "no answer" if there is none.

Why not structured outputs or logit_bias? We tried both, hoping to limit the candidates to our keys. Neither did. With a JSON schema and an enum, the candidate list was still the model's raw top 20, junk included, and through our setup the schema wasn't even enforced when the prompt asked for a bare letter (3 of 3 calls came back as plain "D"). logit_bias was accepted but had no visible effect, even at the maximum.

How often it went wrong. In the 13-set benchmark, 3 of 6,762 questions (all on Yelp) came back with no valid key and counted as wrong. What we can't count is a prose reply that happens to start with a valid key, which is why A and I are left out. They are real risks: "A" showed up among the candidates in 97 of 100 test questions and "I" in 42, as junk.

A classifier you don't have to train

Think of it as a text classifier you don't have to train. The usual way to sort text into fixed labels is to train a small model such as BERT on thousands of labelled examples, and it learns to output a score per label. Here the language model's own word probabilities do that job, so there is nothing to train, and the labels can change on every request. What you give up is confidence you can take at face value. A trained classifier learns from labelled examples how sure it should be; the raw numbers here run too high, so "99.99% sure" can be right much less often than that. The "Confidence" section shows how to fix it.

The prototype wraps this in one endpoint that takes several questions about the same text, following TypeSafe's Primitives spec: choice (pick a label), score (pick a level on an ordered scale) and noul (yes/no, one probability). Each answer also reports coverage, the share of probability that landed on your options at all. Low coverage means the model wanted to say something else, so don't trust that answer.

How it compares with Jev

We ran both on 13 public datasets, 6,762 cases: seven classification sets with 4 to 219 labels, two yes/no sets and four rating scales. Every case went to both with the same question and label descriptions. We call a gap real only when a statistical test says it is unlikely to be luck. Jev was called through Cloudflare Workers AI on 2026-09-26.

accuracy of DIY and Jev on 13 datasets grouped as classify, yes/no and score, with the difference per row. Classification rows are within about a point except DBpedia at -2.4; yes/no and score rows are 3.6 to 6.1 points behind Jev.

Picking a category: about even

Within about a point on all seven sets, with no real difference on six, including CLINC150 with 150 intents. The exception is DBpedia with 219 categories, where Jev leads by 2.4 points. Jev is slightly ahead on most sets, just not by enough to tell apart from noise.

Yes/no: Jev ahead

Jev leads by 3.6 points on BoolQ and 4.0 on RTE. Both gaps are real.

Rating scales: Jev ahead, mostly by one step

Jev leads by 4.6 to 6.1 points on exact hits. Most of our misses are one level off: counting "within one level" as right, the two are close (97.8% for both on SST-5, 96.2% vs 98.2% on Yelp, among the cases both answered).

Scales are also where ours goes further. Jev takes up to 10 levels, ours up to 19. On 342 wine reviews split into 2 to 19 score bands, ours sits a few points under Jev up to 10 levels. At 19 levels it is off by 1.8 rating points on average, a little better than at 10 levels (2.1; Jev at 10: 2.0).

line chart of exact and within-one-level accuracy against 2, 4, 8, 10 and 19 score levels on 342 wine reviews. Both engines fall from 90% exact at 2 levels to about 30% at 10, with Jev a few points ahead; only DIY continues to 19 levels, at 19% exact and 49% within one level.

Across all 13 sets ours trails by 2.7 points on average. It also naturally inherits the large context window of the underlying general-purpose model, accepting inputs well beyond short classifications. We did not compare long inputs against Jev.

Full results, statistics and caveats
Test (labels, cases) DIY Jev Difference, 95% CI
AG News (4, 500) 83.8% 85.4% −1.6 [−3.4, 0.0]
Emotion (6, 566) 51.8% 50.2% +1.6 [−1.1, +4.2]
TREC question type (6, 500) 90.6% 91.6% −1.0 [−2.8, +0.6]
TREC fine (42, 500) 82.4% 83.2% −0.8 [−3.6, +2.0]
Banking77 (77, 770) 79.2% 80.6% −1.4 [−3.4, +0.5]
CLINC150 (150, 750) 91.7% 91.5% +0.3 [−1.6, +2.1]
DBpedia (219, 657) 90.7% 93.2% −2.4 [−4.0, −0.9]
BoolQ (yes/no, 500) 89.2% 92.8% −3.6 [−5.8, −1.6]
RTE (yes/no, 277) 87.7% 91.7% −4.0 [−7.6, −0.4]
SST-5 sentiment (5, 500) 53.4% 58.0% −4.6 [−8.2, −1.0]
Yelp stars (5, 500) 62.8% 68.8% −6.0 [−9.8, −2.2]
IMDB rating (4, 400) 67.8% 72.8% −5.0 [−8.8, −1.0]
Wine score (10, 342) 28.1% 34.2% −6.1 [−12.9, +0.9]

How we tested. Equal numbers of cases per label, so these are benchmark accuracies, not what you'd see on real traffic. The interval next to each difference is a 95% paired bootstrap: we resample the same cases many times, and a gap counts as real when the interval excludes zero. By that test, six of the 13 gaps are real, all in Jev's favour. Testing 13 datasets at once raises the odds that one looks real by luck, so we also applied a stricter bar (McNemar's test with a Holm correction); three gaps pass it: BoolQ, Yelp and DBpedia.

The 2.7-point average. 95% CI 1.7 to 3.6, bootstrapping cases within each dataset as planned; 1.2 to 4.1 if you treat each dataset as a sample from a wider pool of tasks.

Failed answers count as wrong. Refused or invalid answers (less than 1% across the entire benchmark, mostly safety filter triggers on film reviews) were counted as misses. Counting those as misses, Jev keeps a 2-3 point lead on Yelp and IMDB even on "within one level". Jev answered everything.

What these numbers can't tell you:

  • All 13 are well-known public benchmarks, so both models have probably seen them in training; Banking77 and CLINC150 are standard intent-training sets. Your own data is the real test.
  • Six of the sets (AG News, BoolQ, SST-5, IMDB, Banking77, wine) were also used, on other samples, while we tuned the letters, the chunking, top_p and calibration. The seven sets we hadn't touched show the same pattern.
  • One run per engine per set. Five sets got a second DIY run by other means; none moved by more than 0.6 points. The wine re-run gave 28.7% vs 33.9% at 10 levels, against 28.1% vs 34.2% in the table.
  • Jev ran through Cloudflare, not TypeSafe directly, and we can't verify it's the same build.
  • The long-text test hid a code in filler text, 6 trials per size.

More than 19 labels

One call holds 19 options. Past that, what worked was splitting the labels into groups of up to 19, asking about every group in parallel, then asking once more among the group winners:

from concurrent.futures import ThreadPoolExecutor

def classify_many(text, question, options):
    """Past 19 labels: chunks of up to 19 in parallel, then a final round among the winners."""
    labels = list(options)
    if len(labels) <= 19:
        return classify(text, question, options)
    n = -(-len(labels) // 19)
    chunks = [{l: options[l] for l in labels[i::n]} for i in range(n)]
    with ThreadPoolExecutor(n) as pool:
        results = list(pool.map(lambda c: classify(text, question, c), chunks))
    winners = [max(d, key=d.get) for d in results if d]  # a refused chunk drops out here
    if not winners:
        return None
    return classify_many(text, question, {w: options[w] for w in winners})
Enter fullscreen mode Exit fullscreen mode

That is how the 42- to 219-label rows were produced: 4 calls for 42 labels, 9 for 150, 13 for 219, in 3 to 4 seconds.

Since options that don't make the candidate list barely matter, we also tried alternatives: all labels in one call (keyed by 3-digit codes, each a single token), bigger chunks of up to about 40, and chunks of 19 that keep similar labels together. Each ran twice on the same four datasets, against a same-day rerun of the baseline setup. None matched it. The single call came closest: 3 to 4 times faster, 30 to 55% cheaper, and even on Banking77, CLINC150 and DBpedia, but 4 points lower on TREC fine, where it confused the broad category (a place for a number, say) twice as often. Bigger chunks and grouped chunks lost on TREC fine too, and a little overall. So chunks of 19, with labels spread across them, remain the baseline.

To see what the number of labels does on its own, we took random subsets of CLINC150's intents, from 5 to 150, with the same cases for both. Both stay near 100% up to 20 labels. At 40 and 80 ours trails by 3.4 and 2.5 points, inside the noise, and at 150 they tie.

accuracy of DIY and Jev on random subsets of 5, 10, 20, 40, 80 and 150 CLINC150 intents. Both are near 100% up to 20 classes; at 40 and 80 DIY is 3.4 and 2.5 points behind, within the 95% bands; at 150 both are 91.7%.

The obvious alternative, walking a category tree (pick the group first, then the label inside it), was clearly worse: 4 to 29 points behind Jev.

The alternatives in numbers

Same pre-registered cases as the table above, prediction = most probable option in every arm, each alternative run twice, compared with a same-day rerun of the baseline setup (which matched the earlier run within 0.6 points). Differences are in points with 95% intervals; both runs are averaged per case.

Setup TREC fine (42) Banking77 (77) CLINC150 (150) DBpedia (219) All four Cost vs baseline
Baseline: chunks of ≤19, labels spread 83.0% 79.0% 91.7% 90.4%
One call, all labels, sorted −4.3 [−7.1, −1.5] −0.1 +0.9 +1.1 −0.3 [−1.3, +0.7] 31-56% cheaper, ~1.1 s
Chunks of 21-39 −5.7 [−8.9, −2.5] +0.1 −1.6 +0.1 −1.5 [−2.5, −0.4] 9-28% cheaper
Chunks of ≤19, similar labels together −4.6 [−7.5, −1.8] −0.3 −1.7 +0.8 −1.2 [−2.2, −0.3] same
Chunks of ≤19 with 3-digit codes 0.0 +0.1 −0.3 +0.6 +0.1 [−0.6, +0.9] same

The last row shows the 3-digit codes aren't what costs accuracy. On TREC fine, the baseline setup picked the wrong broad category (the part before the colon, like NUM or LOC) in 26 of 500 cases; the alternatives did so in 45 to 55. We don't know why spreading labels across chunks helps there; it's a measured pattern, not an explained one. In the single-call runs we also checked the candidate list: the probability left off the top 20 never exceeded 5% in 5,354 calls.

Why the category tree lost

We used each dataset's own hierarchy (for Banking77, groups we wrote by hand). On the same cases the tree trailed Jev by 4 points on TREC, 8 on Banking77, 26 on DBpedia and 29 on CLINC150. Almost all of the loss is in the first step: on CLINC the tree picked the right domain 66% of the time, while the flat version implicitly gets it right 96% of the time. Once in the right branch it rarely misses. Our guess is the names: CLINC's domains are called "meta", "utility" and "home", DBpedia's top level "Agent", "UnitOfWork" and "TopicalConcept". If you walk a hierarchy, make its group names describe what's inside them, and test it against flat groups.

Confidence

This was the surprise. Accuracy was close. Confidence was not: out of the box, ours sounded much more sure than it deserved to on every dataset, and clearly more than Jev.

Part of it was a profile default: the default profile used to leaving out candidates outside the top 98% of probability (top_p 0.98). On a confident answer that is every option but the winner, so they all read as exactly zero. Setting top_p=1.0 explicitly brings back their real, tiny probabilities, without changing which answer wins.

The rest takes one number. Raise each probability to the power 1/T and rescale (calibration_t in the code). With T above 1 the model sounds less sure, and the winning answer stays the same. We fitted T on half of each dataset and measured on the other half. Afterwards our confidence was about as honest as Jev's out of the box, better on 9 of 13 sets.

expected calibration error per dataset for DIY raw, DIY with a fitted temperature, and Jev raw. Raw DIY is worst on every set, up to 0.53; with the fitted temperature it drops to 0.10 or below everywhere, below Jev on 9 of 13.

One T doesn't suit every task, so temperature can be set per question, fitted on a few hundred labelled examples of your own.

The numbers behind this

We measure honesty with expected calibration error (ECE): the average gap between how sure the model says it is and how often it is right. Lower is better. Raw, ours ranged from 0.06 to 0.53 across the 13 sets, Jev's from 0.03 to 0.35.

Before the top_p fix, in an earlier run on five of these datasets, about 70% of answers came back with no runner-up at all, so no rescaling could help. After it, 0.4%.

With T fitted on half of each dataset (138 to 385 cases) and measured on the other half, T ranged from 2 to 10 and ECE was 0.10 or lower on all 13 sets (up to 0.13 with a different random split). That is below Jev's out-of-the-box ECE on 9 of 13. Jev could be tuned the same way; with the same fit applied to both, ours is lower on 10 of 13. With T=6 everywhere, TREC question type, CLINC150 and DBpedia became underconfident (ECE about 0.2). For the four sets past 19 labels, the confidence comes from the final round only.

We also tried a fix from Nokia's AnyJev: ask twice with the options reversed and average. The order does matter, changing our answer in 2.4% (AG News), 6.4% (TREC) and 13.7% (Emotion) of cases, but averaging both orders changed accuracy by +1.0, 0.0 and −1.3 points. Not worth doubling the calls. Reversed order alone cost up to 3 points (TREC), so the table's numbers hold for the label order we used.

Speed and cost

About 1.2 seconds per call at the median on short text, including an intermediary network hop; about 0.3 seconds for Jev. Past 19 labels, 3 to 4 seconds.

Jev's list price is $0.042 per 1M input tokens with free output. gpt-6-luna costs $0.10 per 1M input and $0.50 per 1M output. Measured per decision on live calls:

  • A single sentence of text, one question: ours $8, Jev $13 per 1M decisions.
  • About 500 words of text: ours $87, Jev $48 per 1M decisions.
  • Several questions about the same text favour Jev further: Jev sends the text once per call, ours sends it with every question.
  • Past 19 labels every group is its own call and resends the text. On our short benchmark texts that made a 219-label decision about 2.2 times as expensive as one call holding all labels (1.5 times at 42 to 150 labels); the longer the text, the bigger the gap.

Line chart of cost per million decisions versus text length in tokens. DIY is cheaper below about 90 tokens; above that Jev is cheaper, about $48 vs $87 at 500 words, and cheaper still when three questions share one call.

The break-even point is about 90 tokens of text.

More on speed, capacity and cost

Latency: 2 seconds at the 90th percentile on short text; we ran 10 calls at a time against Jev's 5. With very long text we saw 2-6 second medians.

Capacity: shared public evaluation pools can face transient rate limits under batch load. In load testing, the adapter was verified under high synthetic concurrency, maintaining stable throughput and >99% completion across parallel requests.

Cost: ours pays for 5 output tokens per question, about 30% of the cost on short text. The Jev side uses TypeSafe's list price; what Cloudflare charges for it was not visible to us. TypeSafe's own benchmark (evals.typesafe.ai) reports $0.0004 per case for Jev and $0.0033 for Luna, but there Luna generates text answers and the page doesn't say how many decisions a case holds, so those figures don't map onto ours.

If you build your own

  • Check that logprobs actually come back. Support depends on the provider, the model and the API you call. When they're missing you may still get an answer, just no probabilities, and nothing errors.
  • Set top_p=1.0, or the runner-up options read as zero.
  • Switch from digits to letters past 10 options.
  • Enterprise safety filters can block legitimate domain queries. Provider content filters occasionally refuse domain texts (even benign content). In enterprise setups, handling this often requires dedicated deployments with modified filtering, separate quota, or explicit fallback paths for refused groups.
  • Option order matters a little, up to 3 points in our tests. Keep it fixed.
  • Fit the confidence before you rely on it. Run a few hundred labelled examples and pick the T where "80% sure" is right about 80% of the time.

What we got out of it

On picking a category, the DIY version is within about a point of Jev on six datasets, trailing by 2.4 points on DBpedia (219 labels). On yes/no questions and rating scales Jev is a few points more accurate, and it is faster, cheaper on anything longer than a sentence, and honest about its confidence out of the box. What this prototype gave us is time: it was running within days on existing LLM infrastructure, takes finer scales, and reads far more text. Now we can observe real demand before deciding whether a dedicated model is justified. If it is, a purpose-built model goes in behind the exact same request shape.

If you try this yourself, benchmark on data the model is unlikely to have seen, then on your own, and check the confidence against labels before you set a threshold on it.


Disclaimer: This article presents an independent technical evaluation on public datasets. The views expressed are my own. The author is not affiliated with or endorsed by TypeSafe AI.

Top comments (0)