<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Stephen Ohakanu</title>
    <description>The latest articles on DEV Community by Stephen Ohakanu (@sohakanu).</description>
    <link>https://dev.to/sohakanu</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4126601%2F9b5e9289-d903-457f-ba1c-8b1fbc54236f.png</url>
      <title>DEV Community: Stephen Ohakanu</title>
      <link>https://dev.to/sohakanu</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sohakanu"/>
    <language>en</language>
    <item>
      <title>One Imperfect Page per Hundred: Testing Scribe on Data I Didn't Curate</title>
      <dc:creator>Stephen Ohakanu</dc:creator>
      <pubDate>Tue, 15 Sep 2026 16:30:25 +0000</pubDate>
      <link>https://dev.to/sohakanu/one-imperfect-page-per-hundred-ambushing-our-pipeline-with-data-nobody-tuned-fh8</link>
      <guid>https://dev.to/sohakanu/one-imperfect-page-per-hundred-ambushing-our-pipeline-with-data-nobody-tuned-fh8</guid>
      <description>&lt;p&gt;I tested Scribe on 100 handwritten prescriptions with annotations supplied by someone else. I had not used these pages to tune the pipeline.&lt;/p&gt;

&lt;p&gt;With the Qwen verifier, &lt;strong&gt;97 pages were sent for human review&lt;/strong&gt;. Of the three accepted automatically, two were correct and one missed a medicine. With the Gemma verifier, 99 pages went to review and the single page accepted automatically was correct.&lt;/p&gt;

&lt;p&gt;The review checks limited the number of unreviewed errors, but the pipeline still needed human help on almost every page. This post walks through what the run showed and what it did not establish.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this started
&lt;/h2&gt;

&lt;p&gt;In &lt;a href="https://dev.to/sohakanu/confidence-is-theater-benchmarking-nine-local-vlm-pipelines-on-handwritten-clinical-forms-38h6"&gt;part 1&lt;/a&gt;, I built Scribe — a local pipeline that turns handwritten clinical forms into structured data and, more importantly, learned when &lt;em&gt;not&lt;/em&gt; to trust a model's answer. The headline was that model confidence carried no signal, and disagreement between two independent reads did.&lt;/p&gt;

&lt;p&gt;But every number in part 1 had a limitation.&lt;/p&gt;

&lt;p&gt;I curated the gold answers. I wrote the extraction schema. I tuned the prompts against the same 23 pages I was scoring on. Even done carefully, that means the system and the test were developed together.&lt;/p&gt;

&lt;p&gt;So I wanted a test set I had no hand in.&lt;/p&gt;

&lt;h2&gt;
  
  
  The external test
&lt;/h2&gt;

&lt;p&gt;I found a &lt;a href="https://huggingface.co/datasets/chaithanyakota/100-handwritten-medical-records" rel="noopener noreferrer"&gt;public dataset of 100 handwritten Indian prescriptions&lt;/a&gt;, where the ground truth — the list of medicines on each page — was written by the dataset's author, not by me.&lt;/p&gt;

&lt;p&gt;The rules I set for myself:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the extraction schema gets written once, from the task description, before seeing results&lt;/li&gt;
&lt;li&gt;the prompts stay frozen exactly as they were in part 1&lt;/li&gt;
&lt;li&gt;the author's answer conventions are the answer key, not mine&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Two things make this set difficult. The handwriting is real doctors' cursive — not the careful volunteer writing my benchmark was built from. And the drug names are Indian brand names the models have seen little of.&lt;/p&gt;

&lt;p&gt;That difficulty is useful: a system built for clinics in one country will eventually meet forms from another.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reading accuracy dropped sharply
&lt;/h2&gt;

&lt;p&gt;Brand names came back right 38% of the time.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;CEPODEM&lt;/em&gt; became &lt;strong&gt;Capelin&lt;/strong&gt;. &lt;em&gt;ESOTAB&lt;/em&gt; became &lt;strong&gt;Esotrib&lt;/strong&gt;. &lt;em&gt;NIVEOLI&lt;/em&gt; became &lt;strong&gt;Ulfah Niveoli&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;These are the same kind of confident misreads that affected patient names in part 1, and they happened more often here: the handwriting is harder and the vocabulary is unfamiliar. (The two evaluations use different measures — field-level accuracy on the development set, brand-name recall here — so I'm not putting a before-and-after number on it.)&lt;/p&gt;

&lt;p&gt;I ran both of my best configurations, which differ only in which model does the second read. They produced &lt;strong&gt;identical reading results&lt;/strong&gt; — same primary model, so the underlying reading was unchanged; the verifier only affected how much got flagged.&lt;/p&gt;

&lt;p&gt;That repeats part 1's central lesson: &lt;strong&gt;verification cannot fix reading. It can only catch it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;No prompt I own fixes the reading either. For the local vision models I tested, hard cursive remains a major limitation — and a &lt;a href="https://arxiv.org/abs/2604.16504" rel="noopener noreferrer"&gt;recent benchmark on South African maternity records&lt;/a&gt; reported similar difficulty even with frontier cloud models.&lt;/p&gt;

&lt;h2&gt;
  
  
  The review gate held
&lt;/h2&gt;

&lt;p&gt;This is the part that mattered to me.&lt;/p&gt;

&lt;p&gt;On these harder pages, the models were much less willing to claim certainty: most confidence scores fell to around 0.60–0.70, and &lt;strong&gt;97–99% of pages were sent for human review.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For this input, "a person must look at this" &lt;em&gt;is&lt;/em&gt; the correct output.&lt;/p&gt;

&lt;h2&gt;
  
  
  Auditing the escapes
&lt;/h2&gt;

&lt;p&gt;With the Qwen verifier, three pages out of a hundred skipped human review. I checked each one by hand.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Two were correct — single-medicine prescriptions where the extracted medicine matched the annotation.&lt;/li&gt;
&lt;li&gt;One was an eleven-medicine page that got ten of them right and &lt;strong&gt;dropped one&lt;/strong&gt; (&lt;code&gt;BETNESOL INJ&lt;/code&gt;).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the imperfect page is one out of 100 pages tested — but also one out of only three pages accepted without review. Those two denominators tell different stories. The result shows the pipeline routed almost everything to a person; it does not demonstrate high accuracy among automatically accepted pages, because there were too few of them to say.&lt;/p&gt;

&lt;p&gt;With the Gemma verifier, one page was accepted automatically, and it was correct.&lt;/p&gt;

&lt;p&gt;Part 1 ended with a rule: &lt;em&gt;count what escapes, and you're grading the system.&lt;/em&gt; This is what that looks like in practice — a specific, audited count of unreviewed errors rather than a single accuracy number.&lt;/p&gt;

&lt;h2&gt;
  
  
  The test also caught a bug in my scorer
&lt;/h2&gt;

&lt;p&gt;One more thing this exercise caught — in my own tooling.&lt;/p&gt;

&lt;p&gt;My first scoring pass penalized extracted rows for containing dose schedules like &lt;code&gt;(1-0-1)&lt;/code&gt;, even when they named the right drug. The scorer's precision check compared token sets in the wrong direction. External data debugs your evaluator, not just your model.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'm taking away
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Reading accuracy didn't transfer to new handwriting; the review-gate behavior did.&lt;/strong&gt; The parts that held up were the system parts — confidence behavior, dual reads, the review queue. The reading itself did not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Neither evaluation establishes performance on the intended target forms.&lt;/strong&gt; Part 1's benchmark is public data; this set is a different form type from a different country. The private held-out evaluation is what will measure the 98% target.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A high flag rate is not failure here.&lt;/strong&gt; Flagging 97% of pages the models genuinely couldn't read is correct behavior. The bad outcome would have been a normal flag rate with wrong drug names passing through unreviewed.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;The held-out evaluation: real target-form registers, hand-filled, photographed, never published, created after every model's training cutoff. Those are the numbers that will actually count, and they'll be part 3 of this series.&lt;/p&gt;

&lt;p&gt;And the checkbox problem from part 1 still needs its own detection stage, likely using the one model in my tests that produced no hallucinated values.&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;
  The numbers, for engineers
  &lt;br&gt;
Both finalist configurations (same Qwen3.8-27B primary; Qwen3-VL-8B vs Gemma-3-12B verifier) against all 100 pages, scored on the medicines list:

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Qwen-8B verifier&lt;/th&gt;
&lt;th&gt;Gemma-3 verifier&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Strict page-level exact&lt;/td&gt;
&lt;td&gt;1%&lt;/td&gt;
&lt;td&gt;1%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Medicine token-match F1&lt;/td&gt;
&lt;td&gt;0.27&lt;/td&gt;
&lt;td&gt;0.27&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Brand-name recall&lt;/td&gt;
&lt;td&gt;38% (136/359)&lt;/td&gt;
&lt;td&gt;38%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Page flag rate&lt;/td&gt;
&lt;td&gt;97%&lt;/td&gt;
&lt;td&gt;99%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two definitions of "correct" appear in this post, and they differ on purpose. &lt;em&gt;Strict page-level exact&lt;/em&gt; requires every extracted medicine to match the annotation's wording and order — so an extraction of "Drop VITAMIN D3 800 IU/ML (Depura)" fails against the annotation "DEPURA" even though it's the same medicine. The hand audit of auto-accepted pages, and the token-match metric, count a medicine as found when the brand name matches (60% gold-token coverage), which is why the audit reports pages as correct that the strict metric does not. Full scoring code: &lt;a href="https://github.com/kod201/scribe/blob/main/scripts/eval_rx100.py" rel="noopener noreferrer"&gt;&lt;code&gt;scripts/eval_rx100.py&lt;/code&gt;&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Dataset: &lt;a href="https://huggingface.co/datasets/chaithanyakota/100-handwritten-medical-records" rel="noopener noreferrer"&gt;100-handwritten-medical-records&lt;/a&gt; (CC-BY-ND-4.0), evaluated as-is and not redistributed.&lt;br&gt;
&lt;/p&gt;

&lt;br&gt;
&lt;p&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;This is part 2 of a series.&lt;/strong&gt; Part 1: &lt;a href="https://dev.to/sohakanu/confidence-is-theater-benchmarking-nine-local-vlm-pipelines-on-handwritten-clinical-forms-38h6"&gt;&lt;em&gt;Confidence Is Theater: What I Learned Building Scribe&lt;/em&gt;&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Code and methodology: &lt;a href="https://github.com/kod201/scribe" rel="noopener noreferrer"&gt;github.com/kod201/scribe&lt;/a&gt; · My own benchmark gold set: &lt;a href="https://huggingface.co/datasets/sohakanu/scribe-flow-gold" rel="noopener noreferrer"&gt;sohakanu/scribe-flow-gold&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;All extraction ran locally on one machine; no page left it.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>llm</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Confidence Is Theater: What I Learned Building Scribe</title>
      <dc:creator>Stephen Ohakanu</dc:creator>
      <pubDate>Tue, 15 Sep 2026 16:29:54 +0000</pubDate>
      <link>https://dev.to/sohakanu/confidence-is-theater-benchmarking-nine-local-vlm-pipelines-on-handwritten-clinical-forms-38h6</link>
      <guid>https://dev.to/sohakanu/confidence-is-theater-benchmarking-nine-local-vlm-pipelines-on-handwritten-clinical-forms-38h6</guid>
      <description>&lt;p&gt;I'm building Scribe, a local pipeline that turns handwritten clinical forms into structured data.&lt;/p&gt;

&lt;p&gt;The obvious problem is reading the handwriting.&lt;/p&gt;

&lt;p&gt;The harder problem turned out to be this:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do you know when the model is wrong?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;My first idea was simple: ask the model how confident it was.&lt;/p&gt;

&lt;p&gt;That failed almost immediately.&lt;/p&gt;

&lt;p&gt;Median confidence when the answer was correct: &lt;strong&gt;0.95&lt;/strong&gt;.&lt;br&gt;
Median confidence when the answer was wrong: &lt;strong&gt;0.95&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;So confidence wasn't giving me a useful signal.&lt;/p&gt;

&lt;p&gt;That sent the project in a different direction: instead of asking a model whether I should trust it, I started comparing two independent reads of each field and treating disagreement as the signal to send it to a person.&lt;/p&gt;

&lt;p&gt;This is the engineering log of what happened next.&lt;/p&gt;
&lt;h2&gt;
  
  
  What Scribe does
&lt;/h2&gt;

&lt;p&gt;Scribe takes scanned or photographed handwritten clinical forms and converts them into structured records.&lt;/p&gt;

&lt;p&gt;The pipeline runs entirely on one machine. That matters for environments where connectivity is unreliable or the data should not leave the device.&lt;/p&gt;

&lt;p&gt;Each extracted value keeps a link back to where it came from on the page. If the system is not comfortable with a value, it sends that field for human review instead of quietly writing it into the record.&lt;/p&gt;

&lt;p&gt;The basic flow is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;page → find fields → read → validate → second read → human review if needed → structured record&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;There are more technical names for those stages, but that is really what the system is doing.&lt;/p&gt;
&lt;h2&gt;
  
  
  A quick note about the numbers
&lt;/h2&gt;

&lt;p&gt;I tested nine pipeline configurations on the same 23 handwritten pages and the same 221 hand-curated field values.&lt;/p&gt;

&lt;p&gt;The dataset used for this development bench is &lt;a href="https://huggingface.co/datasets/Nigeria-Health-data-OCR-pipeline/African-Medical-Records" rel="noopener noreferrer"&gt;public&lt;/a&gt;, which means some models may have seen similar material during training.&lt;/p&gt;

&lt;p&gt;So I'm treating these numbers as engineering comparisons between configurations, not as claims about the absolute performance of the models.&lt;/p&gt;

&lt;p&gt;Production sign-off will use a private held-out set created after model training cutoffs.&lt;/p&gt;
&lt;h2&gt;
  
  
  1. Whole page or crop?
&lt;/h2&gt;

&lt;p&gt;One of the first questions was whether the model should read the whole page or just the area containing the field.&lt;/p&gt;

&lt;p&gt;Both approaches had problems.&lt;/p&gt;

&lt;p&gt;Reading the whole page preserved context. Names, IDs and header-labeled fields were easier to interpret.&lt;/p&gt;

&lt;p&gt;But the extra context also created hallucinations. A model looking for a document date could "borrow" another date from a nearby table.&lt;/p&gt;

&lt;p&gt;Reading only a crop fixed some of that. For example, document-date hallucination dropped from 67% to 17%.&lt;/p&gt;

&lt;p&gt;But crops sometimes removed too much context. Hospital-number accuracy, for example, dropped from 87% to 65%.&lt;/p&gt;

&lt;p&gt;So I stopped trying to choose one.&lt;/p&gt;

&lt;p&gt;Scribe now routes by field type:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;tables, dates, enums and checkbox groups → crop&lt;/li&gt;
&lt;li&gt;scalars and free text → whole page&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That hybrid approach beat both pure modes on the combined quality measures and became the default.&lt;/p&gt;
&lt;h2&gt;
  
  
  2. Confidence didn't work
&lt;/h2&gt;

&lt;p&gt;The next problem was deciding which extracted values could safely skip human review. I initially planned to use the model's own confidence score for this, but as the opening numbers showed, confidence did not separate correct answers from mistakes on this test: even badly misread handwritten names came back at 0.95, and there was no threshold worth tuning. So instead of asking the model how sure it was, I tried reading each field a second time and flagging the ones where the two reads disagreed.&lt;/p&gt;
&lt;h2&gt;
  
  
  3. A second read helped much more than confidence
&lt;/h2&gt;

&lt;p&gt;The first verification approach simply re-read every field that would otherwise have been accepted automatically.&lt;/p&gt;

&lt;p&gt;The second read used the complementary source:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;first read from crop → verify from whole page&lt;/li&gt;
&lt;li&gt;first read from whole page → verify from crop&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If the two reads disagreed, the field went to human review.&lt;/p&gt;

&lt;p&gt;That one change moved the accuracy of automatically accepted fields from 79% to 93%, while silent errors dropped from 21% to 7%.&lt;/p&gt;

&lt;p&gt;The underlying model had not suddenly become a better reader.&lt;/p&gt;

&lt;p&gt;The pipeline had become better at recognizing when it should not trust the answer.&lt;/p&gt;

&lt;p&gt;That became the main design principle of Scribe.&lt;/p&gt;
&lt;h2&gt;
  
  
  4. Then I tried a different model for the second read
&lt;/h2&gt;

&lt;p&gt;There was still an obvious problem.&lt;/p&gt;

&lt;p&gt;If the same model has the same bias twice, two reads can agree on the same wrong answer.&lt;/p&gt;

&lt;p&gt;So I switched the verifier to a smaller, different model.&lt;/p&gt;

&lt;p&gt;That moved auto-accepted accuracy again, from 93% to 95%.&lt;/p&gt;

&lt;p&gt;More importantly, confidently misread patient names were now caught before they passed through automatically.&lt;/p&gt;

&lt;p&gt;The smaller model was not necessarily a better reader.&lt;/p&gt;

&lt;p&gt;Sometimes it was useful precisely because it was wrong in a different way.&lt;/p&gt;

&lt;p&gt;That was enough to surface disagreements.&lt;/p&gt;
&lt;h2&gt;
  
  
  5. Some mistakes were shared by nearly everyone
&lt;/h2&gt;

&lt;p&gt;The most stubborn example involved checkboxes.&lt;/p&gt;

&lt;p&gt;One synthetic form contained a lab-request section where nothing was checked. The correct answer was null.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fhuggingface.co%2Fdatasets%2Fsohakanu%2Fscribe-article-assets%2Fresolve%2Fmain%2Fcheckbox_trap_syn003.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fhuggingface.co%2Fdatasets%2Fsohakanu%2Fscribe-article-assets%2Fresolve%2Fmain%2Fcheckbox_trap_syn003.png" alt="An untouched checkbox row: Specimen type — Blood, Urine, Stool, CSF, all boxes empty" width="800" height="75"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Several model families still confidently chose an option.&lt;/p&gt;

&lt;p&gt;Qwen selected &lt;code&gt;"blood"&lt;/code&gt;. GLM showed the same kind of behavior. Gemma interpreted an empty box as a tick.&lt;/p&gt;

&lt;p&gt;Cross-model verification did not help much because both readers could share the same bias.&lt;/p&gt;

&lt;p&gt;That changed the engineering plan. Checkboxes stopped being treated as "just another prompt." They became their own detection problem.&lt;/p&gt;

&lt;p&gt;Later, Muse Glimmer produced an interesting counterexample: it hallucinated nothing in that run, but performed much worse overall. That makes it potentially useful as a specialist rather than as the main reader.&lt;/p&gt;

&lt;p&gt;That was a useful reminder:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A second model only helps when the models fail differently.&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  6. Changing the verifier eventually stopped helping
&lt;/h2&gt;

&lt;p&gt;After the initial gains, I tested several different verifier families.&lt;/p&gt;

&lt;p&gt;GLM-4.5V. Gemma 3. Gemma 4.&lt;/p&gt;

&lt;p&gt;They all ended up within roughly two fields of the small Qwen verifier on this 221-value set.&lt;/p&gt;

&lt;p&gt;GLM was actually worse operationally because it sent far more fields to review without improving the overall result much.&lt;/p&gt;

&lt;p&gt;That told me something useful. The next improvement probably wasn't hiding in yet another verifier.&lt;/p&gt;

&lt;p&gt;The remaining gains were more likely to come from:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a better primary model&lt;/li&gt;
&lt;li&gt;explicit checkbox detection&lt;/li&gt;
&lt;li&gt;resolving a few annotation conventions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So I stopped testing new verifiers.&lt;/p&gt;
&lt;h2&gt;
  
  
  7. The metric I care about changed too
&lt;/h2&gt;

&lt;p&gt;At first I tracked hallucination as one number: did the model invent a value where the correct answer was empty?&lt;/p&gt;

&lt;p&gt;Later I split that into two questions:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Did the model invent something?&lt;/li&gt;
&lt;li&gt;Did that invention escape all the safety checks?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The second one matters much more to the final system.&lt;/p&gt;

&lt;p&gt;For the routed pipeline without verification, 12% of null fields were invented and 12% escaped.&lt;/p&gt;

&lt;p&gt;With cross-model verification, the model still invented values, but only 2% escaped.&lt;/p&gt;

&lt;p&gt;That distinction changed how I think about evaluation.&lt;/p&gt;

&lt;p&gt;If you only measure raw hallucination, you are grading the model.&lt;/p&gt;

&lt;p&gt;If you measure what survives validation, verification and review, you are grading the system.&lt;/p&gt;

&lt;p&gt;And the system is what actually goes into production.&lt;/p&gt;
&lt;h2&gt;
  
  
  Where the bench ended up
&lt;/h2&gt;

&lt;p&gt;Across the nine configurations, raw exact accuracy only moved from roughly 75% to 81%.&lt;/p&gt;

&lt;p&gt;But the accuracy of fields allowed to bypass human review reached 96% in the current default configuration.&lt;/p&gt;

&lt;p&gt;That is the result I find most interesting.&lt;/p&gt;

&lt;p&gt;The project did not succeed because I found a model that suddenly became perfect at handwriting.&lt;/p&gt;

&lt;p&gt;It improved because the surrounding pipeline got better at recognizing when the model should not be trusted.&lt;/p&gt;
&lt;h2&gt;
  
  
  What still needs to be tested
&lt;/h2&gt;

&lt;p&gt;There are still three important open questions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Contamination.&lt;/strong&gt; The development corpus is public, so the next evaluation needs data the models could not have seen.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Production targets.&lt;/strong&gt; The target is at least 98% accuracy on automatically accepted fields and no more than 2% escaped hallucination, measured only on the private held-out set.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Generalization.&lt;/strong&gt; This one already has a first answer: in &lt;a href="https://dev.to/sohakanu/one-imperfect-page-per-hundred-ambushing-our-pipeline-with-data-nobody-tuned-fh8"&gt;part 2 of this series&lt;/a&gt;, I ran the two best configurations against 100 independently labeled handwritten prescription pages I had no hand in. Reading accuracy dropped sharply; the review system held.&lt;/p&gt;

&lt;p&gt;So the next useful number should come from genuinely unseen data, not from squeezing another decimal point out of this development set.&lt;/p&gt;
&lt;h2&gt;
  
  
  What I'm taking away from it
&lt;/h2&gt;

&lt;p&gt;A few things changed my thinking while building Scribe:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;model confidence was not useful for deciding whether an answer was safe&lt;/li&gt;
&lt;li&gt;disagreement between independent reads was much more useful&lt;/li&gt;
&lt;li&gt;crops and whole-page context solve different problems&lt;/li&gt;
&lt;li&gt;changing models does not fix every failure mode&lt;/li&gt;
&lt;li&gt;some problems need deterministic engineering rather than another prompt&lt;/li&gt;
&lt;li&gt;the most important hallucinations are the ones that escape the system&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Scribe now runs locally, and new form types are defined through schema YAML rather than custom extraction code. The evaluation setup also keeps model revisions, provenance and review decisions tied to each result.&lt;/p&gt;

&lt;p&gt;I'm still testing it.&lt;/p&gt;

&lt;p&gt;But the central lesson so far is simple:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The goal is not to build a model that never makes mistakes. It is to build a system that knows when a model's answer needs another look.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;
  The full scoreboard, for engineers
  &lt;br&gt;
All nine configurations, same 23 pages, same 221-value gold set. &lt;em&gt;Trusted-output acc&lt;/em&gt; = accuracy of fields accepted without human review. &lt;em&gt;Escaped&lt;/em&gt; = invented values that passed every check.

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Config&lt;/th&gt;
&lt;th&gt;Primary&lt;/th&gt;
&lt;th&gt;Verifier&lt;/th&gt;
&lt;th&gt;s/field&lt;/th&gt;
&lt;th&gt;Exact&lt;/th&gt;
&lt;th&gt;Trusted-output acc&lt;/th&gt;
&lt;th&gt;Flags&lt;/th&gt;
&lt;th&gt;Escaped&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;whole-page&lt;/td&gt;
&lt;td&gt;Qwen3-VL-30B-A3B&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;9.8&lt;/td&gt;
&lt;td&gt;75%&lt;/td&gt;
&lt;td&gt;75%&lt;/td&gt;
&lt;td&gt;10%&lt;/td&gt;
&lt;td&gt;20%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;crop-all&lt;/td&gt;
&lt;td&gt;Qwen3-VL-30B-A3B&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;4.0&lt;/td&gt;
&lt;td&gt;70%&lt;/td&gt;
&lt;td&gt;71%&lt;/td&gt;
&lt;td&gt;13%&lt;/td&gt;
&lt;td&gt;18%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;routed&lt;/td&gt;
&lt;td&gt;Qwen3-VL-30B-A3B&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;6.4&lt;/td&gt;
&lt;td&gt;76%&lt;/td&gt;
&lt;td&gt;79%&lt;/td&gt;
&lt;td&gt;14%&lt;/td&gt;
&lt;td&gt;12%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;+ same-model verify&lt;/td&gt;
&lt;td&gt;Qwen3-VL-30B-A3B&lt;/td&gt;
&lt;td&gt;itself&lt;/td&gt;
&lt;td&gt;9.9&lt;/td&gt;
&lt;td&gt;76%&lt;/td&gt;
&lt;td&gt;93%&lt;/td&gt;
&lt;td&gt;32%&lt;/td&gt;
&lt;td&gt;6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;+ cross-model verify&lt;/td&gt;
&lt;td&gt;Qwen3-VL-30B-A3B&lt;/td&gt;
&lt;td&gt;Qwen3-VL-8B&lt;/td&gt;
&lt;td&gt;10.0&lt;/td&gt;
&lt;td&gt;76%&lt;/td&gt;
&lt;td&gt;95%&lt;/td&gt;
&lt;td&gt;34%&lt;/td&gt;
&lt;td&gt;2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GLM verifier&lt;/td&gt;
&lt;td&gt;Qwen3-VL-30B-A3B&lt;/td&gt;
&lt;td&gt;GLM-4.5V&lt;/td&gt;
&lt;td&gt;11.9&lt;/td&gt;
&lt;td&gt;76%&lt;/td&gt;
&lt;td&gt;94%&lt;/td&gt;
&lt;td&gt;38%&lt;/td&gt;
&lt;td&gt;4%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;new primary (shipping)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Qwen3.8-27B&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Qwen3-VL-8B&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;17.3&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;81%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;96%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;28%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 3 verifier&lt;/td&gt;
&lt;td&gt;Qwen3.8-27B&lt;/td&gt;
&lt;td&gt;Gemma-3-12B&lt;/td&gt;
&lt;td&gt;19.8&lt;/td&gt;
&lt;td&gt;81%&lt;/td&gt;
&lt;td&gt;97%&lt;/td&gt;
&lt;td&gt;30%&lt;/td&gt;
&lt;td&gt;4%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 4 verifier&lt;/td&gt;
&lt;td&gt;Qwen3.8-27B&lt;/td&gt;
&lt;td&gt;Gemma-4-12B&lt;/td&gt;
&lt;td&gt;20.0&lt;/td&gt;
&lt;td&gt;81%&lt;/td&gt;
&lt;td&gt;96%&lt;/td&gt;
&lt;td&gt;29%&lt;/td&gt;
&lt;td&gt;2%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A tenth run (Muse Glimmer 30B as primary) read worse across the board (74% exact, 42% flags) but invented nothing: 0% hallucination.&lt;/p&gt;

&lt;p&gt;The pipeline in one picture:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fhuggingface.co%2Fdatasets%2Fsohakanu%2Fscribe-article-assets%2Fresolve%2Fmain%2Fpipeline.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fhuggingface.co%2Fdatasets%2Fsohakanu%2Fscribe-article-assets%2Fresolve%2Fmain%2Fpipeline.png" alt="Scribe pipeline: locate fields, read each from crop or page by type, validate, second-model check, then human review or record" width="800" height="462"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Hardware, pinned model revisions, per-field tables, metric definitions and one documented restatement: &lt;a href="https://github.com/kod201/scribe/blob/main/docs/bench.md" rel="noopener noreferrer"&gt;&lt;code&gt;docs/bench.md&lt;/code&gt;&lt;/a&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;br&gt;
&lt;p&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;This is part 1 of a series.&lt;/strong&gt; Part 2: &lt;a href="https://dev.to/sohakanu/one-imperfect-page-per-hundred-ambushing-our-pipeline-with-data-nobody-tuned-fh8"&gt;&lt;em&gt;One imperfect page per hundred&lt;/em&gt;&lt;/a&gt; — testing Scribe on an external prescription dataset.&lt;/p&gt;

&lt;p&gt;Code and methodology: &lt;a href="https://github.com/kod201/scribe" rel="noopener noreferrer"&gt;github.com/kod201/scribe&lt;/a&gt; · Benchmark gold set: &lt;a href="https://huggingface.co/datasets/sohakanu/scribe-flow-gold" rel="noopener noreferrer"&gt;sohakanu/scribe-flow-gold&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Every name on every sample page shown or discussed is fictional. All extraction ran locally on one machine; no page left it.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>llm</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
