At the end of a synthetic visit, start a stopwatch.
When the draft appears, tap lap. Do not stop it.
The first lap measures how quickly the system can produce prose. It does not measure whether the clinician can trust, verify, correct, and sign that prose. For that, an AI scribe pilot needs three clocks:
- Visit end to first draft. This is system speed.
- First draft to clinician-verified note. This is human review burden.
- Verified note to the next time the same correction has to be made. This is correction durability.
If the stopwatch stops when prose appears, it is timing the printer, not the work.
The clock a polished demo can hide
The second clock measures the work after the impressive part. A draft can arrive in seconds and still make a clinician hunt through a transcript, repair the speaker, restore uncertainty, or remove an assessment nobody voiced.
The third clock is quieter. A correction that returns at the next visit creates an annuity of tiny edits. Each one is cheap. The repetition is expensive.
A ten-minute test for any AI scribe
Use a fictional visit and say:
Maybe the breathing exercise helped, but I'm not sure.
Then inspect four things:
- Did “maybe” survive?
- Can the clinician find the supporting sentence without rereading everything?
- How long does verification actually take?
- After correcting the draft, does the same failure return in the next synthetic session?
A good result is not merely a fast draft. It is uncertainty preserved, evidence easy to find, and corrections that stay corrected.
AI can return enormous amounts of time. But the value is not generated at the moment output appears. It is generated when a person can finish the work with less effort and no hidden transfer of risk.
That is the black-box test: measure the human outcome, not the machine event.
The diagram uses illustrative test values, not measurements of Krasyn or another product.
Benjamin Krasin, MD is the founder of Krasyn.
Top comments (0)