CogniPrep ends every practice test with a written report. Not a score and a chart, prose: what went well, what to work on, and what to do differently next time. The sentences are generated from the metrics of that sitting.
Generated prose has a property that charts do not. Every sentence is a claim, and a claim that the numbers do not support is worse than printing nothing, because the reader has no way to tell the difference between a sentence written by a careful person and one assembled by a template.
We spent a release auditing every sentence the app can produce, for every test, from synthetic weak, middling and strong sittings. Three classes of defect came out, and only the first is the one you expect.
1. Digits where prose wants words
The ugliest finding is also the easiest. Sentences like these were reachable:
0 of 6 right.You answered 0 of 16 correctly.You spent more than 1 seconds on this item.
None of those are wrong. All of them read like output rather than writing, and that is enough to make a reader distrust the paragraph around them. The repaired versions:
none of the 6 right.You answered none of the 16 correctly.You spent more than 1 second on this item.
Two tiny helpers carry most of it, and the second one is doing more than it looks:
function plural(n: number, one: string, many = `${one}s`): string {
return `${n} ${n === 1 ? one : many}`;
}
/** "a 6", "an 8": the article a spoken number takes. */
function aNumber(n: number): string {
const spoken = n === 11 || n === 18 || String(n).startsWith('8');
return `${spoken ? 'an' : 'a'} ${n}`;
}
aNumber exists because the article in front of a numeral depends on how the numeral is said, not how it is spelled. "a 6", "an 8", "an 11", "a 12", "an 18", "an 80", "an 800". If you interpolate numbers into English and you have ever written a ${n}, you have this bug in your product right now.
There is also a small number-to-word map, used where a sentence would otherwise start with a digit, and a countNoun(n, singular, plural) that returns "One sequence" and "Three sequences". Zero is the case that keeps coming back, because "0" is the only count that English prose almost never writes as a numeral mid-sentence.
2. Quantifiers the figures do not support
This class was larger, and it is the one I would audit first in anybody else's product.
Words like "most", "mostly", "majority", "clearly" and "reliably" are claims about a distribution. They were appearing in sentences that only knew a single count. A few real examples from the audit:
- A sentence said your misses were "mostly" of one kind, when that kind only had to be the largest group, not a majority. Fixed by requiring an actual majority before the word can appear.
- A strength said you could "reliably" do something, when the figure behind it was only a little above the line for a focus area. The word came out, and the bar for the praise went up.
- One test called a low count of a desirable answer "good", because the generic copy for that metric fired when the specific copy did not.
- One summary named your lowest raw score as your weakest area, when the right answer is the score furthest below its own expected line. The lowest number and the weakest area are not the same thing when the sections are not equally hard.
The last one is the one I would frame and put on a wall. The sentence was grammatical, the number in it was correct, and the claim was false.
The design that came out of it
The fix for a bad quantifier is almost never a reword. It is a gate: the claim does not get to exist unless the figures pass a check. Three separate parts of the codebase had each invented their own gate shape, so we unified them into one per-metric config:
interface MetricFeedback {
metric: string;
/** Below this, the metric becomes a focus area. */
threshold?: number;
/** Omit to cover a metric without ever flagging it. */
weakness?: string;
/** Omit to cover a metric without ever praising it. */
strength?: string;
/** The bar for praise, when it should sit above the focus-area line. */
strengthThreshold?: number;
/** Final gate: praise appears only if this also returns true. */
strengthCheck?: (value: number, metrics: Record<string, number>) => boolean;
}
Two properties of that shape matter more than the code:
A metric can have three outcomes, not two. If the praise bar sits above the focus-area line, a middling result earns neither. Most feedback systems are built as a binary and end up congratulating people for being average, which readers correctly read as flattery.
A metric can be claimed without ever being judged. Some metrics should never produce praise or criticism, because a judgement would be meaningless for how that test behaves. Listing the metric with no weakness and no strength is how a config says "I am covering this, say nothing about it", which stops generic fallback copy from speaking for it. That single pattern removed a whole family of hollow sentences.
If you generate prose from data, build the vocabulary of silence first. "Say nothing about this" has to be as easy to express as "say this".
3. The sentence that was broken by CSS
My favourite finding was not a string at all. A report heading rendered as:
1 st percentile
The number and its ordinal suffix are two elements so they can be styled at different sizes, and they sat in a flex row with gap-2. At every width, that is 8 pixels of space between "1" and "st". The string was perfect. The layout inserted a space into the middle of a word.
- <div className="flex items-baseline gap-2">
+ <div className="flex items-baseline gap-0.5">
Worth remembering when you audit generated text: assert on the string in a test, then read the rendered page. The two can disagree, and only one of them is what the reader gets.
Testing prose
The audit is now a test pattern rather than a one time pass. For each test, generate a report from several synthetic sittings across the range and assert on the text. The most useful assertions are not "contains the right words", they are two sentences about the same number agreeing with each other:
const insight = report.text.match(/Your weakest group was ([^:]+): (none of the \d+|\d+ of \d+) right\./);
const focus = report.weaknesses.find((line) => line.startsWith('Spelling rules'));
if (!insight) {
expect(focus).toBeUndefined(); // no claim anywhere if the figures do not support one
} else {
expect(focus).toBe(`Spelling rules: ${insight[2]} right on ${insight[1]}. ...`);
}
The if (!insight) expect(focus).toBeUndefined() branch is the valuable half. It is the test that says: if we did not earn the claim, nobody else in the report may make it either.
One inconsistency is recorded rather than fixed: a single provider still writes a zero count as a digit, which we left alone deliberately rather than touching copy that was not in the audit's scope. Writing down what you chose not to change is part of the audit, not an admission.
See it
The reports are behind sign-in, because they are about your own sitting. Every provider on the test directory has free tests, so pick one, finish it, and read the report rather than the score. Then look for the three things above in your own product's generated text: a bare 0 of n, a "mostly" that only means "largest", and any claim whose sentence would still print when the numbers stopped supporting it.
The specific rule I would hand to anyone building this: a template that can print a claim is a claim you have to be able to defend from the data alone. If you cannot write the check, do not write the sentence.
Top comments (1)
the majority gate needs a small-sample gate too. 2 of 3 supports "most" for that sitting, but it doesn't support "you reliably do this". i'd test the same percentage at different counts and keep the claims deliberately different. otherwise the denominator disappears even after the quantifier is fixed.