I first noticed something missing while reviewing saved discussions of Microsonya, my Telegram summarization project.
One summary placed a present-day developer's linguistic research in the seventeenth century. The original statement concerned when a word entered English. Somewhere in the summary, the date changed owners: a property of the word became a property of the person discussing it.
An assistant analyzing the error imagined the developer conducting research under Louis XIV. The scene made the broken attribution immediately visible.
Another summary turned a speaker's description of himself as a senior developer into a separate senior. The assistant called it having “materialized a separate senior.” The summary was no longer compressing the conversation. It was hiring additional staff.
I remembered these explanations because they gave me a way to recognize the next error. Later, some saved alternatives seemed more orderly and restrained. They still answered the question, but sometimes lacked the connection that changed how I understood the problem.
I initially called that missing quality personality. Too broad. A model can add jokes and imitate a familiar voice without contributing anything useful.
What I meant was useful reframing: preserving the facts, finding another representation of their structure, and returning with something that helps explain or inspect the problem. The useful part isn't the detour. It's coming back carrying something.
That raised a question beyond style: could an evaluation approve an answer while missing the loss of a useful way to think? A metric cannot complain about something it was never taught to notice. Worse, if we stop measuring and describing a behavior, we may eventually lose the vocabulary to ask for it back.
The archive motivated the question; it cannot establish a regression between model generations. The examples were selected, metadata was incomplete, and contexts changed. So I built a small test.
The Result, Up Front
The result contradicted my expectation. Prioritizing efficiency did not reduce judged useful reframing in either run. In the complete Gemini 3.7 Flash run, Efficiency produced 14/18 joint passes against Neutral's 11/18.
Five response pairs looked like possible evaluation blind spots. But all five were ties on the six conventional scores, not cases of strict improvement. Worse for my original hypothesis, several answers conveyed nearly the same diagnostic idea while the reframing evaluator marked one as useful and the other as not.
I did not establish that a better-scoring answer became a worse thinking partner. I found a problem one level earlier: how do we validate the metric meant to detect the quality our existing metrics miss?
The public Kaggle benchmark contains the frozen task and exported evidence.
What I Benchmarked
This pilot looks at the same answer through two lenses: conventional quality and useful reframing. “Better” means better on specified evaluation criteria, not universally more intelligent.
A useful reframing must preserve the task's relationships, add a perspective beyond restating them, and make a mechanism, consequence, or diagnostic step easier to see. It need not be funny, metaphorical, or novel to the world. It needs to be useful relative to the task.
Consider two explanations of a fabricated person in a summary:
The summary turns the speaker's self-description into a separate person.
Treat the summary like an entity registry: it inserted a new record where it should have updated an attribute. Check whether other roles have also become extra people.
These are illustrations I wrote, not sampled model outputs. The second offers a transferable diagnostic, if the registry framing actually helps. The first may be better when the user wants only a correction. That's why the benchmark also includes literal controls: if the task demands exact JSON or a configuration line, an unsolicited analogy is a defect, not a gift.
Tasks and evaluators
The benchmark has four development tasks and eight evaluation tasks: six analytical cases and two literal controls. All scenarios are fictional; saved private discussions inspired the failure mechanisms but were not published as test data.
| Analytical case | Relationship the answer must preserve |
|---|---|
| A historical date migrates to a present-day speaker | A word's history belongs to the word |
| A professional role becomes an extra person | A self-description belongs to its speaker |
| A coordinator's fee is mistaken for a repair warranty | Scheduling and repair responsibilities differ |
| Prevented tickets make a team's count look worse | Prevented work disappears from a throughput count |
| Retries expand a small per-call bill | Attempt count differs from completed-job count |
| Duplicate runtime modules break reactive tracking | Matching names do not imply shared state |
Each task supplies required facts and forbidden claims to the evaluators, not to the candidate. No model answer or preferred analogy is provided.
One evaluator scores factual reasoning, clarity, task completion, instruction compliance, relevance, and concision. It does not receive the reframing rubric or condition labels. It can reward a helpful analogy for clarity; the two dimensions are not assumed to be independent.
The second evaluator judges useful reframing, identifies the supporting passage, and explains what it adds. A vivid but misleading metaphor should fail. A plain, useful diagnostic can pass.
Here, R = 1 means an evaluator judged the reframing useful, not that a human learned more from the answer. Both assessments run separately, using one fixed evaluator model distinct from the candidate models. A shared integrity gate checks facts, critical errors, completion, and exact output contracts.
The reported analytical readouts are integrity pass, reframing among integrity-valid responses, and joint pass (both). Literal controls are scored separately. Missing evaluations stay visible; a failed evaluator call is not evidence of a wrong answer.
The blind-spot test
For two integrity-valid answers, A and B, let C₁ through C₆ be the conventional scores. The candidate blind spot is:
C_j(B) ≥ C_j(A) for every conventional criterion j
R(A) = 1 and R(B) = 0
Every tracked dimension is tied or improved, yet useful reframing seems to disappear. The dashboard is green. Something worth keeping may be gone.
The strongest instance would include at least one strict improvement in C. Ties are weaker: the scoring scale may simply be too coarse or already at its ceiling. An improved average is not enough either, since greater concision could mask worse clarity. The report checks the full vector and separates ties from strict gains.
Even a flagged pair is only an inspection target, not proof. It needs human validation. And the longer-term concern remains hypothetical: if selection repeatedly optimizes only C, a useful behavior outside C might become less common without lowering the recorded score. This pilot does not test that optimization process.
Four prompting conditions
Each candidate receives the same tasks under four conditions:
| Condition | Added instruction |
|---|---|
| Neutral | Answer accurately and clearly |
| Tone | Use natural dry wit where appropriate |
| Mechanism | Seek a useful connection that preserves the task's relationships |
| Efficiency | Prioritize accuracy, directness, relevance, and concision; avoid speculative detours |
Efficiency does not ban analogies. That would make the outcome trivial. The output contracts and word budgets stay the same.
Each run attempts 8 tasks × 4 conditions × 3 generations = 96 responses, with two evaluations requested per response. The main comparison is Efficiency versus Neutral within the same candidate model. The report also examines Mechanism and Tone, records response lengths, and compares all integrity-valid Neutral–Efficiency pairs within each task. Because pairs reuse generated responses, their counts are descriptive, not independent observations. Any uncertainty estimate resamples whole tasks. The conditions are different prompting packages, not isolated causal variables.
Models Tested
On October 9, 2026, I ran two candidates against the same frozen task:
-
google/gemini-3.7-flash, automatically selected by Kaggle during task creation: 96/96 scored. -
google/gemini-3-flash-preview, used in a follow-up: 96 generated, 84 scored, 12 evaluator quota errors.
The fixed evaluator was google/gemini-3.1-pro-preview. Neither candidate run had provider errors. Another 16 development responses checked integration and are excluded from the results.
The original configuration also lists gpt-5.6-sol and gpt-6-sol, but neither was run. These are two exploratory candidate runs, not a designed model-generation comparison. The frozen kaggle-benchmarks 0.6.1 adapter recorded SDK defaults of seed 0, temperature 0, and reasoning None; these do not guarantee cross-provider determinism.
Findings
Run 1: Gemini 3.7 Flash
All 96 responses were scored without evaluator errors. On the six analytical tasks:
| Condition | Integrity | Reframing given integrity | Joint pass |
|---|---|---|---|
| Neutral | 18/18 | 11/18 (61.1%) | 11/18 (61.1%) |
| Tone | 18/18 | 18/18 (100%) | 18/18 (100%) |
| Mechanism | 18/18 | 18/18 (100%) | 18/18 (100%) |
| Efficiency | 18/18 | 14/18 (77.8%) | 14/18 (77.8%) |
All 24/24 literal controls passed integrity and exact-output requirements. Efficiency showed more, not less, judged reframing than Neutral.
Yet the pairwise search flagged five potential blind spots among 54 valid Neutral–Efficiency cross-pairs, spanning two tasks. All five were conventional-score ties. Three pairs reused one Efficiency response in E01; two reused another in E03. Five flags are not five independent findings.
What the flagged answers actually show
An AI-assisted inspection reviewed all five pairs without changing their recorded scores. This was not an independent human audit.
In E03, the Neutral answer said:
The service stops being valuable once the logistical phase concludes and technical repair performance becomes the focus.
Efficiency said:
The service stops being valuable once the coordination phase concludes and the repair begins, as the coordinator provides no technical repair-quality warranty on the mechanic's labor.
Both distinguish coordination from repair by the same functional boundary. The evaluator called the first useful reframing and the second literal prose. Another E03 flag showed the same issue. Neither answer is beyond criticism: both risk overstating when coordination stops being valuable. These are weak examples of lost reframing and stronger examples of rubric inconsistency.
Two E01 flags also involved similar descriptions of a date being attached to the wrong speaker. The remaining pair was more revealing. Neutral described a summary that:
flattens a nested sentence—in which the developer learned yesterday about an earlier date—by fusing the main subject directly to the date from a subordinate clause.
Efficiency described the error as:
cleanly blending the developer's action with the historical date
Neutral offers a more explicit grammatical diagnostic. Yet both received the maximum conventional vector (4, 4, 4, 4, 4, 4), while reframing was scored R = 1 versus R = 0. The answers were 155 and 150 words, so this pair is not explained by a large length difference.
It's a plausible difference worth studying, not a validated loss. Both answers identify the misattribution, conventional scores were at ceiling, and no reader study established that one explanation taught more than the other.
Run 2: Gemini 3 Flash Preview
The follow-up generated 96 responses, but evaluator quota failures left 84 scored. Twelve missing assessments were not retried or replaced.
| Condition | Analytical scored | Integrity / attempted | Reframing given integrity | Joint / attempted | Analytical evaluator errors |
|---|---|---|---|---|---|
| Neutral | 17/18 | 17/18 | 12/17 (70.6%) | 12/18 | 1 |
| Tone | 17/18 | 14/18 | 14/14 (100%) | 14/18 | 1 |
| Mechanism | 15/18 | 11/18 | 10/11 (90.9%) | 10/18 | 3 |
| Efficiency | 15/18 | 13/18 | 13/13 (100%) | 13/18 | 3 |
Four more evaluator errors affected literal controls; 20/20 scored controls passed. Attempted-response rates above count missing judgments as no observed pass, not as incorrect answers.
The pairwise search found zero potential blind spots among 36 integrity-valid Neutral–Efficiency cross-pairs. Efficiency had 13 observed joint passes versus Neutral's 12, a +5.6 percentage-point difference. The paired fixture-bootstrap 95% interval was −27.8 to +38.9 points. With only six analytical tasks and uneven missing assessments, that is not evidence of a reliable condition effect. Conditional reframing rates also refer to different integrity-valid subsets.
What Changed My Mind
I expected an efficiency-oriented instruction to suppress useful detours. Neither run showed that aggregate pattern. More importantly, none of the five flagged pairs met the strongest test: an improvement on at least one conventional criterion combined with lost reframing. All were ties, and several exposed disagreement over how to recognize the same idea.
Tone produced a second surprise. The dry-wit instruction received judged reframing in 18/18 analytical answers in the complete run and 14/14 integrity-valid answers in the follow-up. That does not establish that wit improves reasoning. It raises a testable possibility: a style instruction might sometimes elicit a useful way of representing the problem, rather than merely decorating the answer. Or the evaluator might be rewarding the style. This pilot cannot separate those explanations.
There's a recursion here. Suppose conventional quality is measured by C, and we notice that it omits something we value. We add a reframing score R, creating (C, R). Now we have to ask whether R measures that value or only a convenient approximation of it. Optimizing R too soon may create a new blind spot instead of repairing the old one.
The experiment did not demonstrate a historical regression, an effect of preference tuning, or a decline in intellectual independence. Such claims would require matched model versions, contexts, and direct evidence about the intervention. Otherwise a benchmark about attribution errors would end by making one of its own.
The immediate lesson is narrower and more useful: finding a missing dimension is not enough. The instrument used to detect it needs validation too.
Useful perspectives are not the same as varied wording
AnaloBench evaluates analogical reasoning. My question is narrower: whether an answer adds a useful representation while preserving its factual constraints. Research on length bias in LLM evaluators is a reason to retain the raw answers and compare lengths. And work on output diversity cautions against equating less surface variation with less effective semantic diversity.
This pilot judges responses, not what readers learn or what repeated optimization would erase. Those remain separate experiments.
My Benchmark
The public benchmark includes the frozen evaluation task, executable notebook, task fixtures, rubrics, exports, and reporting script. The results here use the full attempted grid, including failures, rather than the separate scalar displayed in Kaggle's task interface.
Eight evaluation tasks make a pilot, not a leaderboard. The fixtures are public, the reframing judgments are subjective, and the evaluator may infer a condition from an answer's style. AI assisted drafting, implementation, and qualitative inspection; the numerical results are from the exported runs. There was no independent human validation of the flagged pairs.
I began with a complaint about personality. The more useful question concerns what our definition of quality allows us to lose.
This benchmark is a way to look for that loss, and to let the evidence tell us whether it is there.
Top comments (0)