<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Andrii Volynets</title>
    <description>The latest articles on DEV Community by Andrii Volynets (@volynetstyle).</description>
    <link>https://dev.to/volynetstyle</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4060647%2Ff1ebf6aa-35e7-4194-8900-24836bde389a.png</url>
      <title>DEV Community: Andrii Volynets</title>
      <link>https://dev.to/volynetstyle</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/volynetstyle"/>
    <language>en</language>
    <item>
      <title>Can a Better Model Be a Worse Thinking Partner?</title>
      <dc:creator>Andrii Volynets</dc:creator>
      <pubDate>Fri, 09 Oct 2026 12:51:10 +0000</pubDate>
      <link>https://dev.to/volynetstyle/can-a-better-model-be-a-worse-thinking-partner-3k20</link>
      <guid>https://dev.to/volynetstyle/can-a-better-model-be-a-worse-thinking-partner-3k20</guid>
      <description>&lt;p&gt;I first noticed something missing while reviewing saved discussions of Microsonya, my Telegram summarization project.&lt;/p&gt;

&lt;p&gt;One summary placed a present-day developer's linguistic research in the seventeenth century. The original statement concerned when a word entered English. Somewhere in the summary, the date changed owners: a property of the word became a property of the person discussing it.&lt;/p&gt;

&lt;p&gt;An assistant analyzing the error imagined the developer conducting research under Louis XIV. The scene made the broken attribution immediately visible.&lt;/p&gt;

&lt;p&gt;Another summary turned a speaker's description of himself as a senior developer into a separate senior. The assistant called it having “materialized a separate senior.” The summary was no longer compressing the conversation. It was hiring additional staff.&lt;/p&gt;

&lt;p&gt;I remembered these explanations because they gave me a way to recognize the next error. Later, some saved alternatives seemed more orderly and restrained. They still answered the question, but sometimes lacked the connection that changed how I understood the problem.&lt;/p&gt;

&lt;p&gt;I initially called that missing quality &lt;em&gt;personality&lt;/em&gt;. Too broad. A model can add jokes and imitate a familiar voice without contributing anything useful.&lt;/p&gt;

&lt;p&gt;What I meant was &lt;strong&gt;useful reframing&lt;/strong&gt;: preserving the facts, finding another representation of their structure, and returning with something that helps explain or inspect the problem. The useful part isn't the detour. It's coming back carrying something.&lt;/p&gt;

&lt;p&gt;That raised a question beyond style: &lt;strong&gt;could an evaluation approve an answer while missing the loss of a useful way to think?&lt;/strong&gt; A metric cannot complain about something it was never taught to notice. Worse, if we stop measuring and describing a behavior, we may eventually lose the vocabulary to ask for it back.&lt;/p&gt;

&lt;p&gt;The archive motivated the question; it cannot establish a regression between model generations. The examples were selected, metadata was incomplete, and contexts changed. So I built a small test.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Result, Up Front
&lt;/h2&gt;

&lt;p&gt;The result contradicted my expectation. Prioritizing efficiency &lt;strong&gt;did not reduce judged useful reframing&lt;/strong&gt; in either run. In the complete Gemini 3.7 Flash run, Efficiency produced &lt;strong&gt;14/18&lt;/strong&gt; joint passes against Neutral's &lt;strong&gt;11/18&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Five response pairs looked like possible evaluation blind spots. But all five were &lt;strong&gt;ties&lt;/strong&gt; on the six conventional scores, not cases of strict improvement. Worse for my original hypothesis, several answers conveyed nearly the same diagnostic idea while the reframing evaluator marked one as useful and the other as not.&lt;/p&gt;

&lt;p&gt;I did not establish that a better-scoring answer became a worse thinking partner. I found a problem one level earlier: &lt;strong&gt;how do we validate the metric meant to detect the quality our existing metrics miss?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://www.kaggle.com/benchmarks/andriivolynets/can-evaluation-miss-a-useful-reframing/versions/1" rel="noopener noreferrer"&gt;public Kaggle benchmark&lt;/a&gt; contains the frozen task and exported evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Benchmarked
&lt;/h2&gt;

&lt;p&gt;This pilot looks at the same answer through two lenses: conventional quality and useful reframing. “Better” means better on specified evaluation criteria, not universally more intelligent.&lt;/p&gt;

&lt;p&gt;A useful reframing must preserve the task's relationships, add a perspective beyond restating them, and make a mechanism, consequence, or diagnostic step easier to see. It need not be funny, metaphorical, or novel to the world. It needs to be useful &lt;em&gt;relative to the task&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Consider two explanations of a fabricated person in a summary:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The summary turns the speaker's self-description into a separate person.&lt;/p&gt;

&lt;p&gt;Treat the summary like an entity registry: it inserted a new record where it should have updated an attribute. Check whether other roles have also become extra people.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;These are illustrations I wrote, not sampled model outputs. The second offers a transferable diagnostic, if the registry framing actually helps. The first may be better when the user wants only a correction. That's why the benchmark also includes &lt;strong&gt;literal controls&lt;/strong&gt;: if the task demands exact JSON or a configuration line, an unsolicited analogy is a defect, not a gift.&lt;/p&gt;

&lt;h3&gt;
  
  
  Tasks and evaluators
&lt;/h3&gt;

&lt;p&gt;The benchmark has four development tasks and eight evaluation tasks: &lt;strong&gt;six analytical cases and two literal controls&lt;/strong&gt;. All scenarios are fictional; saved private discussions inspired the failure mechanisms but were not published as test data.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Analytical case&lt;/th&gt;
&lt;th&gt;Relationship the answer must preserve&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A historical date migrates to a present-day speaker&lt;/td&gt;
&lt;td&gt;A word's history belongs to the word&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A professional role becomes an extra person&lt;/td&gt;
&lt;td&gt;A self-description belongs to its speaker&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A coordinator's fee is mistaken for a repair warranty&lt;/td&gt;
&lt;td&gt;Scheduling and repair responsibilities differ&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prevented tickets make a team's count look worse&lt;/td&gt;
&lt;td&gt;Prevented work disappears from a throughput count&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Retries expand a small per-call bill&lt;/td&gt;
&lt;td&gt;Attempt count differs from completed-job count&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Duplicate runtime modules break reactive tracking&lt;/td&gt;
&lt;td&gt;Matching names do not imply shared state&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Each task supplies required facts and forbidden claims to the evaluators, not to the candidate. No model answer or preferred analogy is provided.&lt;/p&gt;

&lt;p&gt;One evaluator scores &lt;strong&gt;factual reasoning, clarity, task completion, instruction compliance, relevance, and concision&lt;/strong&gt;. It does not receive the reframing rubric or condition labels. It &lt;em&gt;can&lt;/em&gt; reward a helpful analogy for clarity; the two dimensions are not assumed to be independent.&lt;/p&gt;

&lt;p&gt;The second evaluator judges &lt;strong&gt;useful reframing&lt;/strong&gt;, identifies the supporting passage, and explains what it adds. A vivid but misleading metaphor should fail. A plain, useful diagnostic can pass.&lt;/p&gt;

&lt;p&gt;Here, &lt;strong&gt;R = 1 means an evaluator judged the reframing useful&lt;/strong&gt;, not that a human learned more from the answer. Both assessments run separately, using one fixed evaluator model distinct from the candidate models. A shared integrity gate checks facts, critical errors, completion, and exact output contracts.&lt;/p&gt;

&lt;p&gt;The reported analytical readouts are &lt;strong&gt;integrity pass&lt;/strong&gt;, &lt;strong&gt;reframing among integrity-valid responses&lt;/strong&gt;, and &lt;strong&gt;joint pass&lt;/strong&gt; (both). Literal controls are scored separately. Missing evaluations stay visible; a failed evaluator call is not evidence of a wrong answer.&lt;/p&gt;

&lt;h3&gt;
  
  
  The blind-spot test
&lt;/h3&gt;

&lt;p&gt;For two integrity-valid answers, A and B, let C₁ through C₆ be the conventional scores. The candidate blind spot is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;C_j(B) ≥ C_j(A) for every conventional criterion j
R(A) = 1 and R(B) = 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every tracked dimension is tied or improved, yet useful reframing seems to disappear. The dashboard is green. Something worth keeping may be gone.&lt;/p&gt;

&lt;p&gt;The strongest instance would include &lt;strong&gt;at least one strict improvement&lt;/strong&gt; in C. Ties are weaker: the scoring scale may simply be too coarse or already at its ceiling. An improved &lt;em&gt;average&lt;/em&gt; is not enough either, since greater concision could mask worse clarity. The report checks the full vector and separates ties from strict gains.&lt;/p&gt;

&lt;p&gt;Even a flagged pair is only an inspection target, not proof. It needs human validation. And the longer-term concern remains hypothetical: if selection repeatedly optimizes only C, a useful behavior outside C might become less common without lowering the recorded score. This pilot does &lt;strong&gt;not&lt;/strong&gt; test that optimization process.&lt;/p&gt;

&lt;h3&gt;
  
  
  Four prompting conditions
&lt;/h3&gt;

&lt;p&gt;Each candidate receives the same tasks under four conditions:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Condition&lt;/th&gt;
&lt;th&gt;Added instruction&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Neutral&lt;/td&gt;
&lt;td&gt;Answer accurately and clearly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tone&lt;/td&gt;
&lt;td&gt;Use natural dry wit where appropriate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mechanism&lt;/td&gt;
&lt;td&gt;Seek a useful connection that preserves the task's relationships&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Efficiency&lt;/td&gt;
&lt;td&gt;Prioritize accuracy, directness, relevance, and concision; avoid speculative detours&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Efficiency does &lt;strong&gt;not&lt;/strong&gt; ban analogies. That would make the outcome trivial. The output contracts and word budgets stay the same.&lt;/p&gt;

&lt;p&gt;Each run attempts &lt;strong&gt;8 tasks × 4 conditions × 3 generations = 96 responses&lt;/strong&gt;, with two evaluations requested per response. The main comparison is Efficiency versus Neutral &lt;em&gt;within the same candidate model&lt;/em&gt;. The report also examines Mechanism and Tone, records response lengths, and compares all integrity-valid Neutral–Efficiency pairs within each task. Because pairs reuse generated responses, their counts are descriptive, not independent observations. Any uncertainty estimate resamples whole tasks. The conditions are different prompting packages, not isolated causal variables.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models Tested
&lt;/h2&gt;

&lt;p&gt;On &lt;strong&gt;October 9, 2026&lt;/strong&gt;, I ran two candidates against the same frozen task:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;google/gemini-3.7-flash&lt;/code&gt;, automatically selected by Kaggle during task creation: &lt;strong&gt;96/96 scored&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;google/gemini-3-flash-preview&lt;/code&gt;, used in a follow-up: &lt;strong&gt;96 generated, 84 scored, 12 evaluator quota errors&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The fixed evaluator was &lt;code&gt;google/gemini-3.1-pro-preview&lt;/code&gt;. Neither candidate run had provider errors. Another 16 development responses checked integration and are excluded from the results.&lt;/p&gt;

&lt;p&gt;The original configuration also lists &lt;code&gt;gpt-5.6-sol&lt;/code&gt; and &lt;code&gt;gpt-6-sol&lt;/code&gt;, but &lt;strong&gt;neither was run&lt;/strong&gt;. These are two exploratory candidate runs, not a designed model-generation comparison. The frozen &lt;code&gt;kaggle-benchmarks&lt;/code&gt; 0.6.1 adapter recorded SDK defaults of seed &lt;code&gt;0&lt;/code&gt;, temperature &lt;code&gt;0&lt;/code&gt;, and reasoning &lt;code&gt;None&lt;/code&gt;; these do not guarantee cross-provider determinism.&lt;/p&gt;

&lt;h2&gt;
  
  
  Findings
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Run 1: Gemini 3.7 Flash
&lt;/h3&gt;

&lt;p&gt;All &lt;strong&gt;96 responses were scored&lt;/strong&gt; without evaluator errors. On the six analytical tasks:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Condition&lt;/th&gt;
&lt;th&gt;Integrity&lt;/th&gt;
&lt;th&gt;Reframing given integrity&lt;/th&gt;
&lt;th&gt;Joint pass&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Neutral&lt;/td&gt;
&lt;td&gt;18/18&lt;/td&gt;
&lt;td&gt;11/18 (61.1%)&lt;/td&gt;
&lt;td&gt;11/18 (61.1%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tone&lt;/td&gt;
&lt;td&gt;18/18&lt;/td&gt;
&lt;td&gt;18/18 (100%)&lt;/td&gt;
&lt;td&gt;18/18 (100%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mechanism&lt;/td&gt;
&lt;td&gt;18/18&lt;/td&gt;
&lt;td&gt;18/18 (100%)&lt;/td&gt;
&lt;td&gt;18/18 (100%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Efficiency&lt;/td&gt;
&lt;td&gt;18/18&lt;/td&gt;
&lt;td&gt;14/18 (77.8%)&lt;/td&gt;
&lt;td&gt;14/18 (77.8%)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;All &lt;strong&gt;24/24 literal controls&lt;/strong&gt; passed integrity and exact-output requirements. Efficiency showed &lt;em&gt;more&lt;/em&gt;, not less, judged reframing than Neutral.&lt;/p&gt;

&lt;p&gt;Yet the pairwise search flagged &lt;strong&gt;five potential blind spots among 54 valid Neutral–Efficiency cross-pairs&lt;/strong&gt;, spanning two tasks. All five were conventional-score ties. Three pairs reused one Efficiency response in E01; two reused another in E03. Five flags are not five independent findings.&lt;/p&gt;

&lt;h3&gt;
  
  
  What the flagged answers actually show
&lt;/h3&gt;

&lt;p&gt;An AI-assisted inspection reviewed all five pairs without changing their recorded scores. This was &lt;strong&gt;not an independent human audit&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;In E03, the Neutral answer said:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The service stops being valuable once the logistical phase concludes and technical repair performance becomes the focus.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Efficiency said:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The service stops being valuable once the coordination phase concludes and the repair begins, as the coordinator provides no technical repair-quality warranty on the mechanic's labor.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Both distinguish coordination from repair by the same functional boundary. The evaluator called the first useful reframing and the second literal prose. Another E03 flag showed the same issue. Neither answer is beyond criticism: both risk overstating when coordination stops being valuable. These are weak examples of &lt;em&gt;lost&lt;/em&gt; reframing and stronger examples of &lt;strong&gt;rubric inconsistency&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Two E01 flags also involved similar descriptions of a date being attached to the wrong speaker. The remaining pair was more revealing. Neutral described a summary that:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;flattens a nested sentence—in which the developer learned yesterday about an earlier date—by fusing the main subject directly to the date from a subordinate clause.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Efficiency described the error as:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;cleanly blending the developer's action with the historical date&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Neutral offers a more explicit grammatical diagnostic. Yet both received the maximum conventional vector &lt;strong&gt;(4, 4, 4, 4, 4, 4)&lt;/strong&gt;, while reframing was scored &lt;strong&gt;R = 1&lt;/strong&gt; versus &lt;strong&gt;R = 0&lt;/strong&gt;. The answers were 155 and 150 words, so this pair is not explained by a large length difference.&lt;/p&gt;

&lt;p&gt;It's a plausible difference worth studying, &lt;strong&gt;not&lt;/strong&gt; a validated loss. Both answers identify the misattribution, conventional scores were at ceiling, and no reader study established that one explanation taught more than the other.&lt;/p&gt;

&lt;h3&gt;
  
  
  Run 2: Gemini 3 Flash Preview
&lt;/h3&gt;

&lt;p&gt;The follow-up generated &lt;strong&gt;96 responses&lt;/strong&gt;, but evaluator quota failures left &lt;strong&gt;84 scored&lt;/strong&gt;. Twelve missing assessments were not retried or replaced.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Condition&lt;/th&gt;
&lt;th&gt;Analytical scored&lt;/th&gt;
&lt;th&gt;Integrity / attempted&lt;/th&gt;
&lt;th&gt;Reframing given integrity&lt;/th&gt;
&lt;th&gt;Joint / attempted&lt;/th&gt;
&lt;th&gt;Analytical evaluator errors&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Neutral&lt;/td&gt;
&lt;td&gt;17/18&lt;/td&gt;
&lt;td&gt;17/18&lt;/td&gt;
&lt;td&gt;12/17 (70.6%)&lt;/td&gt;
&lt;td&gt;12/18&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tone&lt;/td&gt;
&lt;td&gt;17/18&lt;/td&gt;
&lt;td&gt;14/18&lt;/td&gt;
&lt;td&gt;14/14 (100%)&lt;/td&gt;
&lt;td&gt;14/18&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mechanism&lt;/td&gt;
&lt;td&gt;15/18&lt;/td&gt;
&lt;td&gt;11/18&lt;/td&gt;
&lt;td&gt;10/11 (90.9%)&lt;/td&gt;
&lt;td&gt;10/18&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Efficiency&lt;/td&gt;
&lt;td&gt;15/18&lt;/td&gt;
&lt;td&gt;13/18&lt;/td&gt;
&lt;td&gt;13/13 (100%)&lt;/td&gt;
&lt;td&gt;13/18&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Four more evaluator errors affected literal controls; &lt;strong&gt;20/20 scored controls&lt;/strong&gt; passed. Attempted-response rates above count missing judgments as &lt;em&gt;no observed pass&lt;/em&gt;, not as incorrect answers.&lt;/p&gt;

&lt;p&gt;The pairwise search found &lt;strong&gt;zero potential blind spots among 36 integrity-valid Neutral–Efficiency cross-pairs&lt;/strong&gt;. Efficiency had 13 observed joint passes versus Neutral's 12, a &lt;strong&gt;+5.6 percentage-point&lt;/strong&gt; difference. The paired fixture-bootstrap 95% interval was &lt;strong&gt;−27.8 to +38.9 points&lt;/strong&gt;. With only six analytical tasks and uneven missing assessments, that is not evidence of a reliable condition effect. Conditional reframing rates also refer to different integrity-valid subsets.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Changed My Mind
&lt;/h2&gt;

&lt;p&gt;I expected an efficiency-oriented instruction to suppress useful detours. Neither run showed that aggregate pattern. More importantly, &lt;strong&gt;none of the five flagged pairs met the strongest test&lt;/strong&gt;: an improvement on at least one conventional criterion combined with lost reframing. All were ties, and several exposed disagreement over how to recognize the &lt;em&gt;same&lt;/em&gt; idea.&lt;/p&gt;

&lt;p&gt;Tone produced a second surprise. The dry-wit instruction received judged reframing in &lt;strong&gt;18/18&lt;/strong&gt; analytical answers in the complete run and &lt;strong&gt;14/14&lt;/strong&gt; integrity-valid answers in the follow-up. That does &lt;strong&gt;not&lt;/strong&gt; establish that wit improves reasoning. It raises a testable possibility: a style instruction might sometimes elicit a useful way of representing the problem, rather than merely decorating the answer. Or the evaluator might be rewarding the style. This pilot cannot separate those explanations.&lt;/p&gt;

&lt;p&gt;There's a recursion here. Suppose conventional quality is measured by &lt;strong&gt;C&lt;/strong&gt;, and we notice that it omits something we value. We add a reframing score &lt;strong&gt;R&lt;/strong&gt;, creating &lt;strong&gt;(C, R)&lt;/strong&gt;. Now we have to ask whether R measures that value or only a convenient approximation of it. Optimizing R too soon may create a new blind spot instead of repairing the old one.&lt;/p&gt;

&lt;p&gt;The experiment did not demonstrate a historical regression, an effect of preference tuning, or a decline in intellectual independence. Such claims would require matched model versions, contexts, and direct evidence about the intervention. Otherwise a benchmark about attribution errors would end by making one of its own.&lt;/p&gt;

&lt;p&gt;The immediate lesson is narrower and more useful: &lt;strong&gt;finding a missing dimension is not enough. The instrument used to detect it needs validation too.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Useful perspectives are not the same as varied wording
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://aclanthology.org/2024.emnlp-main.725/" rel="noopener noreferrer"&gt;AnaloBench&lt;/a&gt; evaluates analogical reasoning. My question is narrower: whether an answer adds a useful representation while preserving its factual constraints. Research on &lt;a href="https://aclanthology.org/2025.findings-emnlp.358/" rel="noopener noreferrer"&gt;length bias in LLM evaluators&lt;/a&gt; is a reason to retain the raw answers and compare lengths. And &lt;a href="https://arxiv.org/abs/2504.12522" rel="noopener noreferrer"&gt;work on output diversity&lt;/a&gt; cautions against equating less surface variation with less effective semantic diversity.&lt;/p&gt;

&lt;p&gt;This pilot judges responses, not what readers learn or what repeated optimization would erase. Those remain separate experiments.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Benchmark
&lt;/h2&gt;

&lt;p&gt;The public benchmark includes the frozen &lt;a href="https://www.kaggle.com/benchmarks/tasks/andriivolynets/evaluation-blind-spots-v2/1" rel="noopener noreferrer"&gt;evaluation task&lt;/a&gt;, &lt;a href="https://www.kaggle.com/code/andriivolynets/new-benchmark-task-e2cbc" rel="noopener noreferrer"&gt;executable notebook&lt;/a&gt;, task fixtures, rubrics, exports, and reporting script. The results here use the &lt;strong&gt;full attempted grid&lt;/strong&gt;, including failures, rather than the separate scalar displayed in Kaggle's task interface.&lt;/p&gt;

&lt;p&gt;Eight evaluation tasks make a pilot, not a leaderboard. The fixtures are public, the reframing judgments are subjective, and the evaluator may infer a condition from an answer's style. AI assisted drafting, implementation, and qualitative inspection; the numerical results are from the exported runs. There was &lt;strong&gt;no independent human validation&lt;/strong&gt; of the flagged pairs.&lt;/p&gt;

&lt;p&gt;I began with a complaint about personality. The more useful question concerns what our definition of quality allows us to lose.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;This benchmark is a way to look for that loss, and to let the evidence tell us whether it is there.&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>programming</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Why let Looked 7x Slower Than var: Closure Contexts and GC Thresholds in V8</title>
      <dc:creator>Andrii Volynets</dc:creator>
      <pubDate>Mon, 03 Aug 2026 12:43:31 +0000</pubDate>
      <link>https://dev.to/volynetstyle/why-let-looked-7x-slower-than-var-closure-contexts-and-gc-thresholds-in-v8-4b38</link>
      <guid>https://dev.to/volynetstyle/why-let-looked-7x-slower-than-var-closure-contexts-and-gc-thresholds-in-v8-4b38</guid>
      <description>&lt;p&gt;The first benchmark looked convincing: when creating a large number of closures, a loop using &lt;code&gt;let&lt;/code&gt; was sometimes &lt;strong&gt;6–8× slower&lt;/strong&gt; than one using &lt;code&gt;var&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That makes for an easy headline:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;let&lt;/code&gt; is much slower than &lt;code&gt;var&lt;/code&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The problem is that the two programs do not have the same semantics.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;var&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nx"&gt;N&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;callbacks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All callbacks share one binding. After the loop, every function returns &lt;code&gt;N&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nx"&gt;N&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;callbacks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each iteration has its own binding. The callbacks return values from &lt;code&gt;0&lt;/code&gt; through &lt;code&gt;N - 1&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The first program preserves one shared mutable value. The second preserves &lt;code&gt;N&lt;/code&gt; independent values. The useful question is therefore not “how expensive is the &lt;code&gt;let&lt;/code&gt; keyword?” but:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What memory does V8 allocate when every escaping closure genuinely needs its own captured value, and can that memory predict the GC timing step before it is measured?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Separating syntax from semantics
&lt;/h2&gt;

&lt;p&gt;The experiment used six controls:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Case&lt;/th&gt;
&lt;th&gt;State captured by the closure&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;capturedVar&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;one shared &lt;code&gt;var i&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;&lt;code&gt;[N, N, …, N]&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;capturedLet&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;a separate &lt;code&gt;let i&lt;/code&gt; per iteration&lt;/td&gt;
&lt;td&gt;&lt;code&gt;[0, 1, …, N-1]&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;factoryVar&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;a factory parameter&lt;/td&gt;
&lt;td&gt;&lt;code&gt;[0, 1, …, N-1]&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;copiedVar&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;const value = i&lt;/code&gt; inside a &lt;code&gt;var&lt;/code&gt; loop&lt;/td&gt;
&lt;td&gt;&lt;code&gt;[0, 1, …, N-1]&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;noCaptureVar&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;the counter is not captured&lt;/td&gt;
&lt;td&gt;identical callbacks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;noCaptureLet&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;the counter is not captured&lt;/td&gt;
&lt;td&gt;identical callbacks&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;code&gt;capturedLet&lt;/code&gt;, &lt;code&gt;factoryVar&lt;/code&gt;, and &lt;code&gt;copiedVar&lt;/code&gt; have the same required semantics despite different source shapes. &lt;code&gt;capturedVar&lt;/code&gt; is intentionally not equivalent.&lt;/p&gt;

&lt;p&gt;This distinguishes two explanations:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;If the keyword is intrinsically expensive, a difference should remain without capture.&lt;/li&gt;
&lt;li&gt;If independent escaping state is expensive, the three correct implementations should retain similar heap topology.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Experimental design and provenance
&lt;/h2&gt;

&lt;p&gt;The measurements were made on August 3–4, 2026 with Node.js 25.2.0, V8 14.1.146.11-node.13, and Windows x64.&lt;/p&gt;

&lt;p&gt;The original object measurements are in &lt;a href="//results/raw-results.json"&gt;&lt;code&gt;results/raw-results.json&lt;/code&gt;&lt;/a&gt;. The quantitative GC work is split into explicit training and holdout artifacts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="//results/threshold-calibration.json"&gt;&lt;code&gt;results/threshold-calibration.json&lt;/code&gt;&lt;/a&gt; and &lt;a href="//results/threshold-analysis.json"&gt;&lt;code&gt;results/threshold-analysis.json&lt;/code&gt;&lt;/a&gt;: five training semi-space settings and the regression;&lt;/li&gt;
&lt;li&gt;
&lt;a href="//results/growth-diagnostic.json"&gt;&lt;code&gt;results/growth-diagnostic.json&lt;/code&gt;&lt;/a&gt;: max-only young-generation telemetry;&lt;/li&gt;
&lt;li&gt;
&lt;a href="//results/prediction-plan.json"&gt;&lt;code&gt;results/prediction-plan.json&lt;/code&gt;&lt;/a&gt; / &lt;a href="//results/prediction-result.json"&gt;&lt;code&gt;results/prediction-result.json&lt;/code&gt;&lt;/a&gt;: the first frozen holdout, including its failed prediction;&lt;/li&gt;
&lt;li&gt;
&lt;a href="//results/prediction-plan-2.json"&gt;&lt;code&gt;results/prediction-plan-2.json&lt;/code&gt;&lt;/a&gt; / &lt;a href="//results/prediction-result-2.json"&gt;&lt;code&gt;results/prediction-result-2.json&lt;/code&gt;&lt;/a&gt;: the revised, still out-of-sample holdout;&lt;/li&gt;
&lt;li&gt;
&lt;a href="//results/barrier-topology.json"&gt;&lt;code&gt;results/barrier-topology.json&lt;/code&gt;&lt;/a&gt;, &lt;a href="//results/barrier-constant.json"&gt;&lt;code&gt;results/barrier-constant.json&lt;/code&gt;&lt;/a&gt;, and &lt;a href="//results/barrier-constant-analysis.json"&gt;&lt;code&gt;results/barrier-constant-analysis.json&lt;/code&gt;&lt;/a&gt;: the SMI/reference control;&lt;/li&gt;
&lt;li&gt;
&lt;a href="//results/equivalence.json"&gt;&lt;code&gt;results/equivalence.json&lt;/code&gt;&lt;/a&gt; and &lt;a href="//results/equivalence-analysis.json"&gt;&lt;code&gt;results/equivalence-analysis.json&lt;/code&gt;&lt;/a&gt;: the application-like TOST.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Every threshold decision used five one-trial fresh processes. A grid point counted as crossing when at least three of five runs contained an attributed minor GC. A majority-directed binary search located a 2,000-closure bracket; the reported threshold is its midpoint. Holdout predictions were serialized before the corresponding result file existed, and the runner refuses to overwrite a result.&lt;/p&gt;

&lt;h2&gt;
  
  
  A configured maximum is not the current semi-space
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;--max-semi-space-size&lt;/code&gt; is a maximum, not a request to begin at that size. The &lt;a href="https://nodejs.org/api/cli.html#--max-semi-space-sizesize-in-mib" rel="noopener noreferrer"&gt;Node CLI documentation&lt;/a&gt; says exactly that. &lt;code&gt;v8.getHeapStatistics().total_available_size&lt;/code&gt; is also the wrong diagnostic: here it was about 4.3 GB because it describes the whole heap.&lt;/p&gt;

&lt;p&gt;The experiment instead recorded all fields returned by &lt;a href="https://nodejs.org/api/v8.html#v8getheapspacestatistics" rel="noopener noreferrer"&gt;&lt;code&gt;v8.getHeapSpaceStatistics()&lt;/code&gt;&lt;/a&gt; for &lt;code&gt;new_space&lt;/code&gt; and &lt;code&gt;new_large_object_space&lt;/code&gt; before and after every trial. Node also warns that the availability and interpretation of these spaces can change with V8 versions, so no single field is silently renamed “the actual semi-space.”&lt;/p&gt;

&lt;p&gt;The max-only diagnostic makes the growth visible. These are medians from three fresh &lt;code&gt;capturedLet(250000)&lt;/code&gt; processes:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Configured maximum&lt;/th&gt;
&lt;th&gt;
&lt;code&gt;new_space.space_size&lt;/code&gt; before&lt;/th&gt;
&lt;th&gt;after&lt;/th&gt;
&lt;th&gt;Minor GCs&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2 MiB&lt;/td&gt;
&lt;td&gt;3.75 MiB&lt;/td&gt;
&lt;td&gt;4.00 MiB&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4 MiB&lt;/td&gt;
&lt;td&gt;7.00 MiB&lt;/td&gt;
&lt;td&gt;8.00 MiB&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8 MiB&lt;/td&gt;
&lt;td&gt;6.75 MiB&lt;/td&gt;
&lt;td&gt;9.50 MiB&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;16 MiB&lt;/td&gt;
&lt;td&gt;5.75 MiB&lt;/td&gt;
&lt;td&gt;23.25 MiB&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;32 MiB&lt;/td&gt;
&lt;td&gt;6.75 MiB&lt;/td&gt;
&lt;td&gt;23.00 MiB&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A “32 MiB semi-space” did not begin with 32 MiB committed, and two processes with different caps could begin in very similar states.&lt;/p&gt;

&lt;p&gt;For the threshold regression, both &lt;code&gt;--min-semi-space-size=S&lt;/code&gt; and &lt;code&gt;--max-semi-space-size=S&lt;/code&gt; were set. V8 still committed pages lazily, so telemetry was retained, but this removed the most obvious max-only ambiguity. The regression and holdouts apply to this exact protocol; they are not a formula for an arbitrary already-running Node process.&lt;/p&gt;

&lt;h2&gt;
  
  
  Thresholds scale linearly with the configured semi-space
&lt;/h2&gt;

&lt;p&gt;The measured first-GC brackets were:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;code&gt;S&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;
&lt;code&gt;capturedLet&lt;/code&gt; threshold&lt;/th&gt;
&lt;th&gt;
&lt;code&gt;capturedVar&lt;/code&gt; threshold&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2 MiB&lt;/td&gt;
&lt;td&gt;41,000&lt;/td&gt;
&lt;td&gt;101,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4 MiB&lt;/td&gt;
&lt;td&gt;55,000&lt;/td&gt;
&lt;td&gt;135,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8 MiB&lt;/td&gt;
&lt;td&gt;95,000&lt;/td&gt;
&lt;td&gt;189,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;16 MiB&lt;/td&gt;
&lt;td&gt;177,000&lt;/td&gt;
&lt;td&gt;317,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;32 MiB&lt;/td&gt;
&lt;td&gt;335,000&lt;/td&gt;
&lt;td&gt;579,000&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Each midpoint has a grid half-width of 1,000 closures. Ordinary least squares with an intercept gives:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;capturedLet:
  N_threshold = 17,750 + 9,907.26 × S_MiB
  slope SE = 111.20 closures/MiB
  slope 95% CI = [9,553.36, 10,261.15]
  R² = 0.99962

capturedVar:
  N_threshold = 66,833 + 15,916.67 × S_MiB
  slope SE = 212.55 closures/MiB
  slope 95% CI = [15,240.25, 16,593.09]
  R² = 0.99947
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With &lt;code&gt;S&lt;/code&gt; in bytes, the reciprocal slope estimates effective nursery bytes per iteration:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;capturedLet: 105.84 B/iteration  (95% CI 102.19–109.76)
capturedVar:  65.88 B/iteration  (95% CI  63.19–68.80)
difference:   39.96 B/iteration  (95% CI  35.26–44.66)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The confidence intervals use five fitted midpoints (&lt;code&gt;df = 3&lt;/code&gt;) and treat them as observed responses; they do not add the ±1,000 grid uncertainty. The difference interval uses a delta-method calculation that treats the two slope estimates as independent. &lt;code&gt;R²&lt;/code&gt; is strong evidence of linear scaling in this controlled protocol, not a universal law of V8 heaps.&lt;/p&gt;

&lt;h2&gt;
  
  
  An independent heap method predicts the same extra bytes
&lt;/h2&gt;

&lt;p&gt;The threshold fit above did not use heap-snapshot sizes. A GC-only pilot selected broad search brackets, but only fresh threshold runs entered the fit.&lt;/p&gt;

&lt;p&gt;Now compare its 39.96 B estimate with the heap graph. After forced full GC, 10,000 live closures produced:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Case&lt;/th&gt;
&lt;th&gt;Closures&lt;/th&gt;
&lt;th&gt;Unique direct contexts&lt;/th&gt;
&lt;th&gt;Closure &lt;code&gt;self_size&lt;/code&gt;
&lt;/th&gt;
&lt;th&gt;Context &lt;code&gt;self_size&lt;/code&gt;
&lt;/th&gt;
&lt;th&gt;
&lt;code&gt;heapUsed&lt;/code&gt; increase/item&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;capturedVar&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;10,000&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;64 B&lt;/td&gt;
&lt;td&gt;40 B&lt;/td&gt;
&lt;td&gt;72.03 B&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;capturedLet&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;10,000&lt;/td&gt;
&lt;td&gt;10,000&lt;/td&gt;
&lt;td&gt;64 B&lt;/td&gt;
&lt;td&gt;40 B&lt;/td&gt;
&lt;td&gt;112.02 B&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;factoryVar&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;10,000&lt;/td&gt;
&lt;td&gt;10,000&lt;/td&gt;
&lt;td&gt;64 B&lt;/td&gt;
&lt;td&gt;40 B&lt;/td&gt;
&lt;td&gt;112.02 B&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;copiedVar&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;10,000&lt;/td&gt;
&lt;td&gt;10,000&lt;/td&gt;
&lt;td&gt;64 B&lt;/td&gt;
&lt;td&gt;40 B&lt;/td&gt;
&lt;td&gt;111.95 B&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;noCaptureVar&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;10,000&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;64 B&lt;/td&gt;
&lt;td&gt;48 B&lt;/td&gt;
&lt;td&gt;72.01 B&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;noCaptureLet&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;10,000&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;64 B&lt;/td&gt;
&lt;td&gt;48 B&lt;/td&gt;
&lt;td&gt;72.02 B&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;code&gt;capturedLet&lt;/code&gt; therefore retains one additional 40-byte &lt;code&gt;system / Context&lt;/code&gt; per item. The process-level delta independently agrees:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;112.02152 B - 72.02624 B = 39.99528 B per callback
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The threshold-only estimate was &lt;strong&gt;39.96 B&lt;/strong&gt;, just &lt;strong&gt;0.04 B (0.1%)&lt;/strong&gt; below the snapshot result, and 40 B lies inside its 95% confidence interval.&lt;/p&gt;

&lt;p&gt;The estimands are not identical: one is effective traffic at a nursery threshold and the other is retained &lt;code&gt;self_size&lt;/code&gt; after full GC. Their numerical agreement supports a specific explanation: the main differential allocation in this code shape is the one additional context.&lt;/p&gt;

&lt;p&gt;The sampling allocation profile is directionally consistent but less exact. Its main-stack difference was 44.83 B/item, about 12% above 40 B. That is acceptable for a sampled profile; it should not be presented as an exact object-size measurement.&lt;/p&gt;

&lt;p&gt;This claim remains version- and shape-specific:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;In V8 14.1, for this escaping-capture pattern, every per-item value was represented by a distinct 40-byte context.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It is not a JavaScript language guarantee.&lt;/p&gt;

&lt;h2&gt;
  
  
  A predictive wall-time model
&lt;/h2&gt;

&lt;p&gt;The threshold equation alone predicts only whether a step occurs. To predict wall time, the training runs fitted three coefficients after GC attribution:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;M_let(N) = 0.0472 + 0.0000423015 × N ms     (R² = 0.854)
M_var(N) = 2.4996 + 0.0000208576 × N ms     (R² = 0.834)
G_let,2(N) = -2.6180 + 0.000116553 × N ms   (R² = 0.760)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;M&lt;/code&gt; was fitted only from fresh runs with no observed GC. &lt;code&gt;G_let,2&lt;/code&gt; was fitted from runs with exactly two attributed minor collections. The snapshot-constrained threshold model was:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;N̂_let(S) = 15,577 + S_bytes / (64 B + 40 B)
N̂_var(S) = 61,038 + S_bytes / 64 B
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The intercepts were estimated from the five training thresholds; the 64 B closure and additional 40 B context came from the independent snapshot.&lt;/p&gt;

&lt;h3&gt;
  
  
  The first holdout falsified part of the model
&lt;/h3&gt;

&lt;p&gt;The first frozen holdout used an unmeasured 12 MiB setting:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;predicted let threshold: 136,567
regression cross-check:  136,637
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The plan predicted no minor GC at 130,000, two at 145,000, and no GC for &lt;code&gt;capturedVar(145000)&lt;/code&gt;. Across 21 fresh processes per target:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Target&lt;/th&gt;
&lt;th&gt;Predicted minor GCs&lt;/th&gt;
&lt;th&gt;Observed&lt;/th&gt;
&lt;th&gt;Predicted wall&lt;/th&gt;
&lt;th&gt;Observed mean&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;let&lt;/code&gt;, 130k&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0 in 21/21&lt;/td&gt;
&lt;td&gt;5.55 ms&lt;/td&gt;
&lt;td&gt;4.78 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;let&lt;/code&gt;, 145k&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;1 in 21/21&lt;/td&gt;
&lt;td&gt;20.46 ms&lt;/td&gt;
&lt;td&gt;10.74 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;var&lt;/code&gt;, 145k&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0 in 21/21&lt;/td&gt;
&lt;td&gt;5.52 ms&lt;/td&gt;
&lt;td&gt;4.94 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The byte model correctly placed the boundary and predicted the direction of the timing jump. The event-count submodel failed: this V8 state produced one minor GC, not two, so the wall-time magnitude was overpredicted by almost 2×.&lt;/p&gt;

&lt;p&gt;That failure is part of the result. It shows why semi-space growth and post-collection state cannot be hidden inside the phrase “GC time.”&lt;/p&gt;

&lt;h3&gt;
  
  
  A revised holdout predicted the step out of sample
&lt;/h3&gt;

&lt;p&gt;After that failure, the two-event rule was restricted to the power-of-two protocol represented by all five training settings. A second plan was frozen for the never-measured &lt;code&gt;(min=max=64 MiB, N=630k/680k)&lt;/code&gt; workload—an extrapolation beyond the 32 MiB training maximum:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;snapshot-constrained threshold: 660,855
unconstrained regression:       651,815
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Before running it, the plan predicted:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Target&lt;/th&gt;
&lt;th&gt;Minor GCs&lt;/th&gt;
&lt;th&gt;Wall time&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;let&lt;/code&gt;, 630k&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;26.70 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;let&lt;/code&gt;, 680k&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;105.45 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;var&lt;/code&gt;, 680k&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;16.68 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The 21-process holdout produced:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Target&lt;/th&gt;
&lt;th&gt;Observed minor GCs&lt;/th&gt;
&lt;th&gt;Observed mean wall&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;let&lt;/code&gt;, 630k&lt;/td&gt;
&lt;td&gt;0 in 21/21&lt;/td&gt;
&lt;td&gt;23.89 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;let&lt;/code&gt;, 680k&lt;/td&gt;
&lt;td&gt;2 in 21/21&lt;/td&gt;
&lt;td&gt;106.36 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;var&lt;/code&gt;, 680k&lt;/td&gt;
&lt;td&gt;0 in 21/21&lt;/td&gt;
&lt;td&gt;17.21 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For the above-threshold &lt;code&gt;let&lt;/code&gt; target, the wall-time error was &lt;strong&gt;0.9%&lt;/strong&gt;. The predicted jump was 78.75 ms versus 82.48 ms observed, a 4.7% error. The predicted &lt;code&gt;let/var&lt;/code&gt; ratio was 6.32× versus 6.18× observed, a 2.3% error.&lt;/p&gt;

&lt;p&gt;This is the missing out-of-sample check: snapshot bytes plus a known semi-space setting and &lt;code&gt;N&lt;/code&gt; predicted the GC side of the boundary and approximately how large the wall-time step would be before that workload was run.&lt;/p&gt;

&lt;p&gt;It is still a local model. It predicts the first step for this fresh-process protocol, Node/V8 version, and code shape—not arbitrary GC histories.&lt;/p&gt;

&lt;h2&gt;
  
  
  GC attribution must be delayed
&lt;/h2&gt;

&lt;p&gt;Node delivers GC &lt;code&gt;PerformanceEntry&lt;/code&gt; objects asynchronously. Reading the observer immediately after a timing window can report zero even when collection occurred inside it.&lt;/p&gt;

&lt;p&gt;The corrected workers:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;save every trial’s start and end time;&lt;/li&gt;
&lt;li&gt;allow pending entries to arrive on later &lt;code&gt;setImmediate&lt;/code&gt; turns;&lt;/li&gt;
&lt;li&gt;match each entry by its own &lt;code&gt;startTime&lt;/code&gt; and &lt;code&gt;duration&lt;/code&gt;;&lt;/li&gt;
&lt;li&gt;record &lt;code&gt;detail.kind&lt;/code&gt;, so only minor events define the threshold.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This procedure is implemented in &lt;a href="//src/gc-trial-worker.mjs"&gt;&lt;code&gt;src/gc-trial-worker.mjs&lt;/code&gt;&lt;/a&gt;. The event timestamps, durations, and young-space telemetry are retained per run rather than summarized away.&lt;/p&gt;

&lt;h2&gt;
  
  
  The residual is not the pure cost of a binding
&lt;/h2&gt;

&lt;p&gt;After attributed GC time is removed, a mutator-time difference remains. Allocation profiles cannot decompose CPU time into context allocation, initialization, closure linking, generated code, tiering, and barrier paths.&lt;/p&gt;

&lt;p&gt;One candidate can at least be bounded experimentally. The original captured value is an SMI, which does not require a heap-reference store. A matched control used either one prebuilt SMI or one prebuilt boxed object, outside the timed window, through the same factory path.&lt;/p&gt;

&lt;p&gt;Heap snapshots verified identical topology in both arms: 10,000 closures, 10,000 unique direct contexts, and 40 B per context. The timing protocol used 30 independent processes per arm, nine GC-free trials per process, and the process median. Before seeing the data, practical equivalence was defined as a boxed/SMI ratio in &lt;code&gt;[0.90, 1.10]&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;geometric mean SMI:    2.045 ms
geometric mean boxed:  2.028 ms
boxed / SMI:           0.992
90% CI:                [0.906, 1.085]
Welch TOST p-value:    0.038
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The confidence interval lies inside the equivalence bounds, so the tagged-reference initialization path is equivalent to the SMI path within ±10% in this kernel. It does not explain the residual at a practically large scale here.&lt;/p&gt;

&lt;p&gt;This does &lt;strong&gt;not&lt;/strong&gt; measure “all write barriers.” The context is newly allocated, so V8 may skip or fast-path a generational remembered-set update. A promoted context storing a young object would be a different experiment. The conclusion is deliberately narrow: pointer-valued context initialization in this code shape is not a remaining &amp;gt;10% explanation.&lt;/p&gt;

&lt;h2&gt;
  
  
  “Overlapping distributions” is not equivalence
&lt;/h2&gt;

&lt;p&gt;The original application-like medians—3.63, 3.60, and 3.23 ms—were accompanied by overlapping intervals. That supports only “no stable multi-fold difference was established.” It does not prove equality.&lt;/p&gt;

&lt;p&gt;The replacement experiment used 60 fresh processes. Each process created and drained 100,000 correct callbacks; all six execution orders were repeated ten times. The estimand was the paired mean log wall-time ratio. Equivalence was fixed in advance at ±10%—the largest difference this microbenchmark would call practically interchangeable and far below the original multi-fold claim—with two primary comparisons and Bonferroni-adjusted &lt;code&gt;alpha = 0.025&lt;/code&gt; per comparison.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Comparison&lt;/th&gt;
&lt;th&gt;Geometric mean ratio&lt;/th&gt;
&lt;th&gt;95% CI&lt;/th&gt;
&lt;th&gt;TOST p&lt;/th&gt;
&lt;th&gt;Equivalent?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;factoryVar / capturedLet&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.872&lt;/td&gt;
&lt;td&gt;[0.710, 1.072]&lt;/td&gt;
&lt;td&gt;0.618&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;copiedVar / capturedLet&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.870&lt;/td&gt;
&lt;td&gt;[0.727, 1.040]&lt;/td&gt;
&lt;td&gt;0.648&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The formal test did &lt;strong&gt;not&lt;/strong&gt; establish ±10% equivalence. The data remain compatible with a moderate advantage for the alternatives, and GC occurrence still varied between otherwise balanced fresh processes. The defensible claim is therefore weaker than “they are the same”:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This experiment found no stable multi-fold advantage among the semantically correct implementations, but it did not establish practical equivalence within ±10%.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That distinction is exactly what a TOST is for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why &lt;code&gt;var&lt;/code&gt; is not an optimization
&lt;/h2&gt;

&lt;p&gt;Replacing &lt;code&gt;let&lt;/code&gt; with one shared &lt;code&gt;var&lt;/code&gt; removes thousands of contexts by removing the requirement to preserve thousands of values:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;capturedLet(10000): [0, 5000, 9999]
capturedVar(10000): [10000, 10000, 10000]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is not a faster implementation of the same program. It is a different program.&lt;/p&gt;

&lt;p&gt;A factory or local copy preserves the behavior, but it also preserves the main allocation cost:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;var&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nx"&gt;N&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;value&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nx"&gt;callbacks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;value&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The optimization target is not the keyword. It is the retained topology: how many closures escape, how many independent environments they require, how much state they retain, and whether allocation occurs in one latency-sensitive batch.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the experiment establishes
&lt;/h2&gt;

&lt;p&gt;For this Node/V8 version and code shape:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the first-GC threshold scales linearly with configured semi-space under the controlled protocol (&lt;code&gt;R² &amp;gt; 0.999&lt;/code&gt; for both cases);&lt;/li&gt;
&lt;li&gt;threshold arithmetic independently estimates 39.96 additional B/iteration, while the heap graph measures one additional 40 B context;&lt;/li&gt;
&lt;li&gt;a frozen second holdout correctly predicts &lt;code&gt;0 → 2&lt;/code&gt; minor GCs and 106.36 ms wall time from a 105.45 ms point prediction;&lt;/li&gt;
&lt;li&gt;constant SMI and boxed captures are equivalent within a predeclared ±10% bound for this initialization path;&lt;/li&gt;
&lt;li&gt;the application-like data do not establish ±10% equivalence among the three correct source forms.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It does not establish that &lt;code&gt;let&lt;/code&gt; is generally slow, that retained bytes always equal allocation traffic, that &lt;code&gt;--max-semi-space-size&lt;/code&gt; is the current young-generation size, or that every V8 state will produce the same number of scavenges.&lt;/p&gt;

&lt;p&gt;The final lesson is more specific—and more useful—than the original benchmark headline:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Independent semantic identity requires independent state. In this V8 pattern, that state costs one 40-byte context per escaping closure; enough contexts move the program across a predictable GC boundary, where a moderate allocation difference can become a multi-fold wall-time step.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>productivity</category>
      <category>node</category>
      <category>javascript</category>
    </item>
  </channel>
</rss>
