<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ace-2504</title>
    <description>The latest articles on DEV Community by Ace-2504 (@ace2504).</description>
    <link>https://dev.to/ace2504</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3565080%2F4bdf8322-8b6b-403c-8b9b-8c8de533fdd4.jpg</url>
      <title>DEV Community: Ace-2504</title>
      <link>https://dev.to/ace2504</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ace2504"/>
    <language>en</language>
    <item>
      <title>The model had the right evidence. It still got the answer wrong.

I ran a controlled 60-question RAG experiment to find out whether retrieval or the reader was the real bottleneck. The results surprised me.

When Better Retrieval Doesn't Mean Better Answer</title>
      <dc:creator>Ace-2504</dc:creator>
      <pubDate>Sun, 23 Aug 2026 17:20:20 +0000</pubDate>
      <link>https://dev.to/ace2504/the-model-had-the-right-evidence-it-still-got-the-answer-wrong-i-ran-a-controlled-60-question-1l7m</link>
      <guid>https://dev.to/ace2504/the-model-had-the-right-evidence-it-still-got-the-answer-wrong-i-ran-a-controlled-60-question-1l7m</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/ace2504/fine-tuning-teaches-the-shape-retrieval-supplies-the-facts-4f3d" class="crayons-story__hidden-navigation-link"&gt;When Better Retrieval Doesn't Mean Better Answers.&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/ace2504" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3565080%2F4bdf8322-8b6b-403c-8b9b-8c8de533fdd4.jpg" alt="ace2504 profile" class="crayons-avatar__image" width="800" height="1422"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/ace2504" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Ace-2504
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Ace-2504
                
                
              
              &lt;div id="story-author-preview-content-4465940" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/ace2504" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3565080%2F4bdf8322-8b6b-403c-8b9b-8c8de533fdd4.jpg" class="crayons-avatar__image" alt="" width="800" height="1422"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Ace-2504&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/ace2504/fine-tuning-teaches-the-shape-retrieval-supplies-the-facts-4f3d" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Aug 23&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/ace2504/fine-tuning-teaches-the-shape-retrieval-supplies-the-facts-4f3d" id="article-link-4465940"&gt;
          When Better Retrieval Doesn't Mean Better Answers.
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/rag"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;rag&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/llm"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;llm&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/machinelearning"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;machinelearning&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
            &lt;a href="https://dev.to/ace2504/fine-tuning-teaches-the-shape-retrieval-supplies-the-facts-4f3d#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              &lt;span class="hidden s:inline"&gt;Add&amp;nbsp;Comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            10 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
      <category>ai</category>
      <category>llm</category>
      <category>rag</category>
    </item>
    <item>
      <title>When Better Retrieval Doesn't Mean Better Answers.</title>
      <dc:creator>Ace-2504</dc:creator>
      <pubDate>Sun, 23 Aug 2026 05:50:57 +0000</pubDate>
      <link>https://dev.to/ace2504/fine-tuning-teaches-the-shape-retrieval-supplies-the-facts-4f3d</link>
      <guid>https://dev.to/ace2504/fine-tuning-teaches-the-shape-retrieval-supplies-the-facts-4f3d</guid>
      <description>&lt;h2&gt;
  
  
  Fine-tuning teaches the shape. Retrieval supplies the facts.
&lt;/h2&gt;

&lt;p&gt;Everyone building RAG right now is tuning their retriever — better reranking, deeper retrieval, smarter chunking. I ran six experiments to see how much that helps once retrieval is already good.&lt;/p&gt;

&lt;p&gt;The short version: once the relevant passage was consistently reaching the context window, the retrieval-side changes I tried stopped moving answer quality — and a controlled test pointed at the reader, not the retriever, as the limiting factor. Here's how I got there, with numbers I can defend and the failures reported alongside the wins.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup: three systems, one model
&lt;/h2&gt;

&lt;p&gt;I used Yu-Gi-Oh as the test domain on purpose: thousands of cards, intricate timing rules, dense proper nouns, and a base model that's weak on it out of the box (closed-book, the base model scores just 1.83/5 on correctness and 0.18/2 on groundedness — see below). That weakness is a feature — it lets the three systems separate instead of all scoring the same.&lt;/p&gt;

&lt;p&gt;All three systems are the &lt;strong&gt;same&lt;/strong&gt; &lt;code&gt;google/gemma-2-2b-it&lt;/code&gt;. Only one lever changes at each step:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;System A — base.&lt;/strong&gt; The model as shipped, closed-book. Adapter off.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;System B — fine-tune.&lt;/strong&gt; A QLoRA adapter (rank 16, α 32, 4-bit NF4, all linear modules) trained on 2,683 teacher-distilled, judge-verified Yu-Gi-Oh Q&amp;amp;A pairs, early-stopped at the best validation checkpoint (~1 epoch; 3 epochs overfit). Still closed-book. &lt;em&gt;(Validation perplexity was 3.87. That measures next-token fit on held-out text, not factual correctness or grounding — it's a training-health number, not evidence about answer quality, and I don't use it as such.)&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;System C — fine-tune + retrieval.&lt;/strong&gt; System B, but a hybrid retriever prepends real passages before it answers.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Keeping it one model with an adapter toggle matters: any score difference comes from the lever I moved, not from a different model underneath.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwpzc5tzq55kfi8rvz7gr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwpzc5tzq55kfi8rvz7gr.png" alt=" " width="800" height="344"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  A three-layer way to read the results
&lt;/h2&gt;

&lt;p&gt;Before the numbers, a framing that makes the rest precise. "RAG" is usually drawn as &lt;em&gt;retrieval → answer&lt;/em&gt;, but there are three distinct layers, and they can fail independently:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Retrieval&lt;/strong&gt; — does the relevant evidence make it into the retrieved set at all? (Measured by &lt;a href="mailto:recall@k"&gt;recall@k&lt;/a&gt;.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Evidence utility&lt;/strong&gt; — does the retrieved context actually contain sufficient, usable evidence for &lt;em&gt;this&lt;/em&gt; question — not buried among distractors, not split across passages, not one fact among several competing ones?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reader&lt;/strong&gt; — can the model extract the right clause, reason over it, and produce the correct answer?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;High recall only speaks to Layer 1. It says nothing about whether Layer 2 is sufficient or whether Layer 3 can use it. Keeping these separate is what lets the results below say something defensible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Result 1: fine-tuning shifted the shape; retrieval drove the larger correctness gain
&lt;/h2&gt;

&lt;p&gt;All three systems answer the same 60 held-out questions, each scored 0–10 by a blind, reference-grounded judge (rubric details later). The headline, with paired statistics:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;System&lt;/th&gt;
&lt;th&gt;Mean ± SE&lt;/th&gt;
&lt;th&gt;vs. previous&lt;/th&gt;
&lt;th&gt;95% CI of difference&lt;/th&gt;
&lt;th&gt;p (t / Wilcoxon)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A — base, closed&lt;/td&gt;
&lt;td&gt;3.98 ± 0.39&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B — fine-tune, closed&lt;/td&gt;
&lt;td&gt;5.25 ± 0.54&lt;/td&gt;
&lt;td&gt;+1.27&lt;/td&gt;
&lt;td&gt;[0.42, 2.17]&lt;/td&gt;
&lt;td&gt;0.007 / 0.012&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C — fine-tune + retrieval&lt;/td&gt;
&lt;td&gt;8.05 ± 0.42&lt;/td&gt;
&lt;td&gt;+2.80&lt;/td&gt;
&lt;td&gt;[1.72, 3.87]&lt;/td&gt;
&lt;td&gt;&amp;lt;0.001 / &amp;lt;0.001&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4s2m0nnlsyjatn1qjv7x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4s2m0nnlsyjatn1qjv7x.png" alt=" " width="799" height="501"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Both step-ups are statistically significant on paired tests over the same 60 questions (these were the two comparisons I set out to make; treat them as strong exploratory evidence rather than a confirmatory benchmark).&lt;/p&gt;

&lt;p&gt;The more interesting signal is &lt;em&gt;where&lt;/em&gt; each jump comes from. Breaking the score into its four rubric components:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;System&lt;/th&gt;
&lt;th&gt;Correctness /5&lt;/th&gt;
&lt;th&gt;Completeness /2&lt;/th&gt;
&lt;th&gt;Groundedness /2&lt;/th&gt;
&lt;th&gt;Clarity /1&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A&lt;/td&gt;
&lt;td&gt;1.83&lt;/td&gt;
&lt;td&gt;0.97&lt;/td&gt;
&lt;td&gt;0.18&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B&lt;/td&gt;
&lt;td&gt;2.35&lt;/td&gt;
&lt;td&gt;1.03&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.87&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;3.85&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1.55&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1.65&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqu1z20kn9kawdvm91zqn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqu1z20kn9kawdvm91zqn.png" alt=" " width="800" height="497"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Fine-tuning (A → B) moved groundedness the most (0.18 → 0.87) — the model learned to &lt;em&gt;sound&lt;/em&gt; like a rulings answer and stopped confidently inventing card text. It also nudged correctness up (1.83 → 2.35). So fine-tuning did encode &lt;em&gt;some&lt;/em&gt; facts — I'm not claiming it can't. But the larger correctness gain came from retrieval (B → C: 2.35 → 3.85), where real passages entered the prompt.&lt;/p&gt;

&lt;p&gt;That's the first finding, stated carefully: &lt;strong&gt;in this setup, fine-tuning's main contribution was answer shape and grounding behavior, while retrieval contributed the larger factual-correctness gain.&lt;/strong&gt; "Fine-tuning teaches the shape, retrieval supplies the facts" is a useful mnemonic for that pattern — but it's a metaphor for a difference in &lt;em&gt;magnitude&lt;/em&gt; here, not a literal claim that fine-tuning cannot store facts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Result 2: once retrieval was good, the retrieval-side changes I tried didn't move answer quality
&lt;/h2&gt;

&lt;p&gt;System C averages 8.05/10, not 10 — the study reports 40 of 60 questions already at a perfect 10, so the lost points sit in a short tail. The retriever feeding it is already strong: hybrid dense (&lt;code&gt;all-MiniLM-L6-v2&lt;/code&gt;) + BM25 fused with Reciprocal Rank Fusion, top-5 over 42,412 chunks. &lt;strong&gt;recall@5 = 0.933&lt;/strong&gt;; the gold passage is at rank 1 about 72% of the time (recall@1 = 0.72) and in the top-5 for 93%. Recall does vary by source — rulings 0.96, card facts 0.93, archetype/lore lower on a small sub-sample — so "high recall" is an average, not uniform.&lt;/p&gt;

&lt;p&gt;So I tried to buy back the tail from the retrieval side. Six experiments:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Cross-encoder reranking&lt;/strong&gt; (retrieve top-20 → rerank → top-5). &lt;em&gt;Result:&lt;/em&gt; no help, and slightly worse — 8.05 → 7.55 on the main set; on a balanced 60-question split the drop (8.25 → 7.37) was statistically significant.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Was the reranker mis-ranking?&lt;/strong&gt; &lt;em&gt;Result:&lt;/em&gt; no — recall@5 was identical (0.933) and gold stayed at rank 1 either way. The drop tracked the small reader reacting to reshuffled context, not a ranking defect.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deeper retrieval&lt;/strong&gt; (rerank a wider candidate pool). &lt;em&gt;Result:&lt;/em&gt; recall rose (0.93 → 0.95 → 0.97) but the mean score did not move.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Chunk-repair&lt;/strong&gt; (a card's effect text was split across 1000-char chunks; I reconstructed it with adjacent-chunk expansion). &lt;em&gt;Result:&lt;/em&gt; it fixed the retrieval defect — the stranded clause was back in context — but the answer stayed incomplete; the reader didn't use the recovered clause. (This defect was rare: ~0.2% of cards.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Failure re-diagnosis by question type.&lt;/strong&gt; &lt;em&gt;Result:&lt;/em&gt; the remaining tail is dominated by yes/no reasoning inversions — the model flips the answer with the correct passage present — rather than by missing facts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Six in-model reader fixes&lt;/strong&gt; (self-consistency, self-verification, quote-then-answer, and more), tested on two hard cards. &lt;em&gt;Result:&lt;/em&gt; the in-model tricks didn't fix it; the one change that did was swapping in a stronger reader on the same context (next section).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The pattern across experiments 1–4: each change I tested &lt;em&gt;improved (or held) a retrieval metric while producing no measurable end-to-end gain&lt;/em&gt; on this evaluation — and reranking actually cost a little. That is evidence of &lt;strong&gt;diminishing end-to-end returns from retrieval optimization in this high-recall regime&lt;/strong&gt;, which is a narrower and more defensible statement than "retrieval stops mattering."&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F78f2zbvmyld7db30rxlu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F78f2zbvmyld7db30rxlu.png" alt=" " width="799" height="489"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Why would more recall not help? Because recall is a Layer-1 metric. The tail failures here look like Layer-2/Layer-3 problems: the evidence is present but the reader mis-reads it (the yes/no inversions), or a recovered clause is ignored. &lt;strong&gt;High retrieval recall does not guarantee successful evidence utilization&lt;/strong&gt; — and that gap is where the remaining points went.&lt;/p&gt;

&lt;h2&gt;
  
  
  The reader test — what it does and doesn't establish
&lt;/h2&gt;

&lt;p&gt;Experiment 6 is the most informative, because it holds the retrieved context fixed and changes only the reader:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;same question → same retrieved context → different reader → different answer.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;On the two hardest cards (Blackwing FAM and Endymion), a stronger reader given the &lt;em&gt;identical&lt;/em&gt; passages produced the complete, correct answer that the fine-tuned 2.6B reader kept under-delivering, including the exact effect clause it dropped.&lt;/p&gt;

&lt;p&gt;Here's the honest scope. This is a controlled, same-context comparison — stronger evidence for a reader limitation than simply observing that RAG helps. But it was run on &lt;strong&gt;two cards&lt;/strong&gt;, and the reference-grounded judge is stochastic run to run, so the study reports the result qualitatively (complete vs under-answered) rather than as fixed per-question scores. So the defensible conclusion is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This comparison provides evidence that reader capability is a limiting factor for at least some of the remaining failures — not proof that the reader is the single binding constraint across the whole tail.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The failure taxonomy in experiment 5 (a tail dominated by reasoning inversions with the passage present) is consistent with that reading, but it's a diagnosis, not a controlled manipulation. The clean way to settle it is an experiment I did &lt;strong&gt;not&lt;/strong&gt; run — see "The strongest next test."&lt;/p&gt;

&lt;h2&gt;
  
  
  What this suggests if you're building RAG
&lt;/h2&gt;

&lt;p&gt;Scoped to settings resembling this one (a small reader, a fact-dense domain, already-high recall):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;If recall is already high, &lt;strong&gt;verify that added retrieval effort is actually changing answers before investing in it&lt;/strong&gt; — in this study, reranking/deeper-retrieval/chunk-repair each moved retrieval metrics and not answer quality.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Recall is necessary, not sufficient.&lt;/strong&gt; A passage in the context window is not a fact in the answer; check evidence utility and reader behavior, not just &lt;a href="mailto:recall@k"&gt;recall@k&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;When retrieval is good and quality has plateaued, the higher-leverage move may be the reader — a stronger model, or RAG-aware training (the RAFT / preference-tuned line) — rather than a cleverer retriever.&lt;/li&gt;
&lt;li&gt;This does &lt;strong&gt;not&lt;/strong&gt; say retrieval is unimportant: 5.25 → 8.05 &lt;em&gt;was&lt;/em&gt; retrieval. It says retrieval's &lt;em&gt;marginal&lt;/em&gt; return fell off once recall was already high here.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And a diagnostic workflow that generalizes better than any single number:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1. Measure retrieval recall (Layer 1)
2. Inspect evidence sufficiency on failures (Layer 2)
3. Measure end-to-end answer quality
4. Test with oracle / gold context (isolates the retrieval ceiling)
5. Compare readers on identical context (isolates the reader ceiling)
6. Decide whether retrieval or reader is the dominant limitation — then spend effort there
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  How the evaluation was designed to reduce judge bias
&lt;/h2&gt;

&lt;p&gt;Every answer was scored by an LLM judge (&lt;code&gt;gemini-3.1-flash-lite&lt;/code&gt;) that is reference-grounded (handed the gold answer plus the verbatim evidence and told to grade against &lt;em&gt;that&lt;/em&gt;, not its own knowledge — which is what lets it fairly score a domain it wasn't trained on), blind (never shown which system produced an answer), pointwise (one answer at a time), and run over answers in shuffled order for extra blindness. The rubric sums to 10 (correctness 0–5, completeness 0–2, groundedness 0–2, clarity 0–1) with two guardrails: inventing a card detail forces groundedness to 0, and a correct refusal must outscore a confident wrong answer. Systems are compared paired (same 60 questions) with bootstrap 95% CIs cross-checked by paired t-test and Wilcoxon.&lt;/p&gt;

&lt;p&gt;What this design does &lt;strong&gt;not&lt;/strong&gt; yet establish, and I'd add before calling the measurement airtight: the judge temperature wasn't fixed, each answer was judged once (no repeated-judge consistency measured), and judge scores were never calibrated against human ratings. The study already notes the judge is stochastic per question; that variance is real and unquantified. Treat the 60-question means as reliable and individual per-question scores as noisy.&lt;/p&gt;

&lt;p&gt;The whole study cost about $3.26 in GPU and API spend.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limitations
&lt;/h2&gt;

&lt;p&gt;The boundaries of what this experiment tested, and where I would not extend the conclusions:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Single specialized domain&lt;/strong&gt; (Yu-Gi-Oh). Retrieval behavior and reader difficulty may differ substantially in legal, medical, financial, scientific-literature, enterprise-KB, multi-hop, or open-domain settings.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;60 held-out questions.&lt;/strong&gt; Enough for the paired step-ups to reach significance; too small to characterize RAG in general or to pin a universal recall threshold.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LLM-as-judge&lt;/strong&gt;, single pass, temperature unfixed, no human-agreement calibration.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Entity overlap not excluded.&lt;/strong&gt; The held-out set is built from &lt;em&gt;pages held out of training&lt;/em&gt; (a document-level split) and was decontaminated against training questions — so it tests generalization to unseen pages/questions. But the same cards/entities can appear across training and test pages, so this is not an &lt;em&gt;entity-disjoint&lt;/em&gt; evaluation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The reader finding rests on two cards&lt;/strong&gt; in a controlled swap, reported qualitatively.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Recall ≠ evidence utility&lt;/strong&gt; — I measured recall, and inferred utility failures from error analysis rather than measuring sufficiency directly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One retriever/reranker configuration&lt;/strong&gt; (MiniLM + BM25, RRF, one cross-encoder). Other configs may behave differently.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No oracle-context experiment was run&lt;/strong&gt;, so the retrieval ceiling and reader ceiling aren't cleanly separated numerically.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Findings are strongest for small readers in a high-recall regime&lt;/strong&gt;; they may not hold where recall is the actual bottleneck or where multi-hop reasoning dominates.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No human evaluation or cross-domain replication&lt;/strong&gt; — both would be needed to call this more than a case study.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Takeaway
&lt;/h2&gt;

&lt;p&gt;Fine-tuning changed how this model answered and improved grounding, but retrieval produced the larger factual-correctness gain. Once recall was already high, the retrieval-side changes I tested did not improve end-to-end answer quality, and a same-context reader swap on two hard cards pointed at reader capability as a limiting factor for at least some remaining failures.&lt;/p&gt;

&lt;p&gt;These findings come from a 60-question evaluation in a single specialized domain with one model and one retriever, so they are best read as a controlled engineering case study — not a universal recall threshold and not proof that readers always dominate retrieval. The transferable lesson is diagnostic: &lt;strong&gt;once relevant evidence is consistently reaching the context window, measure whether the reader can actually use it before spending more effort optimizing retrieval.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The strongest next test
&lt;/h2&gt;

&lt;p&gt;The cleanest way to turn the reader story from "suggested" into "demonstrated" is a 2×2 that separates the two ceilings — which I have not run:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Actual retrieved context&lt;/th&gt;
&lt;th&gt;Oracle / gold context&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Weak reader&lt;/strong&gt; (2.6B)&lt;/td&gt;
&lt;td&gt;measured (System C)&lt;/td&gt;
&lt;td&gt;isolates the retrieval ceiling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Strong reader&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;isolates residual reader gains&lt;/td&gt;
&lt;td&gt;upper bound with perfect evidence&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Comparing the columns shows how much quality is lost to imperfect evidence (retrieval ceiling); comparing the rows shows how much is lost to the reader (reader ceiling). Run across all 60 questions with fixed judge settings and repeated judging, it would replace the two-card qualitative result with a quantitative separation — and it's the experiment I'd prioritize before making any stronger claim.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Code, the 60-question eval set, the judge, and all six experiments: &lt;a href="https://github.com/Ace-2504/short-answers-broken-rag" rel="noopener noreferrer"&gt;github.com/Ace-2504/short-answers-broken-rag&lt;/a&gt;. Live arena where all three systems answer and a live judge scores them: &lt;a href="https://harman-ygo-slm.vercel.app" rel="noopener noreferrer"&gt;harman-ygo-slm.vercel.app&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>rag</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>ai</category>
    </item>
    <item>
      <title>The Bugs That Hide Behind "It Works": Debugging a Multithreaded C++ Proxy Server</title>
      <dc:creator>Ace-2504</dc:creator>
      <pubDate>Mon, 08 Jun 2026 08:37:27 +0000</pubDate>
      <link>https://dev.to/ace2504/the-bugs-that-hide-behind-it-works-debugging-a-multithreaded-c-proxy-server-1373</link>
      <guid>https://dev.to/ace2504/the-bugs-that-hide-behind-it-works-debugging-a-multithreaded-c-proxy-server-1373</guid>
      <description>&lt;p&gt;As a student, one of the biggest lessons I've picked up is that a program which compiles and runs is not the same as a program that is correct. It really hit home while I was building a multithreaded C++ network proxy server — a project that authenticates users with SHA-256, enforces role-based website filtering, and caches HTTP responses using a custom LRU cache.&lt;/p&gt;

&lt;p&gt;The happy path worked on pretty early. The real learning started when I went looking for the bugs that &lt;em&gt;don't&lt;/em&gt; announce themselves — the ones that only show up under concurrency, fragmented packets, or heavy load. Here are four of the most interesting and challenging issues I ran into and how I fixed them.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. The Thread Pool That Never Detached
&lt;/h2&gt;

&lt;p&gt;The proxy spawns 20 persistent worker threads at startup. The idea was simple: create the workers, then detach them from the main thread so they run independently.&lt;/p&gt;

&lt;p&gt;The bug was an ordering mistake. My detachment loop ran &lt;em&gt;before&lt;/em&gt; the threads were created:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight cpp"&gt;&lt;code&gt;&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;auto&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="n"&gt;workers&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;detach&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;   &lt;span class="c1"&gt;// workers is still EMPTY here&lt;/span&gt;
&lt;span class="n"&gt;workers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;emplace_back&lt;/span&gt;&lt;span class="p"&gt;(...);&lt;/span&gt;            &lt;span class="c1"&gt;// threads created afterward&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Because the vector was empty when the loop ran, the iteration did nothing — the threads were never actually detached. It's the kind of bug that compiles cleanly, passes a quick test, and then causes resource and lifecycle problems once the server runs for a while.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix:&lt;/strong&gt; I changed the order of the code. First, I created all the worker threads. Then, I detached them. What I learned: when working with threads, the order of setup steps matters. It is part of making the program correct, not just a small coding detail.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. A Data Race Hiding Inside the Logger
&lt;/h2&gt;

&lt;p&gt;My logging system looked thread-safe. Every write to &lt;code&gt;proxy.log&lt;/code&gt; was guarded by a &lt;code&gt;std::lock_guard&amp;lt;std::mutex&amp;gt;&lt;/code&gt;, so 20 threads writing at once could never scramble the file.&lt;/p&gt;

&lt;p&gt;But the timestamps came from &lt;code&gt;ctime(&amp;amp;now)&lt;/code&gt;. Under POSIX, &lt;code&gt;ctime&lt;/code&gt; returns a pointer to a &lt;strong&gt;globally shared static buffer&lt;/strong&gt;. My mutex protected the file stream — it did NOT protect that hidden global buffer. So two threads formatting timestamps at the same time could corrupt each other's strings, even though the file writes themselves were perfectly safe.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix:&lt;/strong&gt; switched to the thread-safe version, &lt;code&gt;ctime_r()&lt;/code&gt;, which writes into a buffer. The lesson was a subtle one: locking the &lt;em&gt;obvious&lt;/em&gt; shared resource isn't enough. You also have to think about the shared state hiding inside the standard library functions you're calling.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. The TCP Read That Assumed Too Much
&lt;/h2&gt;

&lt;p&gt;My request handler made a single call:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight cpp"&gt;&lt;code&gt;&lt;span class="n"&gt;recv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;client_socket&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;buffer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="n"&gt;BUFFER_SIZE&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This assumes that one &lt;code&gt;recv()&lt;/code&gt; gives you one complete HTTP request. TCP makes no such promise. TCP is a &lt;strong&gt;byte stream&lt;/strong&gt;, not a message protocol — a request can arrive split across several packets. If a header got cut in half, my &lt;code&gt;request.find("Host:")&lt;/code&gt; parsing would just fail, and a valid request would get dropped for no obvious reason.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix:&lt;/strong&gt; replaced the single read with a loop that keeps calling &lt;code&gt;recv()&lt;/code&gt; until the &lt;code&gt;\r\n\r\n&lt;/code&gt; header terminator has fully arrived. For my proxy, this handled the request headers; a fuller HTTP implementation would also need to handle bodies using &lt;code&gt;Content-Length&lt;/code&gt; or chunked transfer encoding. This turned out to be one of the most common (and most underestimated) mistakes in network programming: treating a stream like a neat sequence of separate messages.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Defending Against Hanging Connections
&lt;/h2&gt;

&lt;p&gt;Worker threads are a limited resource — I only have 20. A single unresponsive remote server that never closes its connection could tie up a worker forever, and 20 of those stalls would quietly take the whole proxy offline.&lt;/p&gt;

&lt;p&gt;My defense was setting a receive timeout with &lt;code&gt;setsockopt()&lt;/code&gt; using &lt;code&gt;SO_RCVTIMEO&lt;/code&gt;, set to one second. If a remote server goes quiet mid-response, the thread frees itself instead of hanging forever. Combined with proper HTTP status codes sent back to the client — &lt;code&gt;502 Bad Gateway&lt;/code&gt; when DNS resolution fails, &lt;code&gt;504 Gateway Timeout&lt;/code&gt; when the upstream stalls — the proxy fails gracefully instead of dying silently.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Debugging Concurrent Code Taught Me
&lt;/h2&gt;

&lt;p&gt;The common thread across all four bugs is the same: &lt;strong&gt;concurrency and networking break the assumptions that work perfectly in single-threaded, single-packet test runs.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Setup &lt;em&gt;order&lt;/em&gt; is part of correctness, not a detail.&lt;/li&gt;
&lt;li&gt;A mutex protects what it wraps — and nothing else, including hidden global state.&lt;/li&gt;
&lt;li&gt;The network gives you a byte stream, never a tidy message.&lt;/li&gt;
&lt;li&gt;Limited resources need timeouts, or one bad peer takes everything down.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of these bugs threw an error or crashed the compiler. They lived in the gap between "it runs" and "it is correct" — and for me, closing that gap has been the most rewarding part of learning systems programming.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;If you've built low-level networked systems in C++, I'd love to hear which subtle concurrency bug cost you the most time to track down.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>cpp</category>
      <category>systems</category>
      <category>networking</category>
      <category>learning</category>
    </item>
    <item>
      <title>Turning Behaviour into Climate Accountability</title>
      <dc:creator>Ace-2504</dc:creator>
      <pubDate>Tue, 24 Feb 2026 18:57:22 +0000</pubDate>
      <link>https://dev.to/ace2504/turning-behaviour-into-climate-accountability-3pph</link>
      <guid>https://dev.to/ace2504/turning-behaviour-into-climate-accountability-3pph</guid>
      <description>&lt;p&gt;One winter morning in Delhi, the AQI crossed 460 — officially hazardous. Schools shut. Hospitals filled. Transport slowed. Offices continued. Every day, millions of environmental decisions are made. Whether to drive or take public transport. Whether to allow remote work. Whether to burn waste or compost. These choices influence emissions — yet we rarely measure the pollution that did not happen because of them.&lt;/p&gt;

&lt;p&gt;That is the missing piece.&lt;/p&gt;

&lt;p&gt;Most environmental systems measure what exists: current emissions, fuel consumption, air quality levels. They do not formally measure avoided emissions. Counterfactual attribution addresses this gap by comparing two clearly defined scenarios:&lt;/p&gt;

&lt;p&gt;Business-as-usual baseline — what would have happened&lt;br&gt;
Verified alternative — what actually occurred&lt;/p&gt;

&lt;p&gt;Impact = Baseline − Verified Alternative.&lt;/p&gt;

&lt;p&gt;A Practical Example: Urban Commuting&lt;br&gt;
Consider a professional living 14 km from work. A 28 km daily round trip in a petrol car (≈ 120 g CO₂/km) results in about 3.36 kg CO₂ per day. Across 240 working days, that equals approximately 806 kg CO₂ annually. That is the baseline.&lt;/p&gt;

&lt;p&gt;Now introduce two behavioural shifts:&lt;/p&gt;

&lt;p&gt;Work from home two days per week → avoids ~322 kg per year&lt;br&gt;
Shift to electric metro on remaining days (≈ 15 g CO₂/km) → saves ~423 kg&lt;/p&gt;

&lt;p&gt;Total avoided emissions ≈ 745 kg CO₂ Remaining footprint ≈ 61 kg&lt;/p&gt;

&lt;p&gt;That is a reduction of over 90%, derived from a measurable behavioural delta — not offsets, not assumptions.&lt;/p&gt;

&lt;p&gt;If the remaining 61 kg is neutralised through verified sequestration (e.g., monitored urban trees at ~10 kg per tree annually), the commuting footprint approaches net zero.&lt;/p&gt;

&lt;p&gt;Baseline → Reduction → Net Impact. Each step is quantifiable.&lt;/p&gt;

&lt;p&gt;Preventing Double Counting&lt;br&gt;
A critical issue emerges: duplication.&lt;/p&gt;

&lt;p&gt;If an employee reports work-from-home reductions, and their employer's HR department also reports the same reduction in ESG disclosures, total claimed impact exceeds actual impact.&lt;/p&gt;

&lt;p&gt;To prevent this, the framework introduces a non-duplication constraint.&lt;/p&gt;

&lt;p&gt;Let:&lt;/p&gt;

&lt;p&gt;Δ_total = Verified avoided emissions&lt;br&gt;
C_i = Claim attributed to entity i&lt;/p&gt;

&lt;p&gt;The system enforces:&lt;/p&gt;

&lt;p&gt;Σ Cᵢ ≤ Δ_total&lt;/p&gt;

&lt;p&gt;No combination of claims may exceed the verified delta.&lt;/p&gt;

&lt;p&gt;Each reduction event is assigned:&lt;/p&gt;

&lt;p&gt;A unique registry ID&lt;br&gt;
A defined ownership tag&lt;br&gt;
A claim status flag&lt;/p&gt;

&lt;p&gt;Work-from-home emissions may be attributed to the enabling institution, while individual dashboards reflect behavioural contribution — without generating duplicate claimable credits unless formally allocated.&lt;/p&gt;

&lt;p&gt;This ensures:&lt;/p&gt;

&lt;p&gt;Measurability&lt;br&gt;
Additionality&lt;br&gt;
Non-duplication&lt;br&gt;
Audit defensibility&lt;/p&gt;

&lt;p&gt;Without duplication control, avoided-emission accounting collapses under verification.&lt;/p&gt;

&lt;p&gt;Enterprise Implications&lt;br&gt;
Most sustainability reports rely on estimated participation and averaged emission factors. A verification-first counterfactual architecture instead:&lt;/p&gt;

&lt;p&gt;Establishes defined baselines&lt;br&gt;
Verifies behavioural change in a privacy-preserving manner&lt;br&gt;
Models avoided emissions at corridor level&lt;br&gt;
Applies attribution constraints&lt;br&gt;
Produces audit-ready metrics&lt;/p&gt;

&lt;p&gt;Under such a system, work-from-home is not merely HR flexibility. It becomes measurable climate infrastructure.&lt;/p&gt;

&lt;p&gt;The Broader Shift&lt;br&gt;
This logic applies anywhere a conservative baseline can be defined:&lt;/p&gt;

&lt;p&gt;Agricultural burn avoidance&lt;br&gt;
Household waste diversion&lt;br&gt;
Fuel switching in buildings&lt;br&gt;
Distributed renewable adoption&lt;br&gt;
Industrial efficiency upgrades&lt;/p&gt;

&lt;p&gt;For decades, climate systems focused on measuring what we emit.&lt;/p&gt;

&lt;p&gt;The next frontier is measuring what we prevent — without inflation, without duplication, and without ambiguity.&lt;/p&gt;

&lt;p&gt;Climate accountability will not be built on slogans.&lt;/p&gt;

&lt;p&gt;It will be built on baselines, constraints, and verifiable delta.&lt;/p&gt;

&lt;p&gt;Patent Filed, 2026. A technical paper detailing the verification-first MRV architecture, statistical bias bounds, and attribution constraints is under preparation.&lt;/p&gt;

</description>
      <category>analytics</category>
      <category>data</category>
      <category>datascience</category>
      <category>science</category>
    </item>
    <item>
      <title>How McKinsey Frameworks Fixed My Scattered API</title>
      <dc:creator>Ace-2504</dc:creator>
      <pubDate>Sun, 04 Jan 2026 15:51:33 +0000</pubDate>
      <link>https://dev.to/ace2504/how-structured-thinking-fixed-my-api-design-a-case-study-i1n</link>
      <guid>https://dev.to/ace2504/how-structured-thinking-fixed-my-api-design-a-case-study-i1n</guid>
      <description>&lt;p&gt;While building the backend for &lt;strong&gt;Ace Rentals&lt;/strong&gt;, I realized that my authorization logic, although functionally correct, felt increasingly fragile. Ownership checks were scattered across multiple routes, duplicated in several places, and easy to forget when adding new endpoints. Over time, this made the system harder to reason about and increased the risk of subtle security gaps.&lt;/p&gt;

&lt;p&gt;Rather than continuing with incremental fixes, I stepped back and redesigned authorization as a system-level concern using centralized middleware. This article explains how I identified the issue, how I reasoned about the redesign, how it was implemented, and what improved as a result.&lt;/p&gt;

&lt;h2&gt;
  
  
  Previous Authorization Approach
&lt;/h2&gt;

&lt;p&gt;Authorization checks were implemented directly inside individual route handlers. Each protected endpoint contained its own logic to verify whether the current user was allowed to perform the requested action.&lt;/p&gt;

&lt;p&gt;While this approach worked initially, it introduced several challenges as the application grew:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The same ownership logic was copied across multiple routes
&lt;/li&gt;
&lt;li&gt;Route handlers mixed business logic with authorization concerns
&lt;/li&gt;
&lt;li&gt;Adding new endpoints required manual checks, increasing the chance of mistakes
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The system did not fail outright, but correctness increasingly depended on careful and repetitive implementation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Identifying the Design Issue
&lt;/h2&gt;

&lt;p&gt;The problem was not a missing check or a faulty condition—it was a structural issue. Authorization was treated as an implementation detail instead of a rule enforced consistently by the architecture.&lt;/p&gt;

&lt;p&gt;This meant security relied on developer memory rather than system guarantees. As the number of routes increased, so did the effort required to maintain consistency and confidence in the design.&lt;/p&gt;

&lt;h2&gt;
  
  
  Design Analysis and Decision-Making
&lt;/h2&gt;

&lt;p&gt;To avoid patching individual routes, I analyzed the problem at a design level.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;I used &lt;strong&gt;MECE-style analysis&lt;/strong&gt; to verify whether authorization was applied consistently across all relevant routes. This exposed gaps and overlap in enforcement.
&lt;/li&gt;
&lt;li&gt;I compared alternative approaches and found that &lt;strong&gt;centralized middleware&lt;/strong&gt; offered the best balance between reuse, clarity, and maintainability.
&lt;/li&gt;
&lt;li&gt;I set clear constraints for the refactor—no change in behavior, minimal surface area, and focused scope—to avoid unnecessary complexity.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This structured approach helped ensure the redesign addressed the root cause rather than its symptoms.&lt;/p&gt;

&lt;h2&gt;
  
  
  Authorization Redesign and Implementation
&lt;/h2&gt;

&lt;p&gt;Authorization was elevated to a system-level concern and implemented through middleware:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Ownership checks were consolidated into a single, reusable middleware function
&lt;/li&gt;
&lt;li&gt;API endpoints were reorganized around resources rather than actions, improving clarity
&lt;/li&gt;
&lt;li&gt;Authorization was enforced early in the request pipeline, before any business logic executed
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;With this setup, routes automatically inherit authorization rules. Developers no longer need to remember to reapply checks when adding or modifying endpoints.&lt;/p&gt;

&lt;h2&gt;
  
  
  Observed Impact of the Redesign
&lt;/h2&gt;

&lt;p&gt;The redesign produced clear and measurable improvements:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Authorization logic now exists in one place instead of being duplicated
&lt;/li&gt;
&lt;li&gt;All protected routes enforce authorization consistently by default
&lt;/li&gt;
&lt;li&gt;Maintenance effort and regression risk were significantly reduced
&lt;/li&gt;
&lt;li&gt;Resource relationships are clearer through resource-based routing
&lt;/li&gt;
&lt;li&gt;Route handlers are simpler and focus only on business logic
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A qualitative review showed that core domain logic is stable. Remaining risk is isolated to shared authorization middleware, where it is explicit, visible, and easier to reason about.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key Learning
&lt;/h2&gt;

&lt;p&gt;Repeated logic is often a signal of a deeper design issue. Treating security as a structural concern—rather than a manual step—leads to systems that are easier to maintain, extend, and trust over time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Design Decisions and Learnings
&lt;/h2&gt;

&lt;p&gt;This article provides a high-level summary of the decisions and outcomes from redesigning authorization in Ace Rentals.&lt;/p&gt;

&lt;p&gt;For a deeper look into my &lt;strong&gt;learning experience&lt;/strong&gt;—including diagrams, structured reasoning, trade-offs, and implementation details—you can read the full write-up here:&lt;/p&gt;

&lt;p&gt;👉 &lt;strong&gt;&lt;a href="https://ace-2504.github.io/api-design-blog/" rel="noopener noreferrer"&gt;https://ace-2504.github.io/api-design-blog/&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The full version documents the complete thought process and lessons learned while designing and refactoring the system.&lt;/p&gt;

</description>
      <category>api</category>
      <category>backend</category>
      <category>systemdesign</category>
      <category>javascript</category>
    </item>
  </channel>
</rss>
