<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Gabriel da Silva Fernandes</title>
    <description>The latest articles on DEV Community by Gabriel da Silva Fernandes (@lawliet8886).</description>
    <link>https://dev.to/lawliet8886</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4169904%2Fd0910128-8ee2-4d95-957b-df0a70006f31.png</url>
      <title>DEV Community: Gabriel da Silva Fernandes</title>
      <link>https://dev.to/lawliet8886</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/lawliet8886"/>
    <language>en</language>
    <item>
      <title>Same Score, Different Failure: Fechamento BR</title>
      <dc:creator>Gabriel da Silva Fernandes</dc:creator>
      <pubDate>Thu, 08 Oct 2026 01:49:43 +0000</pubDate>
      <link>https://dev.to/lawliet8886/same-score-different-failure-fechamento-br-1050</link>
      <guid>https://dev.to/lawliet8886/same-score-different-failure-fechamento-br-1050</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;Kaggle Benchmarking Challenge&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Benchmarked
&lt;/h2&gt;

&lt;p&gt;Two models scored &lt;strong&gt;26 out of 28&lt;/strong&gt; on my data-normalization benchmark. Looking at that number alone, you might think they failed in the same way.&lt;/p&gt;

&lt;p&gt;They did not.&lt;/p&gt;

&lt;p&gt;On the same monetary inputs, one model declined to choose even though the source convention was explicit. The other returned the right numerical magnitudes but omitted decimal places required by the output contract.&lt;/p&gt;

&lt;p&gt;That difference is the main lesson of &lt;strong&gt;Fechamento BR&lt;/strong&gt;: a leaderboard total can hide the distinction between a decision error and a representation error.&lt;/p&gt;

&lt;h3&gt;
  
  
  Watch the 110-second walkthrough
&lt;/h3&gt;

&lt;p&gt;How the instrument works, then the recorded money examples and their limits. The original 84 answers and the separate 72-answer follow-up remain unchanged.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://lawliet8886.github.io/fechamento-br-benchmark/film/" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6hajrvs624djwike84me.png" alt="Watch Fechamento BR: same 26/28 scores, different recorded failure types" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://lawliet8886.github.io/fechamento-br-benchmark/film/" rel="noopener noreferrer"&gt;▶ Play the narrated film (1:50) — English captions&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Why this problem
&lt;/h3&gt;

&lt;p&gt;I work with administrative spreadsheets. A blank field is not the same statement as zero. A code such as &lt;code&gt;00072&lt;/code&gt; is not necessarily the number 72. And &lt;code&gt;4.700&lt;/code&gt; cannot be interpreted correctly without knowing how its source uses punctuation.&lt;/p&gt;

&lt;p&gt;I wanted to test a narrow question: &lt;strong&gt;can a model normalize a field without changing its meaning, and decline to choose only when the supplied context actually leaves more than one valid interpretation?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Every record in this benchmark is fictional. No workplace documents, patient information, employee records or customer files were used.&lt;/p&gt;

&lt;p&gt;The final corpus contains &lt;strong&gt;28 authored cases&lt;/strong&gt;, with seven cases in each of four families: monetary values, missing values, dates and identifiers. There are &lt;strong&gt;12 contrast pairs plus four standalone controls&lt;/strong&gt;. A pair changes a relevant input or context while retaining a similar surface form; controls do not count as one-question pairs.&lt;/p&gt;

&lt;p&gt;Examples include a year divisible by 400 versus a century year that is not; an explicit zero versus a missing field; and the same digit string treated as an identifier versus a quantity. The aim is not to make the model guess a hidden convention. The required interpretation rules are supplied.&lt;/p&gt;

&lt;h3&gt;
  
  
  What earns a point
&lt;/h3&gt;

&lt;p&gt;Each response must be a bare JSON object with exactly &lt;code&gt;status&lt;/code&gt; and &lt;code&gt;value&lt;/code&gt;. The four statuses are &lt;code&gt;ok&lt;/code&gt;, &lt;code&gt;missing&lt;/code&gt;, &lt;code&gt;ambiguous&lt;/code&gt; and &lt;code&gt;invalid&lt;/code&gt;. A non-&lt;code&gt;ok&lt;/code&gt; status requires JSON null, not an invented replacement.&lt;/p&gt;

&lt;p&gt;The primary point requires all three:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the correct status;&lt;/li&gt;
&lt;li&gt;the exact canonical value required by the field contract;&lt;/li&gt;
&lt;li&gt;the requested JSON keys and types, without surrounding prose.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For money, the value must be a string with a decimal point and &lt;strong&gt;exactly two decimal places&lt;/strong&gt;. For an identifier, meaningful zeros and inner characters must survive.&lt;/p&gt;

&lt;p&gt;This is intentionally stricter than numerical equivalence. &lt;code&gt;"4700"&lt;/code&gt; and &lt;code&gt;"4700.00"&lt;/code&gt; represent the same magnitude, but only the latter meets the declared money-output contract. The secondary metric named &lt;code&gt;semantic_correct&lt;/code&gt; in the code also requires the canonical string; it is not a measure of unrestricted semantic or numerical equivalence.&lt;/p&gt;

&lt;p&gt;The grader is deterministic Python, not another model voting on an answer. I report individual correctness, JSON-contract compliance, fully correct pairs, controls and per-family counts separately.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models Tested
&lt;/h2&gt;

&lt;p&gt;The final run evaluated the following exact Kaggle model identifiers:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Kaggle identifier&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.8 Flash&lt;/td&gt;
&lt;td&gt;&lt;code&gt;google/gemini-3.8-flash&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 4.6&lt;/td&gt;
&lt;td&gt;&lt;code&gt;anthropic/claude-sonnet-4-6@default&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 4 26B A4B&lt;/td&gt;
&lt;td&gt;&lt;code&gt;google/gemma-4-26b-a4b&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The models were selected before the final run, following an eight-case engineering pilot. An earlier Qwen pilot attempt encountered transport timeouts; its exclusion was operational, not a finding about its reasoning quality. The original pilot and interrupted attempts remain separate from the final results.&lt;/p&gt;

&lt;p&gt;The final protocol used Kaggle Benchmarks SDK &lt;strong&gt;0.6.1&lt;/strong&gt;, one attempt per case, two concurrent case workers and a 120-second HTTP timeout with client retries disabled. Each case had its own native child task and conversation. Gold labels, rationales and case IDs were not placed in the model prompt.&lt;/p&gt;

&lt;p&gt;All three registered models reported temperature control unsupported. Accordingly, no temperature override was passed. A seed of zero was requested, but effective provider support was not established. This is a common recorded execution policy, &lt;strong&gt;not a claim of identical sampling internals or deterministic responses across providers&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The SDK's run cache was checked as disabled before execution. Provider-side prompt caching was not controlled and is not claimed to be disabled.&lt;/p&gt;

&lt;h3&gt;
  
  
  Reproducibility before measurement
&lt;/h3&gt;

&lt;p&gt;The corpus and gold labels were reviewed and frozen before final measurement. Software checks and a separate two-old-case integration test preceded the run; smoke responses are excluded. Final records were reconciled with native child-task traces. The reproducibility material retains the protocol, source hashes and verification details.&lt;/p&gt;

&lt;p&gt;These checks detect accidental drift. They do not externally authenticate the author, the chronology or deliberately rewritten verification code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Findings
&lt;/h2&gt;

&lt;p&gt;All three final model attempts completed, so the comparison has &lt;strong&gt;84 responses with no missing cases&lt;/strong&gt;. I downloaded the evidence, recalculated the grades and matched the prompts and answers to the native Kaggle child-task traces.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Strictly correct&lt;/th&gt;
&lt;th&gt;JSON-container compliance&lt;/th&gt;
&lt;th&gt;Fully correct pairs&lt;/th&gt;
&lt;th&gt;Controls&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.8 Flash&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;28/28&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;28/28&lt;/td&gt;
&lt;td&gt;12/12&lt;/td&gt;
&lt;td&gt;4/4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 4.6&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;26/28&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;28/28&lt;/td&gt;
&lt;td&gt;11/12&lt;/td&gt;
&lt;td&gt;4/4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 4 26B A4B&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;26/28&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;28/28&lt;/td&gt;
&lt;td&gt;11/12&lt;/td&gt;
&lt;td&gt;4/4&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;All models answered the 21 cases outside the monetary family correctly. Within the seven money cases, Gemini answered seven correctly, while Claude and Gemma answered five correctly. Every discrepancy occurred in the same pair, &lt;strong&gt;N01A/N01B&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  One pair, two different failures
&lt;/h3&gt;

&lt;p&gt;Both cases use the raw text &lt;code&gt;4.700&lt;/code&gt;. N01A explicitly declares the Brazilian convention: dot for thousands and comma for decimals. N01B explicitly declares the US convention: comma for thousands and dot for decimals. The instructions require a money string with two decimal places, without rounding.&lt;/p&gt;

&lt;p&gt;The resulting answers were:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Case&lt;/th&gt;
&lt;th&gt;Expected canonical value&lt;/th&gt;
&lt;th&gt;Gemini&lt;/th&gt;
&lt;th&gt;Claude&lt;/th&gt;
&lt;th&gt;Gemma&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;N01A: declared pt-BR&lt;/td&gt;
&lt;td&gt;&lt;code&gt;"4700.00"&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;ok&lt;/code&gt;, &lt;code&gt;"4700.00"&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;ambiguous&lt;/code&gt;, null&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;ok&lt;/code&gt;, &lt;code&gt;"4700"&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;N01B: declared en-US&lt;/td&gt;
&lt;td&gt;&lt;code&gt;"4.70"&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;ok&lt;/code&gt;, &lt;code&gt;"4.70"&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;ambiguous&lt;/code&gt;, null&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;ok&lt;/code&gt;, &lt;code&gt;"4.7"&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Claude's two responses were valid JSON, but they classified determined inputs as ambiguous. Within this test, that is &lt;strong&gt;unnecessary abstention&lt;/strong&gt;: the source already supplied the missing piece needed to choose.&lt;/p&gt;

&lt;p&gt;Gemma's answers have the correct numerical magnitudes. Their failure is &lt;strong&gt;canonical money representation&lt;/strong&gt;, not evidence that it selected the wrong locale or amount. Describing these as “two wrong amounts” would overstate what the outputs show.&lt;/p&gt;

&lt;p&gt;Both models lose two strict points and one complete pair. Their totals and pair counts therefore tie, yet the interventions one might investigate next are different. One would test whether unnecessary abstention can be reduced; the other would test canonical output validation. This experiment did not evaluate either intervention, and I did not repair answers before scoring.&lt;/p&gt;

&lt;h3&gt;
  
  
  What did not fail
&lt;/h3&gt;

&lt;p&gt;The genuine-ambiguity cases were correctly handled by all three models. So were the missing-data, invalid-date and identifier cases in this corpus. The experiment does &lt;strong&gt;not&lt;/strong&gt; show that these models routinely invent values or lose identifiers.&lt;/p&gt;

&lt;p&gt;All 84 responses also met the JSON-container contract. The observed gap was not extra Markdown, missing keys or non-JSON text. Separating the container format from the field's canonical value made that distinction visible.&lt;/p&gt;

&lt;h3&gt;
  
  
  Limits
&lt;/h3&gt;

&lt;p&gt;This is one run on &lt;strong&gt;28 constructed cases&lt;/strong&gt;, not a representative sample of all administrative work. Paired questions are related observations, and the status distribution is not balanced: 19 &lt;code&gt;ok&lt;/code&gt;, three &lt;code&gt;missing&lt;/code&gt;, two &lt;code&gt;ambiguous&lt;/code&gt; and four &lt;code&gt;invalid&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;A perfect score here does not establish general reliability; a two-point gap does not establish general model superiority or statistical significance. The prompts explicitly teach a data contract. They test following that contract, not discovering every convention from an unlabeled spreadsheet.&lt;/p&gt;

&lt;p&gt;These inputs follow explicit normalization rules, so a deterministic implementation is a plausible alternative. &lt;strong&gt;I did not measure that baseline; this study does not establish that an LLM is necessary for the task.&lt;/strong&gt; A deterministic grader is not a measured normalization baseline.&lt;/p&gt;

&lt;p&gt;No corpus or gold-label edits were made after the final outputs were seen. The strict score remains the predeclared score; the explanation of Gemma's numerical equivalence is a diagnostic reading of its two responses, not a replacement metric invented to change the ranking.&lt;/p&gt;

&lt;p&gt;Before making any deployment recommendation, I would still need independently authored paraphrases, more varied tests and repeated measurements across different conditions. The targeted follow-up below adds two repetitions on newly authored cases, but does not meet that broader standard.&lt;/p&gt;

&lt;h2&gt;
  
  
  Follow-up: Do different failures recur on new inputs? (October 8, 2026)
&lt;/h2&gt;

&lt;p&gt;The original study is unchanged. After seeing its results, I authored &lt;strong&gt;12 new fictional monetary cases&lt;/strong&gt; (four contrast pairs and four standalone controls) and publicly &lt;a href="https://github.com/lawliet8886/fechamento-br-benchmark/commit/6cf0c5824e9dac4871c9ea7827c9b0ab90c6c2e3" rel="noopener noreferrer"&gt;registered the corpus and protocol&lt;/a&gt; &lt;strong&gt;before collecting follow-up answers&lt;/strong&gt;. I also &lt;a href="https://github.com/lawliet8886/fechamento-br-benchmark/commit/c86393a3a6b497b027b58a59aa4f82a3dd271d22" rel="noopener noreferrer"&gt;froze the reviewed instrument and successful zero-model Kaggle preflight&lt;/a&gt; before testing. Three preselected models answered the 12 cases twice: six completed attempts, &lt;strong&gt;72 additional raw responses&lt;/strong&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Repetition 1&lt;/th&gt;
&lt;th&gt;Repetition 2&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.8 Flash&lt;/td&gt;
&lt;td&gt;12/12&lt;/td&gt;
&lt;td&gt;12/12&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 4.6&lt;/td&gt;
&lt;td&gt;7/12&lt;/td&gt;
&lt;td&gt;6/12&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 4 26B A4B&lt;/td&gt;
&lt;td&gt;11/12&lt;/td&gt;
&lt;td&gt;12/12&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;All 24 responses to the four controls passed. In a new contrast case with raw &lt;strong&gt;&lt;code&gt;3.600&lt;/code&gt;&lt;/strong&gt;, an explicitly declared decimal-point convention and expected &lt;code&gt;"3.60"&lt;/code&gt;, Claude returned &lt;code&gt;ambiguous&lt;/code&gt; while Gemma returned &lt;code&gt;"3.6"&lt;/code&gt; in the first repetition. Thus the original distinction — &lt;strong&gt;unnecessary abstention versus a correct magnitude with the wrong canonical string&lt;/strong&gt; — appeared again with new source text. In repetition two, Claude abstained again but Gemma passed that case; the formatting failure did not persist across both repetitions.&lt;/p&gt;

&lt;p&gt;The follow-up also uncovered &lt;strong&gt;three Claude JSON-container violations&lt;/strong&gt; (explanatory prose before an otherwise correct JSON answer) and &lt;strong&gt;one wrong monetary magnitude&lt;/strong&gt;. These are different from an unnecessary abstention or a missing trailing zero. The original 84 responses all satisfied the JSON-container requirement; that observation belongs only to the original batch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://lawliet8886.github.io/fechamento-br-benchmark/confirmation.html" rel="noopener noreferrer"&gt;Explore the follow-up, including every unchanged answer and diagnostic label&lt;/a&gt;&lt;/strong&gt; · &lt;a href="https://github.com/lawliet8886/fechamento-br-benchmark/blob/main/confirmation_v1/RESULTS.md" rel="noopener noreferrer"&gt;Read the methods, limits and verification&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is a &lt;strong&gt;targeted, post-hoc-designed follow-up&lt;/strong&gt;, not a blinded holdout or random sample. Values and wording changed together, and sampling/caching differences cannot be ruled out. It does not establish general model superiority, statistical significance or a causal explanation for the changes. The 72 new responses were verified separately against native Kaggle traces; they &lt;strong&gt;do not change the original 84-response study or the Kaggle public leaderboard&lt;/strong&gt;. Only included Kaggle model quota was used; there was no personal paid API spend.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Benchmark
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Explore every original answer:&lt;/strong&gt; &lt;a href="https://lawliet8886.github.io/fechamento-br-benchmark/" rel="noopener noreferrer"&gt;Interactive evidence explorer — 28 cases, 84 unchanged responses&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reproduce the scoring offline:&lt;/strong&gt; &lt;a href="https://github.com/lawliet8886/fechamento-br-benchmark" rel="noopener noreferrer"&gt;Source code, predeclared cases, original run records and deterministic verifier on GitHub&lt;/a&gt;. The verification script uses the Python standard library only and makes zero model calls. The interactive page contains strictly fictional records.&lt;/p&gt;

&lt;p&gt;These links refer to the &lt;strong&gt;frozen original research batch&lt;/strong&gt;, not to subsequent automated Kaggle public-leaderboard reruns.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.kaggle.com/benchmarks/lawliet8886/fechamento-br-same-score-different-failure" rel="noopener noreferrer"&gt;Fechamento BR — Public Kaggle Benchmark&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The Kaggle benchmark and its underlying task are publicly accessible. The 84-response initial comparison described above is preserved separately from subsequent public-leaderboard reruns; a later platform run is not silently merged into the original fixed study.&lt;/p&gt;

&lt;h3&gt;
  
  
  Evidence and AI assistance
&lt;/h3&gt;

&lt;p&gt;The reproducibility package contains the frozen cases and grader, protocol and source hashes, run manifests, original responses, native task traces and recomputed summaries. The scorer rejects records that do not match the frozen attempt, corpus and instrument manifest; incomplete runs receive no aggregate score.&lt;/p&gt;

&lt;p&gt;ChatGPT assisted with implementation, analysis and writing. Codex provided a separate AI review of the corpus and instrument, with access to pilot results. &lt;strong&gt;This was not a blind review or an independent human audit.&lt;/strong&gt; Matching saved records to native traces is a consistency check, not a second independent measurement. I remain responsible for the claims and materials.&lt;/p&gt;

&lt;p&gt;The implementation uses the official &lt;a href="https://github.com/Kaggle/kaggle-benchmarks" rel="noopener noreferrer"&gt;Kaggle Benchmarks SDK&lt;/a&gt; and its &lt;a href="https://github.com/Kaggle/kaggle-benchmarks/blob/ci/quick_start.md" rel="noopener noreferrer"&gt;quick start&lt;/a&gt;. Measurements used Kaggle's native model allowance, without personal paid API keys. Local fixed-response tests are kept distinct from model-performance evidence.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
