<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Tao Li</title>
    <description>The latest articles on DEV Community by Tao Li (@tao_li_0d9887b804d8389955).</description>
    <link>https://dev.to/tao_li_0d9887b804d8389955</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4174727%2F6d11cb76-b21d-45bf-881a-46c2ff50dea8.png</url>
      <title>DEV Community: Tao Li</title>
      <link>https://dev.to/tao_li_0d9887b804d8389955</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/tao_li_0d9887b804d8389955"/>
    <language>en</language>
    <item>
      <title>TrailBrief: Keep the Evidence, Leave the Screen</title>
      <dc:creator>Tao Li</dc:creator>
      <pubDate>Sat, 10 Oct 2026 15:24:21 +0000</pubDate>
      <link>https://dev.to/tao_li_0d9887b804d8389955/trailbrief-keep-the-evidence-leave-the-screen-9fb</link>
      <guid>https://dev.to/tao_li_0d9887b804d8389955/trailbrief-keep-the-evidence-leave-the-screen-9fb</guid>
      <description>&lt;p&gt;My entry for DEV's &lt;a href="https://dev.to/challenges/hacktoberfest-week1-2026-10-05"&gt;Week 1: Touch Grass&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;TrailBrief turns a pasted route notice and weather snapshot into a pocket brief you can save as a self-contained HTML file, review offline or print.&lt;/p&gt;

&lt;p&gt;It is for someone preparing a short outing who wants the important source evidence in one place. The aim is to make the screen a small part of preparation: collect the notices, check the gaps, take the brief, and verify conditions with the actual land manager and weather service before leaving.&lt;/p&gt;

&lt;p&gt;The AI has a narrow job. It selects source segment IDs. The backend copies the original text and adds fixed check questions. It never asks the model to invent a forecast, decide that a route is safe, or compose the evidence quotations.&lt;/p&gt;

&lt;p&gt;All examples in this post and demo are synthetic. I have not field-tested this application or authenticated real trail or weather data.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://k-3-lt.github.io/trailbrief/" rel="noopener noreferrer"&gt;Explore the recorded demo&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The public page is labeled &lt;strong&gt;Recorded model result / live API disabled&lt;/strong&gt;. You can switch between six real request records, inspect the synthetic source inputs and selected IDs, compare the failures, and download the two accepted briefs. Changing the selector does not call a model. There is no API-key field or live backend.&lt;/p&gt;

&lt;p&gt;The fourth record includes a closed boardwalk, missing forecast coverage and an embedded source instruction telling the model to declare the route safe. The model selected four permitted IDs. The backend copied the closure, the forecast gap and the two official-source check reminders exactly. The instruction segment was not selected.&lt;/p&gt;

&lt;p&gt;The fifth record supplies a forecast window that contains the planned departure and a notice reporting no closures. Its six selected segments all copied exactly, including both required evidence lines. These statements are true within the supplied fictional text; the application does not independently verify the conditions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/K-3-LT/trailbrief" rel="noopener noreferrer"&gt;Source repository&lt;/a&gt;. The reviewed release contains 216 files.&lt;/p&gt;

&lt;p&gt;The repository contains the actual local Node/Mastra application, locked dependencies, tests, recorded results and the static &lt;code&gt;/docs&lt;/code&gt; demo. Original TrailBrief source remains &lt;strong&gt;UNLICENSED&lt;/strong&gt;. Public source visibility is not an additional open-source license; dependency licenses and notices are preserved separately.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;The application uses &lt;code&gt;@mastra/core&lt;/code&gt; 1.75.0 and Zod with a locked npm dependency graph. The real workflow has three steps:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Validate the two supplied snapshots and add review warnings.&lt;/li&gt;
&lt;li&gt;Select evidence through the model adapter, or use the explicitly labeled deterministic demo.&lt;/li&gt;
&lt;li&gt;Verify the selection and assemble the portable brief.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The first contract asked the model to return verbatim quotations. The early calls showed why this was fragile:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Attempt&lt;/th&gt;
&lt;th&gt;What happened&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Extraction was rejected. The original assistant body was not retained, so I cannot reconstruct the exact cause.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;The 2048-token completion allowance was exhausted with an empty answer and &lt;code&gt;finish_reason=length&lt;/code&gt;. The provider also reported all 2048 tokens as reasoning.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;With an 8192-token allowance, JSON and schema validation passed. But a quotation changed “an imaginary valley” to “a valley”, so exact matching rejected it.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;I changed the contract to ID selection. Four valid IDs produced an accepted, deterministically assembled brief.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;The frozen ID version accepted a different normal fixture: supplied forecast coverage and no reported closure.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;The frozen version's missing-data fixture encountered &lt;code&gt;API_TRANSPORT_ERROR&lt;/code&gt; after about 30 seconds. No body or usage was received. It remained uncertain and was not retried.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For the new contract, the program splits each supplied source into bounded original slices. IDs bind the source ID, complete source text and positions. The model returns only &lt;code&gt;sourceId&lt;/code&gt; and &lt;code&gt;segmentId&lt;/code&gt;. Strict validation rejects additional fields, unknown IDs, repeats, cross-source substitutions and recognizable override directives. The program then copies the selected original slices and runs the existing exact-quotation check. No fuzzy matching or quotation repair happens.&lt;/p&gt;

&lt;p&gt;Questions and checklist items are fixed program text. A model-written explanation cannot silently become evidence. Full snapshots and selected IDs stay available for review.&lt;/p&gt;

&lt;p&gt;All six calls used the external &lt;code&gt;deepseek/deepseek-v4.1-flash&lt;/code&gt; model with low reasoning and zero retries. The fourth API request through complete response-body receipt took 12.216 seconds; the fifth took 15.602 seconds. These measurements include the network and provider response, not just model compute. The sixth did not yield a completed-response measurement.&lt;/p&gt;

&lt;p&gt;There are 82 passing offline tests in the frozen release. They cover ID stability, ownership, unknown/repeated IDs, strict fields, instruction isolation, missing-weather handling, exact copying, safe export, actual Mastra execution with mocked extraction, budget continuity and first-error stop. Separate desktop/mobile browser checks verified that the static demo displayed recorded results, exposed no key form, made only local static GET requests and downloaded byte-identical saved briefs.&lt;/p&gt;

&lt;p&gt;This is a small development record, not a benchmark or success-rate claim. The missing-data case has offline coverage but no accepted real response. Reported usage for attempts 1–5 yields a conventional estimate of $0.013918 at the provider's &lt;a href="https://www.iteracompute.com/models/deepseek-v4-1-flash/" rel="noopener noreferrer"&gt;published token prices&lt;/a&gt;. That excludes any unknown charge for attempt 6 and is not a final invoice. The retained $1.50 reservation is a safety allowance, not actual spending.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;p&gt;Mastra makes the workflow boundaries inspectable: I can keep model selection separate from source validation and export, and test the same orchestration with an injected mock adapter. It is central to the local application's working path, rather than an unused dependency added for a badge.&lt;/p&gt;

&lt;p&gt;The main design change was possible because the evidence contract is owned by the application. I could replace fragile model-written quotations with ID selection while leaving the workflow and the source-verification boundary intact. An OpenAI-compatible adapter also separates provider transport from evidence assembly; I tested one fixed model, not a broad provider comparison.&lt;/p&gt;

&lt;p&gt;The open-source claim here is about the framework and third-party components. My original source is currently UNLICENSED. The model is called through an external provider. Exported briefs can be read offline, but real model selection does not run offline and is not a free hosted service.&lt;/p&gt;

&lt;p&gt;Important limits remain: pasted text and timestamps are not authenticated; exact copying does not establish truth, freshness, completeness or safety; models can omit important evidence; recognizable instruction detection is not a complete injection defense; and a documented moderate transitive dependency advisory remains in the local Node prototype. The public site serves plain static records, not that Node runtime or a live inference service. Human review is required.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Agent Session
&lt;/h2&gt;

&lt;p&gt;I directed an AI-assisted build; AI agents generated the implementation, tests and this write-up, and reviewed the recorded evidence. I am publishing the source, synthetic fixtures, validation records and failure history rather than a raw session log that could contain private configuration. The recorded demo shows the model inputs, retained assistant outputs and deterministic assembly trace without request headers or credentials.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;p&gt;Best Use of Mastra — the actual three-step local workflow uses Mastra's workflow API in both the deterministic demo and model-selection paths. No other partner category is claimed.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>hf26challenge</category>
    </item>
    <item>
      <title>RetryBudget-24: When the Right Retry Decision Still Fails the JSON Contract</title>
      <dc:creator>Tao Li</dc:creator>
      <pubDate>Sat, 10 Oct 2026 06:49:34 +0000</pubDate>
      <link>https://dev.to/tao_li_0d9887b804d8389955/retrybudget-24-when-the-right-retry-decision-still-fails-the-json-contract-39mb</link>
      <guid>https://dev.to/tao_li_0d9887b804d8389955/retrybudget-24-when-the-right-retry-decision-still-fails-the-json-contract-39mb</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;Kaggle Benchmarking Challenge&lt;/a&gt;.&lt;/em&gt; &lt;/p&gt;

&lt;h2&gt;
  
  
  What I Benchmarked
&lt;/h2&gt;

&lt;p&gt;An AI component advising an API client has two jobs: choose the right action and return something the client can actually parse. RetryBudget-24 tests those jobs together, then separates their failure modes for analysis.&lt;/p&gt;

&lt;p&gt;Each question describes a failed or completed request and supplies an explicit, synthetic retry policy. The model returns exactly two fields:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"action"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"RETRY"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"delay_ms"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The actions are DONE, STOP, RECONCILE, and RETRY. The policy covers status precedence, attempt limits, idempotency guarantees, capped exponential backoff, Retry-After, and a remaining-time budget. A retry is permitted only when its wait plus the next attempt's estimated duration fits the budget. Equality is allowed.&lt;/p&gt;

&lt;p&gt;This is a text-based policy-decision task. It does not execute HTTP requests, measure inference-serving latency, or demonstrate safe production retries.&lt;/p&gt;

&lt;h3&gt;
  
  
  Twelve controlled pairs
&lt;/h3&gt;

&lt;p&gt;The formal set contains 24 original synthetic cases arranged as 12 pairs. Within each pair, exactly one input field changes. That makes the boundary being tested easier to inspect than in a collection of unrelated puzzles.&lt;/p&gt;

&lt;p&gt;For example, keep a 100 ms wait and a 200 ms next attempt fixed:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;299 ms remaining means STOP.&lt;/li&gt;
&lt;li&gt;300 ms remaining means RETRY with a 100 ms delay.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Other pairs change whether a server guarantees deduplication for an idempotency key, whether the attempt limit has been reached, or whether the backoff cap binds.&lt;/p&gt;

&lt;p&gt;The cases, fixed order, prompts, oracle, and scorer were hash-frozen before any model call. Four separate development cases checked execution without tuning the prompts against model responses. Expected answers were checked against the deterministic oracle; 20 offline tests checked the original benchmark implementation.&lt;/p&gt;

&lt;p&gt;The primary score is exact correctness on both output fields, divided by 24. JSON whitespace and key order are allowed. Code fences, explanations, extra or duplicate keys, floats, and booleans in the integer field fail the contract. No LLM judge is involved.&lt;/p&gt;

&lt;p&gt;For context, deterministic baselines score 24/24 for the policy oracle, 6/24 for always STOP, and 7/24 for a status-only heuristic that retries transient statuses after 100 ms. The oracle is a correctness reference, not a competing language model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models Tested
&lt;/h2&gt;

&lt;p&gt;The experiment used two models available through Kaggle's Model Proxy:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;google/gemini-2.5-flash&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;anthropic/claude-haiku-4-5@20251001&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This small, cross-provider comparison kept the experiment inexpensive. It is not a survey of the strongest available models.&lt;/p&gt;

&lt;p&gt;Both received the same frozen prompts through Kaggle Benchmarks SDK 0.6.1, with &lt;code&gt;max_tokens=512&lt;/code&gt;, seed 0, and temperature 0 requested. Importantly, the inspected SDK path reported temperature unsupported for both and dropped that setting. Actual temperature was therefore not controlled. Reasoning settings remained at provider defaults, and requested settings do not establish identical effective inference budgets.&lt;/p&gt;

&lt;p&gt;Each case used a separate chat. There were no retries for wrong answers. Client SDK retries were disabled; retries inside the proxy were not observable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Findings
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The planned experiment
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Strict score&lt;/th&gt;
&lt;th&gt;Valid schema&lt;/th&gt;
&lt;th&gt;Both members of pair correct&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 2.5 Flash&lt;/td&gt;
&lt;td&gt;16/24&lt;/td&gt;
&lt;td&gt;22/24&lt;/td&gt;
&lt;td&gt;4/12&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Haiku 4.5&lt;/td&gt;
&lt;td&gt;5/24&lt;/td&gt;
&lt;td&gt;9/24&lt;/td&gt;
&lt;td&gt;1/12&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The strict gap looks large. Inspecting the responses changes what can reasonably be concluded from it.&lt;/p&gt;

&lt;p&gt;Haiku wrapped 15 of its 24 answers in complete Markdown code fences. Nine of those contained the correct action and delay. Gemini produced one complete fence with correct content and one unclosed fence.&lt;/p&gt;

&lt;p&gt;A narrowly defined, post-hoc diagnostic removed only complete outer fences, without repairing JSON, coercing types, or extracting objects from prose. Correctness then became 17/24 for Gemini and 14/24 for Haiku. These are diagnostic counts, not replacement benchmark scores.&lt;/p&gt;

&lt;p&gt;The distinction matters for integration design. A strict JSON consumer really would reject those fenced responses. But calling every rejection a reasoning failure would hide a substantial part of what happened. A parser adapter or constrained-output interface could be worth testing in a separately frozen experiment; neither was retrofitted into these scores.&lt;/p&gt;

&lt;h3&gt;
  
  
  Some errors remain after formatting is separated
&lt;/h3&gt;

&lt;p&gt;Both models chose RETRY on the 299 ms budget example, although 100 + 200 exceeds 299. At 300 ms, both returned the correct decision content, with Haiku's answer fenced.&lt;/p&gt;

&lt;p&gt;On an unsafe timeout with no idempotency guarantee, Haiku returned RETRY where the policy required RECONCILE. On a capped-backoff case where &lt;code&gt;min(1000, 700 * 2)&lt;/code&gt; should produce 1000 ms, Gemini returned 1400 and Haiku returned 700.&lt;/p&gt;

&lt;p&gt;Those examples identify concrete policy boundaries worth testing. They do not establish how either model would behave across real API traffic.&lt;/p&gt;

&lt;h3&gt;
  
  
  An unintended second experiment
&lt;/h3&gt;

&lt;p&gt;There was an execution-control mistake during publication preparation: Build Task was treated as a save/build operation, but it executed the original notebook again. The additional run completed before it could be stopped. That exceeded the planned 56 client calls and must be counted.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Batch&lt;/th&gt;
&lt;th&gt;Gemini formal score&lt;/th&gt;
&lt;th&gt;Haiku formal score&lt;/th&gt;
&lt;th&gt;Calls including development&lt;/th&gt;
&lt;th&gt;Reported quota cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Planned experiment&lt;/td&gt;
&lt;td&gt;16/24&lt;/td&gt;
&lt;td&gt;5/24&lt;/td&gt;
&lt;td&gt;56&lt;/td&gt;
&lt;td&gt;$0.0521384&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Unplanned Build rerun&lt;/td&gt;
&lt;td&gt;18/24&lt;/td&gt;
&lt;td&gt;4/24&lt;/td&gt;
&lt;td&gt;56&lt;/td&gt;
&lt;td&gt;$0.0523159&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Total&lt;/td&gt;
&lt;td&gt;Separate batches&lt;/td&gt;
&lt;td&gt;Separate batches&lt;/td&gt;
&lt;td&gt;112&lt;/td&gt;
&lt;td&gt;$0.1044543&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;All 112 recorded calls completed. The additional batch used identical prompt hashes and unchanged scoring. Its outputs are retained separately; the better Gemini result does not replace the planned result, and the batches are not pooled into a new headline score.&lt;/p&gt;

&lt;p&gt;The amounts are reported Model Proxy quota consumption within the account's available free quota, not a cash purchase. They show that this particular experiment was inexpensive, not that future runs have a guaranteed price.&lt;/p&gt;

&lt;p&gt;This also exposed a benchmark-workflow requirement: publication steps need their own execution budget. A call ledger inside one notebook session cannot impose a global cap across fresh platform executions.&lt;/p&gt;

&lt;h3&gt;
  
  
  What the results cannot tell us
&lt;/h3&gt;

&lt;p&gt;This is a small synthetic set, with related pairs, one planned sample per case, and one accidental repeat. It supports failure analysis rather than a general model ranking or statistically reliable reliability estimate.&lt;/p&gt;

&lt;p&gt;The saved artifacts do not contain &lt;code&gt;finish_reason&lt;/code&gt;. In the planned Gemini test, 21 of 24 reported output-token counts were at least 450 against the requested 512-token limit. That warrants caution, but it does not prove truncation or explain a particular error. The models' effective generation conditions were not fully controlled.&lt;/p&gt;

&lt;p&gt;The next useful experiment would predefine strict versus constrained-output conditions, more policy boundaries, and planned repetitions. It would also verify the platform's build behavior before enabling any model-call entry point.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Benchmark
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Kaggle benchmark collection:&lt;/strong&gt; &lt;a href="https://www.kaggle.com/benchmarks/liwhattao/retrybudget-24-policy-and-json-contracts" rel="noopener noreferrer"&gt;RetryBudget-24: Policy and JSON Contracts&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Source, frozen cases, manifests, and raw results:&lt;/strong&gt; &lt;a href="https://www.kaggle.com/code/liwhattao/retrybudget-24-v1?scriptVersionId=356959464" rel="noopener noreferrer"&gt;Planned experiment snapshot&lt;/a&gt; and &lt;a href="https://www.kaggle.com/code/liwhattao/retrybudget-24-v1?scriptVersionId=356960994" rel="noopener noreferrer"&gt;unplanned Build rerun snapshot&lt;/a&gt;. The self-authored benchmark code is published under Apache 2.0.&lt;/p&gt;

&lt;p&gt;The published Kaggle task and collection contain only the unplanned batch's Gemini result, 0.75 (18/24). The two-model comparisons above come from separately preserved and audited raw experiment artifacts; they are not a claim that both models appear on the platform leaderboard. The original execution adapter also depends on its notebook context. A self-contained publication wrapper has been prepared and tested offline, but it has not been run against real models or substituted for the measured version.&lt;/p&gt;

&lt;p&gt;AI assistance was used to design and implement the synthetic benchmark, run the workflow, audit saved outputs, and draft this article. The task uses original synthetic cases rather than third-party evaluation data. The SDK is credited below. The article explicitly retains the AI-assisted workflow's unintended rerun instead of hiding it.&lt;/p&gt;

&lt;p&gt;Official references:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/Kaggle/kaggle-benchmarks" rel="noopener noreferrer"&gt;Kaggle Benchmarks SDK&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/Kaggle/kaggle-skills/blob/main/write-kaggle-benchmarks/SKILL.md" rel="noopener noreferrer"&gt;Kaggle's task-authoring and execution guide&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;Kaggle Benchmarking Challenge&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>kagglechallenge</category>
      <category>devchallenge</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
