<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: 十方空烬 · Ember ☄️</title>
    <description>The latest articles on DEV Community by 十方空烬 · Ember ☄️ (@widechaos).</description>
    <link>https://dev.to/widechaos</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4155333%2F0a208f72-8cde-493a-bbb0-9fc745edea42.png</url>
      <title>DEV Community: 十方空烬 · Ember ☄️</title>
      <link>https://dev.to/widechaos</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/widechaos"/>
    <language>en</language>
    <item>
      <title>Shelf Friend helps my friend compare supermarket prices and product care</title>
      <dc:creator>十方空烬 · Ember ☄️</dc:creator>
      <pubDate>Fri, 02 Oct 2026 06:39:11 +0000</pubDate>
      <link>https://dev.to/widechaos/shelf-friend-helps-my-friend-compare-supermarket-prices-and-product-care-2g9o</link>
      <guid>https://dev.to/widechaos/shelf-friend-helps-my-friend-compare-supermarket-prices-and-product-care-2g9o</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-weekend-2026-10-01"&gt;Hacktoberfest Weekend Challenge: Build for a Friend&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;A friend needs help while shopping at a supermarket: compare prices, understand how a kitchen or household item should be used, and weigh its advantages against the maintenance it brings home.&lt;/p&gt;

&lt;p&gt;I built &lt;strong&gt;Shelf Friend&lt;/strong&gt;, a small Chinese-language web app with two connected decisions: &lt;strong&gt;what costs less per usable unit, and what fits the intended use?&lt;/strong&gt; My friend stays anonymous. I have not yet collected their feedback, so the results below are my implementation checks, not a claimed user study.&lt;/p&gt;

&lt;p&gt;The calculator takes prices from the shelf, net quantity and pack count. The AI retrieves a relevant use guide or buying checklist from a small, visible corpus. A cheaper pan may still be a poor purchase if its weight, stove compatibility or care routine does not fit the person buying it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://widechaos.github.io/shelf-friend/" rel="noopener noreferrer"&gt;Open the live demo&lt;/a&gt;&lt;/strong&gt; — no account or API key needed.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1dd5c9d8n2m9kz36a351.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1dd5c9d8n2m9kz36a351.png" alt="Shelf Friend showing a unit-price comparison beside a locally retrieved cast-iron care guide" width="800" height="569"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Try this short walkthrough:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Click &lt;strong&gt;试试演示价格&lt;/strong&gt; (“Try demo prices”). The fictional rice offers are 2 kg for 2.50 OMR and 500 g for 0.75 OMR. They become 1.25 and 1.50 OMR/kg; the first is 16.7% cheaper per kg. These are demonstration prices, not retailer quotes.&lt;/li&gt;
&lt;li&gt;Click &lt;strong&gt;铸铁锅保养&lt;/strong&gt; (“Cast-iron care”). The browser downloads the model on first use, then retrieves a guide scoped to Lodge's seasoned cast iron, with a source link and a maintenance tradeoff.&lt;/li&gt;
&lt;li&gt;Try &lt;strong&gt;保鲜盒怎么用&lt;/strong&gt; (“Food-container use”) or ask in English, “How should I clean and oil a cast iron skillet?” The interface and guides are Chinese, but query matching is multilingual.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/widechaos/shelf-friend" rel="noopener noreferrer"&gt;Public repository&lt;/a&gt;&lt;/strong&gt; · MIT application code and original checklists.&lt;/p&gt;

&lt;p&gt;I started this project and its repository on October 2, 2026, inside the weekend challenge window. It is separate from my weekly challenge work and uses no employer resources.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;The AI core is &lt;a href="https://github.com/huggingface/transformers.js" rel="noopener noreferrer"&gt;Transformers.js&lt;/a&gt; running the quantized ONNX conversion of the open-weight &lt;a href="https://huggingface.co/Xenova/paraphrase-multilingual-MiniLM-L12-v2" rel="noopener noreferrer"&gt;multilingual MiniLM model&lt;/a&gt;. A web worker embeds the question and the guide descriptions, then ranks them by cosine similarity. This is model inference, not a keyword lookup. The app displays only the highest-ranking guide above its threshold, or says the corpus has no sufficiently relevant content. Similarity is labelled as similarity, not a probability or factual-accuracy score.&lt;/p&gt;

&lt;p&gt;The eight-guide corpus covers cookware, pantry packaging, cleaners, food containers, small appliances, tissue multipacks and batteries. Most entries are my original buying checklists. The cast-iron entry is a concise, linked summary of &lt;a href="https://www.lodgecastiron.com/pages/discover-cleaning-and-care-cast-iron" rel="noopener noreferrer"&gt;Lodge's care instructions&lt;/a&gt;, with its scope explicit. Unknown product specifications are left unknown.&lt;/p&gt;

&lt;p&gt;Price arithmetic is deterministic JavaScript. It normalizes grams/kilograms and millilitres/litres, accounts for multipacks, and rejects comparisons between mass, volume and count. The model never calculates the winning price.&lt;/p&gt;

&lt;p&gt;In Edge, I checked Chinese cast-iron, bulk-rice and food-container queries, an English cast-iron query, and an unrelated football question that returned no guide. I also checked the public deployment, a small-phone layout and landscape layout. Automated tests cover unit conversion, incompatible quantities, invalid inputs, zero prices, ties and similarity arithmetic. These checks are limited; I have not benchmarked retrieval accuracy or tested every phone.&lt;/p&gt;

&lt;p&gt;The first model download is about 118 MB. It took 13.3 seconds in my public-demo test; one later query completed in about 0.1 seconds. That is one browser observation, not a speed promise. I would load it before entering a supermarket with weak reception. Caches may help later use, but I do not promise offline availability on every device.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;p&gt;Open weights make the semantic search usable inside the shopper's own browser. I can inspect and change the corpus, pin the model revision, and keep shopping questions away from a remote inference service. I do not need a subscription, an API secret on the page, or a server that stores my friend's questions. Initial runtime and model downloads still contact public CDNs and produce ordinary request metadata.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://huggingface.co/sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2" rel="noopener noreferrer"&gt;base model&lt;/a&gt; and Transformers.js use Apache-2.0 licenses. Their open distribution made this small, reproducible deployment possible. The application and checklists can be inspected and improved independently of the model.&lt;/p&gt;

&lt;p&gt;This first version supports &lt;strong&gt;manual shelf-price comparison and a small guide library&lt;/strong&gt;. It does not scrape live retailer prices, recognize product photos or barcodes, certify materials, or invent a product manual. My next useful step is to let my friend try it on a real shopping trip, then use their feedback to decide which products and instructions to add.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
    </item>
    <item>
      <title>VoiceStack Atlas: version-aware speech deployment with Astro and Sanity</title>
      <dc:creator>十方空烬 · Ember ☄️</dc:creator>
      <pubDate>Thu, 01 Oct 2026 18:46:35 +0000</pubDate>
      <link>https://dev.to/widechaos/voicestack-atlas-version-aware-speech-deployment-with-astro-and-sanity-2la2</link>
      <guid>https://dev.to/widechaos/voicestack-atlas-version-aware-speech-deployment-with-astro-and-sanity-2la2</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/sanity-2026-09-16"&gt;Sanity Challenge, Path Two: Vibe-Code Something Strange&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;I built VoiceStack Atlas, a deployment planner that keeps old speech-tool documentation in its own era. It asks for a runtime profile, reads structured compatibility rules from Sanity, and returns decisions tied to original, commit-pinned sources.&lt;/p&gt;

&lt;p&gt;The interesting case is CUDA 12 with cuDNN 8. A generic “GPU requirements” search can surface different documentation eras. The current faster-whisper README documents a specific CTranslate2 workaround; the v1.0.0 README describes a different CUDA stack. I model the condition and documentation status separately, so a historical fact cannot become today's installation instruction.&lt;/p&gt;

&lt;p&gt;The same distinction applies to audio: built-in file decoding does not mean an 8 kHz NumPy array will be resampled. The planner separates those paths and reminds the caller that transcription starts when the segment generator is iterated.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://widechaos.github.io/voicestack-atlas/" rel="noopener noreferrer"&gt;Open the project&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Try CUDA 12 / cuDNN 8 / Raw NumPy array, then change CUDA to 13. The second profile should ask for more evidence rather than inventing an installation command. Switch to CPU and file input to see GPU requirements disappear.&lt;/p&gt;

&lt;p&gt;The deployed Astro app reads the public Sanity dataset without a login or browser token. I verified the current CUDA 12 / cuDNN 8 / array profile, the unsupported CUDA 13 profile, and CPU / file input against real Content Lake documents.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://widechaos.github.io/voicestack-atlas/screenshots/live-planner.png" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flte7qnkio6gsndrfc6el.png" alt="Live planner resolving CUDA 12, cuDNN 8, and raw array constraints" width="800" height="510"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://widechaos.github.io/voicestack-atlas/screenshots/live-planner.png" rel="noopener noreferrer"&gt;Open the full-size screenshot&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/widechaos/voicestack-atlas" rel="noopener noreferrer"&gt;Source code and reproducible checks&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Original compatibility annotations, constraints, interface, and evaluation are my work. Public faster-whisper documentation and source code are credited and preserved under their upstream MIT license. No employer code, recordings, datasets, or credentials are used.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Build Process
&lt;/h2&gt;

&lt;p&gt;I started by asking for conditions, conclusions, actions, and provenance to be separate fields. A polished answer is not useful if it mixes a release-era requirement with a current requirement.&lt;/p&gt;

&lt;p&gt;I used the Codex desktop IDE to build a narrow planner and source validation first, then an Astro interface. The build instructions evolved from a real Context-connected agent to a directly verifiable structured-content app. This is a summary of the iteration, not a verbatim prompt transcript: separate conditions from conclusions; preserve literal evidence from pinned files; refuse unsupported profiles; and verify Sanity reads before publishing. An early assumption about the historical README failed its literal evidence check. I corrected the annotation against the actual pinned file. That failure shaped the design: source checks are part of the workflow, not a decorative citation after the answer.&lt;/p&gt;

&lt;p&gt;I initially attempted Path One with a Sanity Context Knowledge Base. Public website ingestion succeeded, but the completed build exposed no readable entries, and endpoint creation returned a not-found page. I did not present the offline planner as a working MCP agent. I prepared a direct Content Lake version in Astro instead, because its structured decision model is valuable without a live model loop. The Context harness remains experimental code, not a claimed completed agent.&lt;/p&gt;

&lt;p&gt;The second correction was operational. I created source and compatibility-rule documents in a personal public Content Lake dataset, added a native source reference, and configured credential-free browser reads. An import token stays in a local ignored environment file. The frontend has no token and fetches the live rules on each resolution. I checked three profiles against the live dataset before switching the public deployment from the earlier offline prototype.&lt;/p&gt;

&lt;p&gt;I kept the decision engine deterministic: a runtime profile must satisfy every applicable condition, and historical records only appear in a separate context section. This makes the result traceable and gives unsupported combinations a useful failure state. I did not use App SDK or Workflows; the scope is an Astro frontend over structured Sanity content.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Used Sanity
&lt;/h2&gt;

&lt;p&gt;The Content Lake model has two document types. &lt;code&gt;source&lt;/code&gt; stores repository, ref, commit, original URL, and license. &lt;code&gt;compatibilityRule&lt;/code&gt; stores &lt;code&gt;when&lt;/code&gt;, &lt;code&gt;status&lt;/code&gt;, &lt;code&gt;claim&lt;/code&gt;, &lt;code&gt;action&lt;/code&gt;, and a native &lt;code&gt;sourceRef&lt;/code&gt; reference to the pinned source document. Conditions combine device, CUDA major, cuDNN major, and audio input format. The historical/current status is independent of those conditions.&lt;/p&gt;

&lt;p&gt;The Astro frontend reads those documents through public GROQ queries and applies all conditions together. It displays an error when Sanity is unavailable; no cached compatibility answer is silently substituted. Tokens used for import stay on the server side and are never shipped to the browser.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sanity Project Details
&lt;/h2&gt;

&lt;p&gt;Project ID: &lt;strong&gt;jud8maoc&lt;/strong&gt;. Dataset: &lt;strong&gt;production&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://widechaos.github.io/voicestack-atlas/sources/" rel="noopener noreferrer"&gt;Public source corpus&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://jud8maoc.api.sanity.io/v2026-09-01/data/query/production?query=%2A%5B_type%20%3D%3D%20%22compatibilityRule%22%20%7C%7C%20_type%20%3D%3D%20%22source%22%5D" rel="noopener noreferrer"&gt;Inspect the public dataset&lt;/a&gt;. It contains four source documents and eight original compatibility rules. No credentials are required.&lt;/p&gt;

&lt;p&gt;For example, a rule combines &lt;code&gt;device: cuda&lt;/code&gt;, &lt;code&gt;cuda: 12&lt;/code&gt;, and &lt;code&gt;cudnn: 8&lt;/code&gt;, with &lt;code&gt;status: current&lt;/code&gt;, an evidence literal, an action, and &lt;code&gt;sourceRef&lt;/code&gt; pointing to a commit-pinned document. A source reference is shared provenance; the historical/current field controls whether the rule can drive a current decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  Validation and Limits
&lt;/h2&gt;

&lt;p&gt;Local checks cover current versus historical dependency selection, missing versions, unsupported combinations, CPU separation, file versus array handling, and missing or unknown citations. Each curated evidence literal was checked against its pinned upstream file.&lt;/p&gt;

&lt;p&gt;I have not run GPU inference or claimed speed, memory, or speech accuracy results. This tool is a deployment planning aid, not proof that a dependency stack works on every machine. The Context prototype has not passed a real retrieval run, and is not the finished submission.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>sanitychallenge</category>
      <category>sanity</category>
      <category>ai</category>
    </item>
    <item>
      <title>Valid JSON is not enough: testing bilingual patch contracts on Kaggle</title>
      <dc:creator>十方空烬 · Ember ☄️</dc:creator>
      <pubDate>Thu, 01 Oct 2026 18:25:07 +0000</pubDate>
      <link>https://dev.to/widechaos/valid-json-is-not-enough-testing-bilingual-patch-contracts-on-kaggle-3690</link>
      <guid>https://dev.to/widechaos/valid-json-is-not-enough-testing-bilingual-patch-contracts-on-kaggle-3690</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;Kaggle Benchmarking Challenge&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Benchmarked
&lt;/h2&gt;

&lt;p&gt;A JSON response can parse successfully and still change the wrong state. A response can also contain the right state but arrive wrapped in Markdown that breaks a strict consumer. I built &lt;strong&gt;Bilingual Patch Contracts&lt;/strong&gt; to keep those two failure modes visible.&lt;/p&gt;

&lt;p&gt;The task has twelve handcrafted state-update scenarios, each with an English, Chinese, and code-switched instruction body: &lt;strong&gt;36 prompts&lt;/strong&gt;. Each triplet shares an initial state and expected answer. The contract prefix stays in English, and the output keys stay canonical English. This tests changing the instruction body's language under a shared contract; it is not a fully Chinese interaction benchmark.&lt;/p&gt;

&lt;p&gt;Every case begins with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"route"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"email"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"delay_minutes"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"active"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"tags"&lt;/span&gt;&lt;span class="p"&gt;:[&lt;/span&gt;&lt;span class="s2"&gt;"alpha"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s2"&gt;"beta"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="nl"&gt;"note"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"seed"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The scenarios cover later corrections, negation, null versus empty values, ordered and case-sensitive tags, hours-to-minutes conversion, sequential conditions, instruction-like literal data, and exact copying of Unicode, backslashes, quotes, and a newline.&lt;/p&gt;

&lt;p&gt;A pass requires the entire response to be one JSON object with exactly five keys, valid types, and every expected value. I do not strip Markdown, repair outputs, or ask another model to judge. Whitespace, object key order, and equivalent Unicode escapes are accepted. Duplicate keys, extra fields, nonfinite values, and booleans or floats in integer fields fail. Array order matters.&lt;/p&gt;

&lt;p&gt;I use ordinary text generation, with temperature 0 and seed 0 requested through the SDK, and a fresh isolated conversation for each case. There is no constrained JSON decoding, schema enforcement, or tool use. Provider behavior may still vary between runs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models Tested
&lt;/h2&gt;

&lt;p&gt;I ran the complete suite on Kaggle on &lt;strong&gt;October 1, 2026&lt;/strong&gt;, using task &lt;strong&gt;version 2&lt;/strong&gt;. I selected three lightweight models from different providers available in the platform catalog:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Exact Kaggle model identifier&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.7 Flash&lt;/td&gt;
&lt;td&gt;&lt;code&gt;gemini-3.7-flash&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.4 nano&lt;/td&gt;
&lt;td&gt;&lt;code&gt;gpt-5.4-nano-2026-03-17&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Haiku 4.5&lt;/td&gt;
&lt;td&gt;&lt;code&gt;claude-haiku-4-5-20251001&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I also attempted &lt;code&gt;qwen3-next-80b-a3b-instruct&lt;/code&gt;. Both the pilot and version 2 attempt stopped with HTTP 429 and a provider message about heavy load. It has &lt;strong&gt;no complete score&lt;/strong&gt;, so it is excluded from the leaderboard rather than assigned zero.&lt;/p&gt;

&lt;p&gt;Version 2 fixes task registration so Kaggle selects the whole-suite aggregate instead of a helper function. The prompts, fixtures, and scorer were unchanged. All results below come from complete version 2 runs. I downloaded the raw responses, checked all 36 unique case IDs against the frozen prompts and answers, and independently recalculated the saved scores.&lt;/p&gt;

&lt;h2&gt;
  
  
  Findings
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Strict exact match&lt;/th&gt;
&lt;th&gt;Valid JSON&lt;/th&gt;
&lt;th&gt;Valid schema&lt;/th&gt;
&lt;th&gt;English&lt;/th&gt;
&lt;th&gt;Chinese&lt;/th&gt;
&lt;th&gt;Mixed&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.7 Flash&lt;/td&gt;
&lt;td&gt;36/36 (100%)&lt;/td&gt;
&lt;td&gt;36/36&lt;/td&gt;
&lt;td&gt;36/36&lt;/td&gt;
&lt;td&gt;12/12&lt;/td&gt;
&lt;td&gt;12/12&lt;/td&gt;
&lt;td&gt;12/12&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.4 nano&lt;/td&gt;
&lt;td&gt;24/36 (66.7%)&lt;/td&gt;
&lt;td&gt;36/36&lt;/td&gt;
&lt;td&gt;36/36&lt;/td&gt;
&lt;td&gt;7/12&lt;/td&gt;
&lt;td&gt;8/12&lt;/td&gt;
&lt;td&gt;9/12&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Haiku 4.5&lt;/td&gt;
&lt;td&gt;0/36 (0%)&lt;/td&gt;
&lt;td&gt;0/36&lt;/td&gt;
&lt;td&gt;0/36&lt;/td&gt;
&lt;td&gt;0/12&lt;/td&gt;
&lt;td&gt;0/12&lt;/td&gt;
&lt;td&gt;0/12&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Valid JSON does not prove a correct update
&lt;/h3&gt;

&lt;p&gt;GPT-5.4 nano produced a valid object with valid field types on every case, but twelve answers contained wrong values.&lt;/p&gt;

&lt;p&gt;For &lt;code&gt;case-sensitive-en&lt;/code&gt;, the instruction body is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Tag strings are case-sensitive. Add Beta. Remove beta. Add ALPHA. Add alpha.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Starting with &lt;code&gt;["alpha","beta"]&lt;/code&gt;, the expected tags are &lt;code&gt;["alpha","Beta","ALPHA"]&lt;/code&gt;. The raw model response was:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"route"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"email"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"delay_minutes"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"active"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"tags"&lt;/span&gt;&lt;span class="p"&gt;:[&lt;/span&gt;&lt;span class="s2"&gt;"alpha"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s2"&gt;"beta"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s2"&gt;"Beta"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s2"&gt;"ALPHA"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="nl"&gt;"note"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"seed"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The lowercase &lt;code&gt;beta&lt;/code&gt; should have been removed. A JSON parser or type validator alone would accept this answer.&lt;/p&gt;

&lt;h3&gt;
  
  
  A presentation choice can break the contract
&lt;/h3&gt;

&lt;p&gt;Claude Haiku 4.5 wrapped every answer in a Markdown code fence despite the explicit instruction “No Markdown or explanations.” For &lt;code&gt;last-write-en&lt;/code&gt;, its full response was:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;```json
{"route": "sms", "delay_minutes": 0, "active": true, "tags": ["alpha", "beta"], "note": "seed"}
```
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The values in this example are correct, but the full response is not a JSON document. Its zero is a strict interface score, not evidence that it understood none of the requests.&lt;/p&gt;

&lt;p&gt;As a separate diagnostic, removing only the complete outer code fences would make &lt;strong&gt;33/36&lt;/strong&gt; answers pass the same value checks. That counterfactual is not the benchmark score and does not repair any leaderboard output. It shows why formatting compliance and state correctness deserve separate reporting.&lt;/p&gt;

&lt;h3&gt;
  
  
  Language totals need paired inspection
&lt;/h3&gt;

&lt;p&gt;Nano's mixed total exceeded its English total by two cases. Within the twelve matched scenarios, seven passed in both English and mixed, three failed in both, and two passed only in mixed. For English versus Chinese, six passed in both, three failed in both, one passed only in English, and two passed only in Chinese.&lt;/p&gt;

&lt;p&gt;These observations locate cases to inspect; they do not establish that the model is generally stronger in Chinese or code-switching. The bodies were hand-authored, and phrasing and token lengths are not perfectly controlled.&lt;/p&gt;

&lt;h3&gt;
  
  
  A ceiling is a prompt to extend the suite
&lt;/h3&gt;

&lt;p&gt;Gemini passed every case in this run. This suite therefore cannot distinguish its reliability beyond these examples. I would next add longer instruction chains, more literal-copy edge cases, and repeated runs, while keeping new cases separate from this frozen result.&lt;/p&gt;

&lt;p&gt;This is a small diagnostic benchmark, not a general model ranking. The three language variants are paired observations: there are twelve independent semantic scenarios, not thirty-six independent problems. A single run does not establish production reliability, and shared English contract instructions limit the multilingual claim. I did not benchmark latency, cost, or tool calling.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Benchmark
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.kaggle.com/benchmarks/xover2022/bilingual-patch-contracts" rel="noopener noreferrer"&gt;Bilingual Patch Contracts on Kaggle&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The collection uses one numeric task whose score is strict exact matches divided by 36. Its overall score is the average of task scores; with one task, it equals that task's score.&lt;/p&gt;

&lt;p&gt;Inspect the &lt;a href="https://www.kaggle.com/benchmarks/tasks/xover2022/bilingual-patch-contracts/2" rel="noopener noreferrer"&gt;version 2 task and model outputs&lt;/a&gt; and the &lt;a href="https://www.kaggle.com/code/xover2022/new-benchmark-task-0a5f1/output" rel="noopener noreferrer"&gt;public backing notebook&lt;/a&gt; for the embedded cases, expected states, scorer, and run artifacts (&lt;code&gt;contract_results.json&lt;/code&gt; and &lt;code&gt;contract_summary.json&lt;/code&gt;). Infrastructure errors abort the suite rather than silently shrinking the denominator.&lt;/p&gt;

&lt;p&gt;I authored the prompts and answer fixtures for this personal project. The implementation uses the &lt;a href="https://github.com/Kaggle/kaggle-benchmarks" rel="noopener noreferrer"&gt;Kaggle Benchmarks SDK&lt;/a&gt;, following the &lt;a href="https://www.kaggle.com/docs/benchmarks" rel="noopener noreferrer"&gt;official platform documentation&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
