<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: simon levy</title>
    <description>The latest articles on DEV Community by simon levy (@simon_levy_e0b34a39073843).</description>
    <link>https://dev.to/simon_levy_e0b34a39073843</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4157081%2F21d85301-6a32-4103-b3c4-77f4e9164d25.jpeg</url>
      <title>DEV Community: simon levy</title>
      <link>https://dev.to/simon_levy_e0b34a39073843</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/simon_levy_e0b34a39073843"/>
    <language>en</language>
    <item>
      <title>Rehearsal Mirror helps you find your own story before an interview</title>
      <dc:creator>simon levy</dc:creator>
      <pubDate>Fri, 02 Oct 2026 21:43:58 +0000</pubDate>
      <link>https://dev.to/simon_levy_e0b34a39073843/rehearsal-mirror-helps-you-find-your-own-story-before-an-interview-1d1n</link>
      <guid>https://dev.to/simon_levy_e0b34a39073843/rehearsal-mirror-helps-you-find-your-own-story-before-an-interview-1d1n</guid>
      <description>&lt;p&gt;Entry for &lt;a href="https://dev.to/challenges/hacktoberfest-weekend-2026-10-01"&gt;Hacktoberfest Weekend Challenge: Build for a Friend&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;Rehearsal Mirror is a browser-local interview practice tool. You write down experiences in Situation, Task, Action and Result cards, choose a practice question, and use a small local AI model to find relevant cards. Then you rehearse an answer in your own words.&lt;/p&gt;

&lt;p&gt;I built Rehearsal Mirror for someone close to me who wants help organising real experiences into interview answers.&lt;/p&gt;

&lt;p&gt;The problem it targets is specific: having an experience worth talking about, but struggling to recall and organise it when a question arrives. The app gives those experiences a place to live and a way to find them again. It has 12 questions across personal-assistant, customer-support and administration roles, plus a 60- or 90-second practice timer.&lt;/p&gt;

&lt;p&gt;The person practising controls the content. The model retrieves existing notes and quotes a related field directly. It cannot create a past achievement, rewrite a result or claim that an experience happened.&lt;/p&gt;

&lt;p&gt;This is a technical pilot. It has not yet been handed to its intended recipient, so there is no recipient feedback or claim of improved interview outcomes. Every example in the demo is fictional.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://rehearsal-mirror-private-pilot.simonlevy00.chatgpt.site/" rel="noopener noreferrer"&gt;Try the live demo&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Try the following sequence:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Select &lt;strong&gt;Load fictional demo&lt;/strong&gt;. This loads three clearly labelled examples: a calendar collision, a delayed delivery and a duplicate invoice. If cards already exist, the app asks for confirmation before replacing them; export any notes you want to keep first.&lt;/li&gt;
&lt;li&gt;Choose the first personal-assistant question and select &lt;strong&gt;Download &amp;amp; load local AI&lt;/strong&gt;. The first load needs internet and downloads approximately 50 MB of model/runtime assets, depending on compression.&lt;/li&gt;
&lt;li&gt;Select &lt;strong&gt;Find relevant stories&lt;/strong&gt;. The calendar card should rank first. The result includes an excerpt copied from that card and a link to the full story.&lt;/li&gt;
&lt;li&gt;Switch to the upset-customer question or the administration question about correcting a record. In the recorded checks, the delivery and invoice cards respectively ranked first.&lt;/li&gt;
&lt;li&gt;Write a practice answer and start the timer. Pause, resume or restart whenever needed. Export a JSON backup if you want to keep your notes outside this browser.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The first model download is an explicit choice. Before it completes, the app clearly says it is in manual mode and lets you browse your cards yourself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/simon-levy01/rehearsal-mirror" rel="noopener noreferrer"&gt;Source code on GitHub&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The original application code is MIT-licensed. Model, library and runtime licenses are recorded separately in the dependency notices. The built application serves static files and has no application backend, account system or upload endpoint.&lt;/p&gt;

&lt;p&gt;The repository includes the application source, tests, model provenance, dependency notices and licenses. Generated build output is not committed. To build and run it locally, use Node.js 22.12+ or 24 and Python 3:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm ci
npm run build
npm &lt;span class="nb"&gt;test
&lt;/span&gt;python tools/serve.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For a browser-only Linux install, the README documents how to skip the unused native CUDA download.&lt;/p&gt;

&lt;p&gt;Then open &lt;code&gt;http://127.0.0.1:4173&lt;/code&gt; and keep the local server running.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;The retrieval path uses &lt;a href="https://huggingface.co/docs/transformers.js/index" rel="noopener noreferrer"&gt;Transformers.js&lt;/a&gt; 3.8.1 and the &lt;a href="https://huggingface.co/Xenova/all-MiniLM-L6-v2" rel="noopener noreferrer"&gt;Xenova ONNX conversion of all-MiniLM-L6-v2&lt;/a&gt;, pinned to revision &lt;code&gt;751bff37182d3f1213fa05d7196b954e230abad9&lt;/code&gt;. The &lt;a href="https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2" rel="noopener noreferrer"&gt;upstream Sentence Transformers model&lt;/a&gt; and the conversion declare Apache-2.0 licensing.&lt;/p&gt;

&lt;p&gt;The quantized model runs in a dedicated browser worker through WebAssembly. It turns the selected question and each story into normalized, 384-dimensional embeddings. Cosine similarity ranks the stories. A second comparison selects the most related STAR field from each of the top three cards, and the interface displays that field unchanged.&lt;/p&gt;

&lt;p&gt;This keeps the AI's job narrow and inspectable. A related result is a prompt to review a memory. It is not a measure of English proficiency, employability or the quality of an answer. The interface never presents similarity as a hiring score.&lt;/p&gt;

&lt;p&gt;Practice text stays in browser local storage; embeddings stay in worker memory. During the model download, Hugging Face and its file hosts receive ordinary connection metadata. The application does not send stories or answers to a remote inference service.&lt;/p&gt;

&lt;p&gt;The production-build verification recorded four passing unit tests and 18 passing browser checks in Chrome on Windows. Those checks covered real model loading and ranking, timers, validated backup import/export, deletion, download failure, labels and mobile layout. They also inserted distinct synthetic markers into a story and an answer, then checked page and worker requests for either marker. Neither appeared in a request URL or body in the recorded flows.&lt;/p&gt;

&lt;p&gt;Another check edited a synthetic story and ran retrieval with browser networking disabled after model loading. That succeeded without additional requests. This supports the specific claim that a loaded model can run inference offline. It does not establish offline startup after closing or reloading the app.&lt;/p&gt;

&lt;p&gt;There are important limits. The three ranking examples are smoke tests, not a retrieval benchmark. Similar experiences may be harder to distinguish. Long text is truncated by the model, so concise cards work best. Browser storage and exported backups are unencrypted, and clearing browser data can remove notes. Other browsers, low-memory phones and full screen-reader accessibility have not been independently verified.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;p&gt;The open model is the part that makes semantic retrieval work. Without it, this build has manual card browsing, but no automatic story matching.&lt;/p&gt;

&lt;p&gt;Local inference also makes the privacy design practical. The app can compare a question with personal notes without forwarding those notes to a hosted model endpoint. Once loaded, the model has no per-query API charge and can keep matching while the browser is offline. That still requires a suitable device, an initial download and a working app page.&lt;/p&gt;

&lt;p&gt;The weights, runtime and application code can be inspected separately. Pinning the model revision makes the current behaviour easier to reproduce, while the source can be adapted to try another embedding model later. The small model also makes the trade-offs visible: limited context, modest retrieval tests and no generated coaching.&lt;/p&gt;

&lt;p&gt;The next useful evaluation would be whether the intended recipient can add a truthful experience and find it when they need it. That remains untested. For this version, the demonstrated result is a working, local-first practice tool with a bounded AI feature and explicit limits.&lt;/p&gt;

&lt;h3&gt;
  
  
  AI assistance
&lt;/h3&gt;

&lt;p&gt;OpenAI Codex and OpenAI writing tools provided substantial assistance with this application's implementation, automated verification and article draft. The technical claims above are grounded in the source and recorded synthetic tests. No recipient testimonial was generated or inferred.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
    </item>
    <item>
      <title>A CSV benchmark with two zero scores and a useful diagnostic</title>
      <dc:creator>simon levy</dc:creator>
      <pubDate>Fri, 02 Oct 2026 20:45:27 +0000</pubDate>
      <link>https://dev.to/simon_levy_e0b34a39073843/a-csv-benchmark-with-two-zero-scores-and-a-useful-diagnostic-n9e</link>
      <guid>https://dev.to/simon_levy_e0b34a39073843/a-csv-benchmark-with-two-zero-scores-and-a-useful-diagnostic-n9e</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;Kaggle Benchmarking Challenge&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Benchmarked
&lt;/h2&gt;

&lt;p&gt;Inventory reconciliation sounds straightforward: match a count to an exported product row and replace its new quantity. The difficult part is preserving everything else and recognizing when an update should be blocked.&lt;/p&gt;

&lt;p&gt;CSV Contract Integrity tests this with synthetic, offline inputs. Each model receives an export CSV, a count CSV, an explicit column mapping, and a reconciliation contract. It must return one JSON object containing a candidate CSV, row-level issues, and any fatal file error.&lt;/p&gt;

&lt;p&gt;Small differences matter. The SKU &lt;code&gt;0005&lt;/code&gt; is a string identifier. Locations &lt;code&gt;North&lt;/code&gt; and &lt;code&gt;north&lt;/code&gt; are distinct. A blank new-quantity cell differs from &lt;code&gt;0&lt;/code&gt;. Quoted fields can contain commas or newlines. Unknown columns must survive unchanged. Duplicate keys, conflicting values, and incomplete product information can block an update.&lt;/p&gt;

&lt;p&gt;The scorer allows harmless representation differences: reordered headers, CSV quote style, LF or CRLF line endings, and equivalent integer spellings in the editable quantity field. Every protected cell still requires exact string preservation. Candidate row order matters too.&lt;/p&gt;

&lt;p&gt;The source contains 72 fixtures: 48 original cases and 24 compositions. Twelve original cases are reserved for development, leaving 36 original evaluation cases plus 24 compositions in each reported run. Those compositions come from six recipes with four related transformations each. The suite is small and deliberately inspectable, with correlated cases rather than 60 independent samples of inventory work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models Tested
&lt;/h2&gt;

&lt;p&gt;The current comparison uses &lt;a href="https://www.kaggle.com/code/simonlevy42/csv-inventory-contract-benchmark/output?scriptVersionId=354731097" rel="noopener noreferrer"&gt;Gemini 3.7 Flash&lt;/a&gt; (&lt;code&gt;google/gemini-3.7-flash&lt;/code&gt;) and &lt;a href="https://www.kaggle.com/code/simonlevy42/csv-inventory-contract-benchmark/output?scriptVersionId=354732397" rel="noopener noreferrer"&gt;Claude Haiku 4.5&lt;/a&gt; (&lt;code&gt;anthropic/claude-haiku-4-5@20251001&lt;/code&gt;) through Kaggle Model Proxy. Comparing two providers gives this narrow task more than one model response pattern to inspect.&lt;/p&gt;

&lt;p&gt;Both completed the same 60 selected cases under the frozen contract and scorer. Each case received a fresh chat. The task requested temperature 0, seed 0, no tools, and a maximum of 8,000 output tokens per response. These were text-generation calls, without an enforced JSON output schema or generated-code execution.&lt;/p&gt;

&lt;p&gt;An inspection of the installed SDK confirmed that its string-response path adds no formatting instruction and returns model text unchanged. The reported results therefore include the behavior of the configured model and provider interface, not an SDK operation that adds or removes code fences.&lt;/p&gt;

&lt;h2&gt;
  
  
  Findings
&lt;/h2&gt;

&lt;p&gt;Both models returned 60 nonempty responses. Both scored &lt;strong&gt;0/60 on the strict contract&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The primary score requires the raw response to satisfy the JSON interface and then match the allowed CSV semantics and issue report. A response wrapped in Markdown fails the raw JSON requirement even when its contents look correct to a reader.&lt;/p&gt;

&lt;p&gt;A separate post hoc diagnostic removed exactly one complete outer &lt;code&gt;json&lt;/code&gt; code fence. It accepted no surrounding prose, repaired no incomplete JSON, and made no other correction. It then applied the same scorer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Gemini 3.7 Flash: &lt;strong&gt;59/60&lt;/strong&gt;, comprising 35/36 original cases and 24/24 compositions&lt;/li&gt;
&lt;li&gt;Claude Haiku 4.5: &lt;strong&gt;14/60&lt;/strong&gt;, comprising 14/36 original cases and 0/24 compositions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Eight Haiku responses contained fenced JSON with trailing explanation or multiple blocks. The diagnostic therefore did not apply to them. They remain in its denominator.&lt;/p&gt;

&lt;p&gt;These numbers answer two questions. The strict result measures whether the specified interface can consume the response directly and obtain a correct result. The diagnostic measures correctness after one narrowly defined formatting adjustment.&lt;/p&gt;

&lt;p&gt;Reporting only the zero scores would conceal the content differences that the diagnostic reveals. Presenting 59/60 or 14/60 as the official score would conceal the interface failures. Both views belong in the report.&lt;/p&gt;

&lt;p&gt;Gemini’s remaining case, &lt;code&gt;preservation-05&lt;/code&gt;, contains an unusually long identifier. Its response had incomplete JSON and 448 visible characters, while usage reported 7,996 output tokens against the requested 8,000 cap. That token total is not a visible-text length. The cause is unestablished; a larger budget is a follow-up experiment, not a demonstrated fix.&lt;/p&gt;

&lt;p&gt;The comparison is descriptive. It does not establish a general ranking of these model families. A different prompt, structured-output interface, token budget, provider configuration, or repeat run could produce different results.&lt;/p&gt;

&lt;h3&gt;
  
  
  What the benchmark cannot establish
&lt;/h3&gt;

&lt;p&gt;This is conservative offline validation. No Shopify import or real inventory change was performed. A failed contract verdict does not demonstrate inventory damage.&lt;/p&gt;

&lt;p&gt;The emitted-row mutation diagnostic also has a narrower meaning: it counts unexpected emitted rows relative to the expected candidate. Omitted rows fail overall correctness but are not counted as emitted-row mutations. For both models, all 60 primary responses failed JSON parsing and mutation remained unassessed. Unknown must never become “zero damage.”&lt;/p&gt;

&lt;p&gt;Coverage is incomplete. Several error branches in the written contract have no evaluation fixture in this version. Expected outputs and grading code also need scrutiny: passing expected answers through the scorer checks consistency, not the independent correctness of every label.&lt;/p&gt;

&lt;h3&gt;
  
  
  What to measure next
&lt;/h3&gt;

&lt;p&gt;A useful next experiment would predeclare separate interface conditions: the existing raw-text prompt, a one-fence adapter, and supported structured-output modes. Their scores should remain separate rather than retroactively changing this benchmark's primary metric.&lt;/p&gt;

&lt;p&gt;Further work should vary the output budget for long preservation cases, add independently authored cases, fill the coverage gaps, and repeat runs. The current results are one completed run per model, without a formal statistical generalization claim.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Benchmark
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://www.kaggle.com/benchmarks/simonlevy42/csv-contract-integrity" rel="noopener noreferrer"&gt;CSV Contract Integrity on Kaggle&lt;/a&gt; contains the &lt;a href="https://www.kaggle.com/benchmarks/tasks/simonlevy42/csv-inventory-contract-integrity-v6" rel="noopener noreferrer"&gt;CSV Inventory Contract Integrity task&lt;/a&gt;. The collection and task are now published publicly on Kaggle.&lt;/p&gt;

&lt;p&gt;The benchmark's primary result is the mean of 60 strict binary verdicts. Source materials retain the contract, fixtures, expected outputs, scorer, request hashes, raw responses, usage, and separate diagnostics. These make the evaluation inspectable; requested seeds do not guarantee identical future provider output. There is no hidden-test claim or production-safety guarantee.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI disclosure: Fully Autonomous.&lt;/strong&gt; AI assistance was used to create the prototype, synthetic fixtures, grading code, tests, documentation, and this article. The article was drafted primarily by an AI agent. The grader itself does not use an LLM judge.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
