<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Jaye Nichols</title>
    <description>The latest articles on DEV Community by Jaye Nichols (@jayenichols).</description>
    <link>https://dev.to/jayenichols</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4158197%2Fb362196f-ba34-45d8-b154-0e5060c1d337.jpg</url>
      <title>DEV Community: Jaye Nichols</title>
      <link>https://dev.to/jayenichols</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/jayenichols"/>
    <language>en</language>
    <item>
      <title>Access Atlas: keyboard repair plans from structured W3C guidance</title>
      <dc:creator>Jaye Nichols</dc:creator>
      <pubDate>Fri, 02 Oct 2026 20:10:01 +0000</pubDate>
      <link>https://dev.to/jayenichols/access-atlas-keyboard-repair-plans-from-structured-w3c-guidance-3ii5</link>
      <guid>https://dev.to/jayenichols/access-atlas-keyboard-repair-plans-from-structured-w3c-guidance-3ii5</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/sanity-2026-09-16"&gt;Sanity Challenge, Path One: Ship an Agent That Queries Real Content&lt;/a&gt;.&lt;/em&gt; &lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;Access Atlas is a small keyboard and focus repair agent. Give it a symptom such as “Tab leaves my modal” or “arrowing through tabs loads every slow panel,” and it selects source-linked guidance from Sanity. The product plan brings together the repair steps, prerequisites, constraints, suggested manual checks, expected observations, and original W3C source.&lt;/p&gt;

&lt;p&gt;The corpus contains seven repair topics. Its boundaries matter: a menu trigger needs an actual ARIA menu; automatic tab activation depends on panel timing; visible focus is a different question from contrast; and Focus Not Obscured (Minimum) does not demand complete component exposure. The records preserve those distinctions instead of turning a search hit into a generic “accessibility fix.”&lt;/p&gt;

&lt;p&gt;The model chooses which records to retrieve. The displayed repair plan is assembled directly from the retrieved published fields. A model draft remains inspectable in the trace, but it cannot silently add a new repair technique to the product plan. Every proposed manual check is marked &lt;code&gt;notRun&lt;/code&gt;. This is implementation guidance, not a completed interface audit or conformance certificate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://access-atlas-jaye-2026.netlify.app/" rel="noopener noreferrer"&gt;Open Access Atlas&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhxkwigk90ftkiiqtp2gx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhxkwigk90ftkiiqtp2gx.png" alt="Desktop local workbench showing the recorded modal run" width="800" height="1756"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Local workbench showing actual recorded modal run — desktop.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc12f3ezii1wg401fhf18.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc12f3ezii1wg401fhf18.png" alt="Mobile local workbench showing the recorded modal run" width="375" height="1800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Local workbench showing actual recorded modal run — mobile.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The public explorer replays actual saved local Qwen2.5:7b + live Sanity Context MCP runs. Choose a recorded question, inspect the retrieved plan, open the source links, and expand the trace to see the model-selected tools, generated GROQ, real results, timestamps, and call counts. The evidence desk reads the public dataset's current counts and provenance separately.&lt;/p&gt;

&lt;p&gt;Replaying a recorded question does not run new inference. Fresh questions use the source's local server at &lt;code&gt;http://127.0.0.1:4340&lt;/code&gt;, local Ollama, and an authorized server-side Context Viewer token. The README explains setup and how to reproduce the corpus in a Sanity project you control.&lt;/p&gt;

&lt;p&gt;The actual nine-question evaluation recorded &lt;strong&gt;9/9 structural passes&lt;/strong&gt;: seven source-linked plans and two abstentions for color-contrast calculation and whole-site/legal certification. Across the nine runs there were &lt;strong&gt;26 local model calls and 25 real Context MCP calls&lt;/strong&gt;, including nine schema bootstraps. &lt;a href="https://access-atlas-jaye-2026.netlify.app/runs/index.json" rel="noopener noreferrer"&gt;Recorded-run manifest&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Those two abstentions are enforced by explicit corpus guards after real retrieval. Both original model drafts had citation errors; the product withheld the drafts. The trace preserves &lt;code&gt;draftStatus&lt;/code&gt;, the applicable boundary rule, and the final presentation so this distinction is inspectable.&lt;/p&gt;

&lt;p&gt;Each case specifies expected guide IDs, required source URLs, concepts to review, and claims to avoid. The automated checks measure execution, disposition, guide retrieval, and URL membership; they do not prove every model sentence is entailed by a source. Seven Node regression checks also passed. These results apply to the stated nine questions, with no claim of an external benchmark or whole-site conformance.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://access-atlas-jaye-2026.netlify.app/code/index.html" rel="noopener noreferrer"&gt;Browse the public repository&lt;/a&gt;, &lt;a href="https://access-atlas-jaye-2026.netlify.app/code/README.md.html" rel="noopener noreferrer"&gt;read the setup guide&lt;/a&gt;, or &lt;a href="https://access-atlas-jaye-2026.netlify.app/access-atlas-source.zip" rel="noopener noreferrer"&gt;download the source ZIP&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Clone the repository:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://access-atlas-jaye-2026.netlify.app/access-atlas.git
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The source includes the Sanity schemas, public seed, model/tool adapter, MCP client, local server, replay UI, golden questions, evaluator, and regression checks. Node.js 24 and local Ollama run the agent; a server-side organization Context Viewer token authorizes retrieval. The public demo requires no account to inspect its replays and public content.&lt;/p&gt;

&lt;p&gt;Original application code is MIT licensed. The guidance is paraphrased from W3C WAI APG pattern pages and WCAG 2.2 Understanding documents; it retains its own source terms, copyright, and modification notices. The full &lt;a href="https://access-atlas-jaye-2026.netlify.app/licenses.html" rel="noopener noreferrer"&gt;source attribution and W3C notice&lt;/a&gt; is visible in the app. Qwen and Ollama are credited there under their upstream licenses.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Used Sanity
&lt;/h2&gt;

&lt;p&gt;I used the full dataset with embeddings enabled, through a real read-only Sanity Context MCP endpoint. This is Context's live dataset/GROQ mode, which the challenge allows, rather than a separately built website/file Knowledge Base.&lt;/p&gt;

&lt;p&gt;The 21 public documents form seven related triples:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Type&lt;/th&gt;
&lt;th&gt;What it contributes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;guide&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Problem/symptoms, pattern, prerequisites, constraints, repair steps, unsupported questions, source/check references&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;source&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Primary W3C URL, verified sections, checked time, informative status, license/copyright/derivation metadata&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;check&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Proposed manual procedure, expected observations, limitations, &lt;code&gt;notRun&lt;/code&gt; status, reverse guide/source references&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The harness first initializes MCP and calls the real &lt;code&gt;initial_context&lt;/code&gt; tool to retrieve the schema overview. Then the local model selects &lt;code&gt;search_guides&lt;/code&gt; and &lt;code&gt;read_guide&lt;/code&gt;. These are narrow application tools: the adapter executes real Context &lt;code&gt;groq_query&lt;/code&gt; calls behind both. Search ranks semantic similarity against the symptom. Reading a guide dereferences &lt;code&gt;references[0]&lt;/code&gt; and &lt;code&gt;checks[0]&lt;/code&gt;, joining its source and review procedure in the same response.&lt;/p&gt;

&lt;p&gt;That relationship is the useful part. A guide alone does not carry its source provenance and review expectations. A source URL alone does not carry the scope constraints or a proposed procedure. The joined response gives the renderer specific fields for each, including the fact that no check has been performed. Keyword matching could find a title; this plan depends on retrieving and composing the related records.&lt;/p&gt;

&lt;p&gt;Two real development failures shaped the implementation. Giving the model raw GROQ produced a parsing error and an answer with a URL that had not been retrieved; the grounding gate rejected it. Narrow search/read tools fixed the query construction. A subsequent free-form answer could still include a technique absent from the retrieved guide. The product now renders the published steps, constraints, and checks directly, retaining the model draft separately for inspection. A search-only response is retried for a detailed relational read. Explicit contrast/certification guards enforce those corpus boundaries after retrieval. Citation membership is useful evidence, but it is not a semantic correctness certificate.&lt;/p&gt;

&lt;p&gt;The runtime cannot write to the dataset or change a user's interface. The corpus can be edited through Sanity Studio; the next local retrieval reads current published content. Saved public runs preserve the original retrieval results as historical evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sanity Project Details
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Project ID: &lt;code&gt;2mflxxa8&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Public dataset: &lt;code&gt;accessatlas&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Document types: &lt;code&gt;guide&lt;/code&gt;, &lt;code&gt;source&lt;/code&gt;, &lt;code&gt;check&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Corpus: 21 documents across seven topics&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://access-atlas-jaye-2026.sanity.studio/" rel="noopener noreferrer"&gt;Studio&lt;/a&gt; — authorized editing; no login is needed for public data inspection&lt;/li&gt;
&lt;li&gt;&lt;a href="https://2mflxxa8.api.sanity.io/v2026-10-02/data/query/accessatlas?query=%7B%22guides%22%3Acount%28*%5B_type%3D%3D%22guide%22%5D%29%2C%22sources%22%3Acount%28*%5B_type%3D%3D%22source%22%5D%29%2C%22checks%22%3Acount%28*%5B_type%3D%3D%22check%22%5D%29%7D" rel="noopener noreferrer"&gt;Public document counts&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is a separate dataset and application from my Path Two entry. It contains public technical guidance and provenance, with no private mission information or invented experience/testing records.&lt;/p&gt;

&lt;h2&gt;
  
  
  Agent Session
&lt;/h2&gt;

&lt;p&gt;The replay explorer exposes the task agent's actual model/tool trace alongside the displayed plan. It distinguishes the session bootstrap's &lt;code&gt;initial_context&lt;/code&gt; call from the model's selected search/read calls, records the actual Context query/results, and keeps the original model draft inspectable. These runtime traces are not a DEV Agent Sessions upload.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>sanitychallenge</category>
      <category>sanity</category>
      <category>ai</category>
    </item>
    <item>
      <title>FrameFlip48: four models, one frame switch, two kinds of failure</title>
      <dc:creator>Jaye Nichols</dc:creator>
      <pubDate>Fri, 02 Oct 2026 18:59:38 +0000</pubDate>
      <link>https://dev.to/jayenichols/frameflip48-four-models-one-frame-switch-two-kinds-of-failure-41gf</link>
      <guid>https://dev.to/jayenichols/frameflip48-four-models-one-frame-switch-two-kinds-of-failure-41gf</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;Kaggle Benchmarking Challenge&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Benchmarked
&lt;/h2&gt;

&lt;p&gt;FrameFlip48 measures whether a model tracks the translation frame in a short robot-motion program. Each of 24 matched pairs has identical numbers and commands. One prompt says &lt;code&gt;ACTIVE_FRAME = LOCAL&lt;/code&gt;; its counterpart says &lt;code&gt;ACTIVE_FRAME = WORLD&lt;/code&gt;. Cases run in separate conversations, so the model never receives the counterpart's answer.&lt;/p&gt;

&lt;p&gt;The robot lives on an integer grid. Its turns are multiples of 90 degrees. In WORLD mode, a translation uses the fixed east/north axes. In LOCAL mode, those axes rotate with its current heading. Turns rotate around the robot's own center; they do not move the center.&lt;/p&gt;

&lt;p&gt;For example, start at &lt;code&gt;(2, -1)&lt;/code&gt; facing east, turn 90 degrees counterclockwise, then &lt;code&gt;MOVE(3, 2)&lt;/code&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Translation frame&lt;/th&gt;
&lt;th&gt;Final world position&lt;/th&gt;
&lt;th&gt;Final heading&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;WORLD&lt;/td&gt;
&lt;td&gt;&lt;code&gt;(5, 1)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;90°&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LOCAL&lt;/td&gt;
&lt;td&gt;&lt;code&gt;(0, 2)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;90°&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Coordinates can look plausible while being expressed in the wrong frame. Exact arithmetic gives this test a ground truth that needs no LLM judge or subjective rating.&lt;/p&gt;

&lt;p&gt;The corpus has 48 cases, balanced across four initial headings and programs of 2, 4 and 8 commands. Each pair has distinct final LOCAL/WORLD positions. Fixed-seed generation used only oracle properties; no model outputs selected the cases. The corpus was frozen before evaluation, with SHA-256 &lt;code&gt;590c916ee01d408b580f40a8f1d00262e79360cec46c18194ba8c6eb883d8e87&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models Tested
&lt;/h2&gt;

&lt;p&gt;The planned lineup used three labs' small models, selected from Kaggle's available catalog before seeing outputs. All ran on the same fixed task version 1 on October 2, 2026:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Gemini 3.1 Flash-Lite Preview: &lt;code&gt;google/gemini-3.1-flash-lite-preview&lt;/code&gt;, run &lt;code&gt;3955369&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;GPT-5.4 nano: &lt;code&gt;openai/gpt-5.4-nano-2026-03-17&lt;/code&gt;, run &lt;code&gt;3955370&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Claude Haiku 4.5: &lt;code&gt;anthropic/claude-haiku-4-5@20251001&lt;/code&gt;, run &lt;code&gt;3955371&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Kaggle task creation also automatically ran Gemini 3.7 Flash (&lt;code&gt;google/gemini-3.7-flash&lt;/code&gt;, run &lt;code&gt;3955292&lt;/code&gt;). That additional result is included rather than discarded.&lt;/p&gt;

&lt;p&gt;Every case gets one response in a fresh conversation, temperature zero, with no supplied tools. Reasoning uses each model's platform default, so this is not an equal-compute comparison. The prompt requests &lt;strong&gt;only one JSON object&lt;/strong&gt; with the final position and normalized heading. The frozen parser accepts plain or fenced JSON and integer-valued numbers, but rejects surrounding explanation. Formatting failures are reported separately. Transport/quota failures abort a run instead of shrinking the denominator; all four runs completed.&lt;/p&gt;

&lt;p&gt;There was also one earlier real SDK validation with Flash-Lite: 10/48 cases and 1/24 pairs. It is separate from the server comparison below, with no prompt/scorer changes or pooling of the two runs. Mock SDK checks and deterministic baselines are not model evaluations. The full workflow used $0.2543933 of Kaggle's free inference allowance; actual paid spend was $0.&lt;/p&gt;

&lt;h2&gt;
  
  
  Findings
&lt;/h2&gt;

&lt;p&gt;The primary score is the fraction of 24 pairs for which &lt;strong&gt;both&lt;/strong&gt; variants have the exact final pose and satisfy the output contract. These are the &lt;a href="https://www.kaggle.com/benchmarks/tasks/jayenichols/frameflip48?compare=true" rel="noopener noreferrer"&gt;public task results&lt;/a&gt;, independently re-scored from the downloaded responses:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Correct pairs / 24&lt;/th&gt;
&lt;th&gt;Strict correct cases / 48&lt;/th&gt;
&lt;th&gt;Format failures&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.7 Flash&lt;/td&gt;
&lt;td&gt;24&lt;/td&gt;
&lt;td&gt;48&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Haiku 4.5&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;48&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.1 Flash-Lite Preview&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.4 nano&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Pose correctness and output compliance need separate diagnostics.&lt;/strong&gt; Claude's zero is dominated by its format: all 48 responses explain the calculation before presenting a final fenced JSON object. For example, its answer to &lt;code&gt;n8-h90-r0-world&lt;/code&gt; explains eight commands and ends with &lt;code&gt;{"x":9,"y":-13,"heading_deg":270}&lt;/code&gt;. That pose is correct, but the complete response violates the requested output contract.&lt;/p&gt;

&lt;p&gt;After observing this, an independent reviewer applied the &lt;strong&gt;same post-hoc extraction rule to all 192 saved responses&lt;/strong&gt;: require exactly one valid pose JSON object in the entire response and require it to be terminal, optionally fenced; preceding explanation is ignored. This was analysis of existing outputs, not a model rerun, and it does &lt;strong&gt;not&lt;/strong&gt; replace the frozen leaderboard score:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Post-hoc correct poses / 48&lt;/th&gt;
&lt;th&gt;Post-hoc correct pairs / 24&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.7 Flash&lt;/td&gt;
&lt;td&gt;48&lt;/td&gt;
&lt;td&gt;24&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Haiku 4.5&lt;/td&gt;
&lt;td&gt;44&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.1 Flash-Lite Preview&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.4 nano&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Claude's four remaining pose errors are all LOCAL position errors; every heading and all 24 WORLD positions are correct. Flash-Lite's strict successes are all WORLD cases, and all are two-command programs. Thus the weak strict scores have different causes: output compliance for Claude, and substantial pose errors for the two smaller models in these runs. Gemini 3.7 Flash reaches the ceiling of this corpus; a harder extension would be needed to measure its margin.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pair accuracy protects against a fixed-frame shortcut.&lt;/strong&gt; A deterministic program that always uses WORLD scores 24/48 individual cases but solves 0/24 pairs. Always-LOCAL has the same pattern; the exact oracle solves 48/48 cases and 24/24 pairs. These are local code baselines, not LLM results.&lt;/p&gt;

&lt;p&gt;The outputs do not establish that the weaker models used that shortcut. GPT-5.4 nano has just one exact opposite-frame match: for &lt;code&gt;n2-h180-r0-local&lt;/code&gt;, it returns &lt;code&gt;(0,6,90)&lt;/code&gt; instead of &lt;code&gt;(0,4,90)&lt;/code&gt;, matching the WORLD counterfactual. Other errors include incorrect headings and unrelated positions. An opposite-frame match is an observable signature, not proof of internal reasoning. WORLD is computationally simpler than LOCAL; the paired design controls numbers and commands, not equal reasoning difficulty.&lt;/p&gt;

&lt;p&gt;This is a small synthetic test of 2D quarter-turn motion, not a robotics reliability certificate. Pair members are mathematically dependent, and one response per case cannot estimate run-to-run variance. Different platform defaults and the Flash-Lite validation/server discrepancy also limit broad model rankings. A future version could add mixed per-command frames, 3D rotations and a separately reported tool-enabled condition, frozen before inspecting its outputs.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Benchmark
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.kaggle.com/benchmarks/jayenichols/frameflip48" rel="noopener noreferrer"&gt;FrameFlip48 public Kaggle benchmark collection&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The collection contains the owned &lt;a href="https://www.kaggle.com/benchmarks/tasks/jayenichols/frameflip48/1" rel="noopener noreferrer"&gt;task version 1&lt;/a&gt; and all four completed models. It aggregates the numeric task with &lt;em&gt;Average of task scores&lt;/em&gt;. Reproduce or inspect it using the &lt;a href="https://www.kaggle.com/code/jayenichols/new-benchmark-task-62aca/notebook" rel="noopener noreferrer"&gt;source notebook&lt;/a&gt;, &lt;a href="https://www.kaggle.com/code/jayenichols/new-benchmark-task-62aca/output" rel="noopener noreferrer"&gt;run outputs&lt;/a&gt;, and &lt;a href="https://www.kaggle.com/datasets/jayenichols/frameflip48-original-cases" rel="noopener noreferrer"&gt;original cases, prompts and gold answers&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Original code is MIT; original synthetic data is CC0-1.0. The public collection is also under Kaggle's Apache 2.0 publication terms. No external dataset, private information or participant responses were used. The task uses the &lt;a href="https://github.com/Kaggle/kaggle-benchmarks" rel="noopener noreferrer"&gt;Kaggle Benchmarks SDK&lt;/a&gt;. Seven local tests passed; an independent integer-matrix oracle checked all 48 gold and 48 opposite-frame poses. A separate review reproduced all strict metrics and verified that all 192 scored prompts/responses match the native SDK traces.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Autonomous workflow disclosure:&lt;/strong&gt; Codex agents designed, implemented, evaluated, reviewed and wrote this entry on my behalf. I completed the account's required personal verification. The benchmark, outputs and interpretation are public so readers can assess the work directly.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Deadline Loom: a schedule with a memory</title>
      <dc:creator>Jaye Nichols</dc:creator>
      <pubDate>Fri, 02 Oct 2026 17:57:37 +0000</pubDate>
      <link>https://dev.to/jayenichols/deadline-loom-a-schedule-with-a-memory-a9k</link>
      <guid>https://dev.to/jayenichols/deadline-loom-a-schedule-with-a-memory-a9k</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/sanity-2026-09-16"&gt;Sanity Challenge, Path Two: Vibe-Code Something Strange&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;Deadline Loom asks a small, uncomfortable question: &lt;strong&gt;which version of a deadline leaves enough time to finish?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It is an Astro planning workbench for someone comparing several pieces of time-sensitive work. A source-linked deadline is one input; estimated effort, daily capacity, and prerequisites are others. Moving an estimate or selecting a proposed source version changes the schedule. A saved local planning decision remembers its tracked planning inputs and becomes stale when those inputs change.&lt;/p&gt;

&lt;p&gt;The strange part is treating a schedule as an argument with a memory. A reassuring timeline should still show where its dates came from, what its assumptions are, and what could invalidate an earlier decision.&lt;/p&gt;

&lt;p&gt;The sample content contains three real DEV challenges, with official source links checked on October 2. Build-hour and human-minute estimates are illustrative. The earlier-deadline revision is explicitly simulated. No green timeline represents eligibility, an organizer update, a likely win, or money earned.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://deadline-loom-jaye-2026.netlify.app/" rel="noopener noreferrer"&gt;Open Deadline Loom&lt;/a&gt;. No login is needed. When the query completes, the header shows &lt;strong&gt;CONNECTED / PUBLIC SANITY CONTENT&lt;/strong&gt;. A failed read instead shows an explicit sample-mode notice.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7d0f1a59yif0vadgdd4i.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7d0f1a59yif0vadgdd4i.png" alt="Deadline Loom desktop workbench" width="800" height="1462"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdeadline-loom-jaye-2026.netlify.app%2Fmobile.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdeadline-loom-jaye-2026.netlify.app%2Fmobile.png" alt="Deadline Loom on a narrow screen" width="375" height="4545"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Try this short walkthrough:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Change the clock from Chicago to UTC. The start input changes from 13:00 CDT to 18:00 UTC; the planned instant stays the same.&lt;/li&gt;
&lt;li&gt;Select the simulated earlier source version for the first thread. The selected deadline moves seven hours earlier, while the recorded organizer date remains visible.&lt;/li&gt;
&lt;li&gt;Add a reason and record a &lt;strong&gt;local planning decision&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Increase that thread's estimated work to 30 hours. The plan shows missed deadlines, and your decision becomes &lt;strong&gt;stale&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Export the scenario as an ICS calendar file. Its events use UTC instants and include the source and planning disclaimer.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Scenario edits and notes stay in your browser. The separate authenticated editorial workflow runs in Sanity Studio. The public decision memory displays a real persisted &lt;strong&gt;simulated approval&lt;/strong&gt; performed during an authorized agent-operated Studio rehearsal; it does not imply that a human clicked or an organizer changed the deadline. Judges can inspect that published decision without editor credentials.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://deadline-loom-jaye-2026.netlify.app/code/" rel="noopener noreferrer"&gt;Browse the public source repository&lt;/a&gt; or &lt;a href="https://deadline-loom-jaye-2026.netlify.app/deadline-loom-source.zip" rel="noopener noreferrer"&gt;download the complete source ZIP&lt;/a&gt;. Original code is MIT licensed; third-party packages retain their own licenses.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://deadline-loom-jaye-2026.netlify.app/deadline-loom.git
&lt;span class="nb"&gt;cd &lt;/span&gt;deadline-loom
pnpm &lt;span class="nb"&gt;install
&lt;/span&gt;pnpm dev
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The project uses Astro 7.3.5 and Sanity 6.17.0. The repository includes the dependency lockfile, &lt;a href="https://deadline-loom-jaye-2026.netlify.app/code/README.md.html" rel="noopener noreferrer"&gt;setup README&lt;/a&gt;, six document schemas, Studio document actions, seed validator, and planner checks. For CONNECTED mode locally, copy &lt;code&gt;.env.example&lt;/code&gt; to &lt;code&gt;.env&lt;/code&gt;, set &lt;code&gt;PUBLIC_SANITY_PROJECT_ID=2mflxxa8&lt;/code&gt; and &lt;code&gt;PUBLIC_SANITY_DATASET=production&lt;/code&gt;, restart the dev server, and open &lt;code&gt;http://127.0.0.1:4321&lt;/code&gt;. Those values are public; no write token is needed. To use a different Sanity project, configure its public routing values and allow its browser origin through CORS with credentials disabled.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Build Process
&lt;/h2&gt;

&lt;p&gt;This was built on October 2 using &lt;strong&gt;Codex as the AI-native development environment&lt;/strong&gt;. Delegated Codex agents researched primary rules, developed the content model, generated the interface and review actions, operated the authorized setup, and ran the checks. This account's owner authorized the work; the coding, browser rehearsal, and verification described here were agent-operated. I am not presenting them as manual coding or human user testing.&lt;/p&gt;

&lt;p&gt;The initial direction was broader: a source-linked opportunity review desk. Reviewing current entries revealed similar territory, including &lt;a href="https://dev.to/di_wang_3516db206ab336792/proof-desk-receipts-before-rewards-3m7n"&gt;Proof Desk: receipts before rewards&lt;/a&gt;. The project narrowed to a distinct interaction: source versions alter the time available, effort consumes it, and planning decisions remember their context. No competitor code, text, or images were copied.&lt;/p&gt;

&lt;p&gt;These are condensed build directives, rather than verbatim transcript excerpts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Keep official dates, proposed revisions, and effort assumptions separate; make the selected revision change a schedule without overwriting the accepted fact.&lt;/li&gt;
&lt;li&gt;Give every local decision a fingerprint of its tracked planning inputs; changing those inputs must make its previous decision stale.&lt;/li&gt;
&lt;li&gt;Build a public reader without a write token, and put authenticated editorial transitions in Studio with a reason and optimistic revision checks.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those constraints shaped the schema and UI more than a visual prompt did. The cream-and-green workbench puts the clock, threads, timeline, source versions, and decision memory in one place. The mobile layout stacks those controls while keeping the timeline readable.&lt;/p&gt;

&lt;p&gt;The main integration correction came from a real false positive: &lt;strong&gt;an anonymous Node query succeeded while the browser still showed sample mode&lt;/strong&gt;. The origin was missing from Sanity's CORS configuration. Adding the exact local and deployed origins with credentials disabled made the hosted browser show live content and the persisted approval. A successful script read was insufficient evidence of a working public demo.&lt;/p&gt;

&lt;p&gt;A second correction separated the static build from content loading. The build now embeds the labeled fixture without a network request; the browser fetches published content and refreshes once a minute while visible. That gives the public view an honest failure state.&lt;/p&gt;

&lt;p&gt;Final independent review caught another concrete gap: the decision fingerprint covered explicit browser estimate edits but omitted source-provided effort and human-minute defaults. A refreshed default could therefore change the schedule without invalidating its old decision. The fix fingerprints effective work values, including the local override when present. Regression checks compare old and changed source records: changed effort or human minutes changes both the schedule and fingerprint, while an unused effort default preserves the fingerprint under an explicit override. No public dataset mutation was needed for these checks.&lt;/p&gt;

&lt;p&gt;The custom Studio actions approve or reject an evidence revision and save its decision in one transaction. A real approval can apply the proposed accepted value; a simulated approval only records the rehearsal. The observed rehearsal approved &lt;code&gt;revision-simulated-deadline&lt;/code&gt; with a reason, created a public decision, and retained the official deadline &lt;code&gt;2026-10-05T06:59:00Z&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;This uses a custom document-review process built with Sanity's Document Actions API. It does not use the App SDK or the Sanity Workflows product.&lt;/p&gt;

&lt;p&gt;Checks actually run:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Production Astro and Studio builds, plus TypeScript checking, passed.&lt;/li&gt;
&lt;li&gt;Planner checks covered equivalent Chicago/Pacific/India deadline instants, a winter offset, earliest-deadline order, capacity, revision impact, UTC calendar times, text escaping, UTF-8 line folding, remote effort/minute changes, and preservation of explicit overrides.&lt;/li&gt;
&lt;li&gt;Browser interactions confirmed the timezone-preserving start, seven-hour revision impact, local decision recording, late schedules, and stale decision memory.&lt;/li&gt;
&lt;li&gt;Anonymous public queries and the hosted browser confirmed the Sanity content and simulated approval. The exported ICS file was downloaded and inspected.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The scheduling model remains deliberately small: serial work using average daily capacity, with human minutes added separately. It does not model hourly availability or holidays. Prerequisites are a checklist, not an eligibility validator. The estimates are assumptions, and the rejection action has not been independently rehearsed in the live dataset.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sanity Project Details
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Project ID:&lt;/strong&gt; &lt;code&gt;2mflxxa8&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dataset:&lt;/strong&gt; &lt;code&gt;production&lt;/code&gt; (public)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Published document types:&lt;/strong&gt; &lt;code&gt;competition&lt;/code&gt;, &lt;code&gt;evidenceSource&lt;/code&gt;, &lt;code&gt;sourceClaim&lt;/code&gt;, &lt;code&gt;evidenceRevision&lt;/code&gt;, &lt;code&gt;reviewDecision&lt;/code&gt;, &lt;code&gt;projectNote&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;sourceClaim → competition + evidenceSource
evidenceRevision → competition + evidenceSource
reviewDecision → exact evidenceRevision
projectNote → estimate, verification, or simulation context
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Sources remain independent documents with publisher, URL, and checked time. Claims preserve what was reported. A revision represents a proposed change, and a decision points to the exact revision reviewed. This avoids treating a copied date as both evidence and acceptance.&lt;/p&gt;

&lt;p&gt;The initial seed contains 28 documents. The Studio rehearsal added one persisted simulated decision, bringing the public dataset to 29. The browser reads published content without an Authorization header. Public routing values are included; write credentials, personal account details, identity documents, and financial records are excluded from the dataset and public source.&lt;/p&gt;

&lt;h2&gt;
  
  
  Agent Session
&lt;/h2&gt;

&lt;p&gt;No public agent transcript is attached. The build account above describes the work actually completed, and the repository contains the implementation and reproducible planner checks.&lt;/p&gt;

&lt;p&gt;Submitted by &lt;a href="https://dev.to/jayenichols"&gt;@jayenichols&lt;/a&gt;. AI agents were development tools; there are no additional human teammates.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>sanitychallenge</category>
      <category>sanity</category>
      <category>ai</category>
    </item>
  </channel>
</rss>
