<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Wishbone-Data</title>
    <description>The latest articles on DEV Community by Wishbone-Data (@wishbone_data).</description>
    <link>https://dev.to/wishbone_data</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4151012%2F1595d61d-9945-463c-bcc8-a755162ad5cf.webp</url>
      <title>DEV Community: Wishbone-Data</title>
      <link>https://dev.to/wishbone_data</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/wishbone_data"/>
    <language>en</language>
    <item>
      <title>1:30 AM Happens Twice on November 1. Most AI Models Picked One and Moved On.</title>
      <dc:creator>Wishbone-Data</dc:creator>
      <pubDate>Wed, 30 Sep 2026 00:52:53 +0000</pubDate>
      <link>https://dev.to/wishbone_data/130-am-happens-twice-on-november-1-most-ai-models-picked-one-and-moved-on-k51</link>
      <guid>https://dev.to/wishbone_data/130-am-happens-twice-on-november-1-most-ai-models-picked-one-and-moved-on-k51</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;Kaggle Benchmarking Challenge&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Benchmarked
&lt;/h2&gt;

&lt;p&gt;Time zones look like arithmetic. Add some hours, maybe cross midnight, done.&lt;/p&gt;

&lt;p&gt;They aren't. Twice a year a local time disappears or happens twice. The US and Europe change their clocks on different weekends, so for a few weeks "New York is five hours behind London" is wrong. Some places sit on :30 or :45 offsets. Some countries changed their rules in the last few years, so a model trained on older text can be confidently out of date.&lt;/p&gt;

&lt;p&gt;So I built &lt;strong&gt;Wall-Clock Traps&lt;/strong&gt;: 59 scenarios, each asked two ways.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Clean:&lt;/strong&gt; "Local date and time in Chicago, USA: 2026-03-08 02:30. What is the local date and time in London, UK at that same moment?"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Messy:&lt;/strong&gt; "Backup job on the Chicago server is set for 2:30 am local on Sunday March 8th. London team wants to watch it run. What time is that in London?"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Same facts, same answer. (The answer is that 2:30 am never happens in Chicago that night. The clocks jump from 2:00 to 3:00.)&lt;/p&gt;

&lt;p&gt;That gives 118 items in nine groups: basic conversions, US/EU gap weeks, odd offsets (Nepal, Chatham Islands, Lord Howe Island), the date line, flights that land on a clock-change day, countries that changed their rules recently, times that never happen, times that happen twice, and controls that look like traps but aren't.&lt;/p&gt;

&lt;p&gt;Design choices that mattered:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The answer key is code, not me.&lt;/strong&gt; Python's &lt;code&gt;zoneinfo&lt;/code&gt; computes every answer from the IANA time zone database. I checked 25 by hand.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Exact-match grading.&lt;/strong&gt; Every prompt asks for a last line like &lt;code&gt;ANSWER: 2026-03-16 14:00&lt;/code&gt;, or &lt;code&gt;ANSWER: NONEXISTENT&lt;/code&gt; / &lt;code&gt;ANSWER: AMBIGUOUS&lt;/code&gt;. No LLM judge.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Wrong answers get sorted.&lt;/strong&gt; Off by exactly an hour is a DST mistake. Off by a day is a date-line mistake. Matching a country's &lt;em&gt;old&lt;/em&gt; rule is a stale-knowledge mistake. Flagging a valid time as impossible is a false alarm.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Four messy prompts carry a bad hint on purpose&lt;/strong&gt;, like a CFO who "always just adds five hours."&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Models Tested
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Gemini 3.7 Flash (Kaggle's default model)&lt;/li&gt;
&lt;li&gt;Gemma 4 26B A4B (small open-weights model)&lt;/li&gt;
&lt;li&gt;Claude Sonnet 4.5&lt;/li&gt;
&lt;li&gt;Gemini 2.5 Flash&lt;/li&gt;
&lt;li&gt;Claude Haiku 4.5&lt;/li&gt;
&lt;li&gt;GPT-5.4 mini&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Same prompts for everyone, one attempt per item. I ran the benchmark twice: once on the Kaggle leaderboard, and once in an analysis notebook that keeps every answer so I could see &lt;em&gt;why&lt;/em&gt; models missed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Findings
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Leaderboard (run 1):&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Accuracy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.7 Flash&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 4 26B A4B&lt;/td&gt;
&lt;td&gt;96%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 4.5&lt;/td&gt;
&lt;td&gt;91%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 2.5 Flash&lt;/td&gt;
&lt;td&gt;88%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Haiku 4.5&lt;/td&gt;
&lt;td&gt;73%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.4 mini&lt;/td&gt;
&lt;td&gt;66%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Full breakdown (run 2):&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Accuracy&lt;/th&gt;
&lt;th&gt;Clean&lt;/th&gt;
&lt;th&gt;Messy&lt;/th&gt;
&lt;th&gt;Traps caught&lt;/th&gt;
&lt;th&gt;False alarms&lt;/th&gt;
&lt;th&gt;Stale-rule errors&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.7 Flash&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;18/18&lt;/td&gt;
&lt;td&gt;0/12&lt;/td&gt;
&lt;td&gt;0/16&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 4 26B A4B&lt;/td&gt;
&lt;td&gt;97%&lt;/td&gt;
&lt;td&gt;97%&lt;/td&gt;
&lt;td&gt;97%&lt;/td&gt;
&lt;td&gt;17/18&lt;/td&gt;
&lt;td&gt;0/12&lt;/td&gt;
&lt;td&gt;2/16&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 2.5 Flash&lt;/td&gt;
&lt;td&gt;91%&lt;/td&gt;
&lt;td&gt;88%&lt;/td&gt;
&lt;td&gt;93%&lt;/td&gt;
&lt;td&gt;11/18&lt;/td&gt;
&lt;td&gt;0/12&lt;/td&gt;
&lt;td&gt;4/16&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Haiku 4.5&lt;/td&gt;
&lt;td&gt;78%&lt;/td&gt;
&lt;td&gt;80%&lt;/td&gt;
&lt;td&gt;76%&lt;/td&gt;
&lt;td&gt;9/18&lt;/td&gt;
&lt;td&gt;1/12&lt;/td&gt;
&lt;td&gt;10/16&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.4 mini&lt;/td&gt;
&lt;td&gt;68%&lt;/td&gt;
&lt;td&gt;75%&lt;/td&gt;
&lt;td&gt;61%&lt;/td&gt;
&lt;td&gt;7/18&lt;/td&gt;
&lt;td&gt;1/12&lt;/td&gt;
&lt;td&gt;5/16&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Claude Sonnet 4.5 is missing from run 2. 102 of its 118 calls came back as API errors, so the row would measure the API, not the model. Its leaderboard score from run 1 stands.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. "Happens twice" is the real trap. "Never happens" mostly isn't.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When a time doesn't exist (spring forward), nearly every model noticed: 100% for three of the five, 80% for Haiku, 60% for GPT-5.4 mini. When a time happens twice (fall back), the picture flips. Gemini 3.7 Flash got all of them and Gemma got 88%. Gemini 2.5 Flash, Haiku and GPT-5.4 mini each got &lt;strong&gt;12%&lt;/strong&gt;. They did the math on one of the two possible moments and gave a confident answer. The hardest single item in the set was 1:45 am on Lord Howe Island, where the clock only moves back 30 minutes. Almost nobody flagged it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Models still carry old rules for Egypt, Greenland and Paraguay.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The "changed rules" group split the field too. Haiku gave the &lt;em&gt;old&lt;/em&gt; answer on 10 of 16 items, GPT-5.4 mini on 5, and Gemini 2.5 Flash on 4. The worst was Paraguay, which dropped its winter clock change in 2024: most models got the July time in Asunción wrong. Egypt (DST back in 2023) and Nuuk, Greenland (new offset in 2023) tripped up several models too. These aren't math mistakes. The models know a rule, just not the current one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Messy wording hurt the weak models and didn't touch the strong ones.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The top two scored the same clean and messy. GPT-5.4 mini dropped 14 points (75% to 61%) and Haiku 4. Oddly, Gemini 2.5 Flash did &lt;em&gt;better&lt;/em&gt; on messy (93% vs 88%). The planted hints worked on some: the "just add five hours" message (New York to London during the gap week) flipped two models from right to wrong.&lt;/p&gt;

&lt;p&gt;Accuracy by group (run 2):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Basic&lt;/th&gt;
&lt;th&gt;US/EU gap&lt;/th&gt;
&lt;th&gt;Odd offset&lt;/th&gt;
&lt;th&gt;Date line&lt;/th&gt;
&lt;th&gt;Flight&lt;/th&gt;
&lt;th&gt;Changed rules&lt;/th&gt;
&lt;th&gt;Never happens&lt;/th&gt;
&lt;th&gt;Happens twice&lt;/th&gt;
&lt;th&gt;Control&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.7 Flash&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 4 26B A4B&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;94%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;88%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;88%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 2.5 Flash&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;75%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;12%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Haiku 4.5&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;88%&lt;/td&gt;
&lt;td&gt;83%&lt;/td&gt;
&lt;td&gt;94%&lt;/td&gt;
&lt;td&gt;38%&lt;/td&gt;
&lt;td&gt;80%&lt;/td&gt;
&lt;td&gt;12%&lt;/td&gt;
&lt;td&gt;83%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.4 mini&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;58%&lt;/td&gt;
&lt;td&gt;69%&lt;/td&gt;
&lt;td&gt;75%&lt;/td&gt;
&lt;td&gt;75%&lt;/td&gt;
&lt;td&gt;69%&lt;/td&gt;
&lt;td&gt;60%&lt;/td&gt;
&lt;td&gt;12%&lt;/td&gt;
&lt;td&gt;58%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Where the wrong answers came from (run 2, counts):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Missed trap&lt;/th&gt;
&lt;th&gt;Off by 1 hour&lt;/th&gt;
&lt;th&gt;Stale rule&lt;/th&gt;
&lt;th&gt;Wrong day&lt;/th&gt;
&lt;th&gt;False alarm&lt;/th&gt;
&lt;th&gt;Other&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 4 26B A4B&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 2.5 Flash&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;1 (format)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Haiku 4.5&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.4 mini&lt;/td&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;What surprised me&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Thinking time bought accuracy almost one for one. GPT-5.4 mini wrote about 80 tokens per answer and was the cheapest run ($0.06 for all 118 items). It was also the only model that failed plain US/EU gap-week conversions, the kind of thing you'd put in a calendar invite. Gemma, a much smaller open model, wrote about 2,500 tokens per answer and scored 97%. The other surprise was run-to-run noise: Haiku scored 73% on the leaderboard and 78% on the second pass, and GPT-5.4 mini 66% then 68%, with the same prompts. A single run isn't the whole story.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What I'd measure next&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Give the models a time zone tool and see whether they call it or still do the math in their head.&lt;/li&gt;
&lt;li&gt;Ask about dates in 2028 and 2030, where a rule change could land after the training data.&lt;/li&gt;
&lt;li&gt;Run each item several times. A model that's right 60% of the time on a trap isn't the same as one that's right every time.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  My Benchmark
&lt;/h2&gt;

&lt;p&gt;Kaggle benchmark: &lt;a href="https://www.kaggle.com/benchmarks/wishbonedata/wall-clock-traps" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/wishbonedata/wall-clock-traps&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Task (code, answer key and grader): &lt;a href="https://www.kaggle.com/benchmarks/tasks/wishbonedata/wall-clock-traps" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/tasks/wishbonedata/wall-clock-traps&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Everything is deterministic. The answer key is rebuilt from the time zone database every time the notebook runs, and the notebook prints a fingerprint so you can check whether your machine's time zone data gives the same key.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
