DEV Community

Cover image for 1:30 AM Happens Twice on November 1. Most AI Models Picked One and Moved On.
Wishbone-Data
Wishbone-Data

Posted on

1:30 AM Happens Twice on November 1. Most AI Models Picked One and Moved On.

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge

What I Benchmarked

Time zones look like arithmetic. Add some hours, maybe cross midnight, done.

They aren't. Twice a year a local time disappears or happens twice. The US and Europe change their clocks on different weekends, so for a few weeks "New York is five hours behind London" is wrong. Some places sit on :30 or :45 offsets. Some countries changed their rules in the last few years, so a model trained on older text can be confidently out of date.

So I built Wall-Clock Traps: 59 scenarios, each asked two ways.

  • Clean: "Local date and time in Chicago, USA: 2026-03-08 02:30. What is the local date and time in London, UK at that same moment?"
  • Messy: "Backup job on the Chicago server is set for 2:30 am local on Sunday March 8th. London team wants to watch it run. What time is that in London?"

Same facts, same answer. (The answer is that 2:30 am never happens in Chicago that night. The clocks jump from 2:00 to 3:00.)

That gives 118 items in nine groups: basic conversions, US/EU gap weeks, odd offsets (Nepal, Chatham Islands, Lord Howe Island), the date line, flights that land on a clock-change day, countries that changed their rules recently, times that never happen, times that happen twice, and controls that look like traps but aren't.

Design choices that mattered:

  • The answer key is code, not me. Python's zoneinfo computes every answer from the IANA time zone database. I checked 25 by hand.
  • Exact-match grading. Every prompt asks for a last line like ANSWER: 2026-03-16 14:00, or ANSWER: NONEXISTENT / ANSWER: AMBIGUOUS. No LLM judge.
  • Wrong answers get sorted. Off by exactly an hour is a DST mistake. Off by a day is a date-line mistake. Matching a country's old rule is a stale-knowledge mistake. Flagging a valid time as impossible is a false alarm.
  • Four messy prompts carry a bad hint on purpose, like a CFO who "always just adds five hours."

Models Tested

  • Gemini 3.7 Flash (Kaggle's default model)
  • Gemma 4 26B A4B (small open-weights model)
  • Claude Sonnet 4.5
  • Gemini 2.5 Flash
  • Claude Haiku 4.5
  • GPT-5.4 mini

Same prompts for everyone, one attempt per item. I ran the benchmark twice: once on the Kaggle leaderboard, and once in an analysis notebook that keeps every answer so I could see why models missed.

Findings

Leaderboard (run 1):

Model Accuracy
Gemini 3.7 Flash 100%
Gemma 4 26B A4B 96%
Claude Sonnet 4.5 91%
Gemini 2.5 Flash 88%
Claude Haiku 4.5 73%
GPT-5.4 mini 66%

Full breakdown (run 2):

Model Accuracy Clean Messy Traps caught False alarms Stale-rule errors
Gemini 3.7 Flash 100% 100% 100% 18/18 0/12 0/16
Gemma 4 26B A4B 97% 97% 97% 17/18 0/12 2/16
Gemini 2.5 Flash 91% 88% 93% 11/18 0/12 4/16
Claude Haiku 4.5 78% 80% 76% 9/18 1/12 10/16
GPT-5.4 mini 68% 75% 61% 7/18 1/12 5/16

Claude Sonnet 4.5 is missing from run 2. 102 of its 118 calls came back as API errors, so the row would measure the API, not the model. Its leaderboard score from run 1 stands.

1. "Happens twice" is the real trap. "Never happens" mostly isn't.

When a time doesn't exist (spring forward), nearly every model noticed: 100% for three of the five, 80% for Haiku, 60% for GPT-5.4 mini. When a time happens twice (fall back), the picture flips. Gemini 3.7 Flash got all of them and Gemma got 88%. Gemini 2.5 Flash, Haiku and GPT-5.4 mini each got 12%. They did the math on one of the two possible moments and gave a confident answer. The hardest single item in the set was 1:45 am on Lord Howe Island, where the clock only moves back 30 minutes. Almost nobody flagged it.

2. Models still carry old rules for Egypt, Greenland and Paraguay.

The "changed rules" group split the field too. Haiku gave the old answer on 10 of 16 items, GPT-5.4 mini on 5, and Gemini 2.5 Flash on 4. The worst was Paraguay, which dropped its winter clock change in 2024: most models got the July time in Asunción wrong. Egypt (DST back in 2023) and Nuuk, Greenland (new offset in 2023) tripped up several models too. These aren't math mistakes. The models know a rule, just not the current one.

3. Messy wording hurt the weak models and didn't touch the strong ones.

The top two scored the same clean and messy. GPT-5.4 mini dropped 14 points (75% to 61%) and Haiku 4. Oddly, Gemini 2.5 Flash did better on messy (93% vs 88%). The planted hints worked on some: the "just add five hours" message (New York to London during the gap week) flipped two models from right to wrong.

Accuracy by group (run 2):

Model Basic US/EU gap Odd offset Date line Flight Changed rules Never happens Happens twice Control
Gemini 3.7 Flash 100% 100% 100% 100% 100% 100% 100% 100% 100%
Gemma 4 26B A4B 100% 100% 94% 100% 100% 88% 100% 88% 100%
Gemini 2.5 Flash 100% 100% 100% 100% 100% 75% 100% 12% 100%
Claude Haiku 4.5 100% 100% 88% 83% 94% 38% 80% 12% 83%
GPT-5.4 mini 100% 58% 69% 75% 75% 69% 60% 12% 58%

Where the wrong answers came from (run 2, counts):

Model Missed trap Off by 1 hour Stale rule Wrong day False alarm Other
Gemma 4 26B A4B 1 0 2 0 0 1
Gemini 2.5 Flash 6 0 4 0 0 1 (format)
Claude Haiku 4.5 9 4 10 0 1 2
GPT-5.4 mini 11 11 5 1 1 9

What surprised me

Thinking time bought accuracy almost one for one. GPT-5.4 mini wrote about 80 tokens per answer and was the cheapest run ($0.06 for all 118 items). It was also the only model that failed plain US/EU gap-week conversions, the kind of thing you'd put in a calendar invite. Gemma, a much smaller open model, wrote about 2,500 tokens per answer and scored 97%. The other surprise was run-to-run noise: Haiku scored 73% on the leaderboard and 78% on the second pass, and GPT-5.4 mini 66% then 68%, with the same prompts. A single run isn't the whole story.

What I'd measure next

  • Give the models a time zone tool and see whether they call it or still do the math in their head.
  • Ask about dates in 2028 and 2030, where a rule change could land after the training data.
  • Run each item several times. A model that's right 60% of the time on a trap isn't the same as one that's right every time.

My Benchmark

Kaggle benchmark: https://www.kaggle.com/benchmarks/wishbonedata/wall-clock-traps

Task (code, answer key and grader): https://www.kaggle.com/benchmarks/tasks/wishbonedata/wall-clock-traps

Everything is deterministic. The answer key is rebuilt from the time zone database every time the notebook runs, and the notebook prints a fingerprint so you can check whether your machine's time zone data gives the same key.

Top comments (1)

Collapse
 
reidmarlow profile image
Reid Marlow •

The asymmetry between spring forward and fall back shows up everywhere in agent scheduling. When a time skips forward, models catch the syntax-like absence because the sequence has a missing number. When a time repeats, models treat the mapping as a clean one-to-one function and pick whatever offset dominated their training tokens.

That gap is why letting an agent calculate UTC offsets in-prompt for future jobs is so risky. If an agent tries to schedule a cron task or a calendar hold with mental arithmetic, it completely misses PEP 495 fold transitions and gets burned by recent boundary changes like Asuncion dropping winter time. The only reliable boundary is forcing the agent to output the raw wall-clock string and an IANA identifier into a deterministic tool, then letting zoneinfo raise an error on ambiguous hours.