This is a submission for the Kaggle Benchmarking Challenge
Last weekend I built a model to predict my triathlon splits. It's called Race Card, it uses TabPFN, and I checked it against every race I've done since 2014 before I trusted it. I was quite pleased with it.
Then the obvious question turned up, a bit later than it should have. Could I have just pasted my training log into a chatbot and asked?
So I built a benchmark to find out.
What I Benchmarked
79 real race legs (swim, bike and run) from my own racing between 2014 and 2026: half Ironmans, sea swims, half marathons. For each one, the model gets my training from before that race and nothing after it, and has to predict how fast I went. It answers with a median speed and an 80% range: a low and a high it expects the real answer to fall between eight times out of ten.
What the model sees, for each leg:
- the leg itself: sport, distance and how much it climbs
- my training load on race day: hours in the last 7 and 42 days, and kilometres in that sport
- a table of earlier sessions in the same sport: every race, plus everything from the 42 days before race day
Every row is the same data Race Card's TabPFN was trained on, from the same cut-off. That's the point: same information, same question, so the comparison is fair.
There are two versions of the task, and the difference between them is a single line of text:
- Blind. No names, no dates. Just numbers.
- Told which race. Same numbers, plus the event's name and year, for this race and for earlier ones. "Ironman 70.3 Weymouth (2019)."
I added the second one because of something Race Card got wrong. It predicted my Weymouth bike too fast every time, by 6 to 11%. It knows how much the course climbs, but it doesn't know what Weymouth's hills and wind actually feel like. A language model might.
Two scores per model. Typical miss: how far the median prediction was from my actual time, averaged over all 79 legs. Range: how many of my actual times landed inside the model's own 80% range. If the range is honest, that should be about 63 of 79.
The baselines are the same ones Race Card was checked against. TabPFN with all its rows, TabPFN with exactly the rows the language models see, and the one most of us actually use: "same as my last race".
Models Tested
11 models from five labs, all through Kaggle's model proxy: GPT-6 Astra, GPT-5.6 Terra, GPT-5.4 mini, Claude Opus 5, Claude Sonnet 5, Claude Haiku 4.5, Gemini 3.1 Pro, Gemini 3.8 Flash, Gemini 3.7 Flash, Grok 4.20 Reasoning and GLM-5.
I ran 10 model and task pairs twice to see how much a model moves between runs. The biggest gap was 0.7 points of typical miss, so the differences below that are bigger than that are worth reading. The ones smaller than that, I wouldn't lean on.
Findings
| Typical miss, blind | Typical miss, told the race | Actual time inside its 80% range (blind) | |
|---|---|---|---|
| TabPFN in Race Card, all rows | 6.8% | n/a | 63 of 79 |
| TabPFN, same rows as the models | 6.3% | n/a | 67 of 79 |
| GPT-6 Astra | 6.0% | 5.3% | 65 of 79 |
| Claude Sonnet 5 | 6.5% | 6.5% | 56 of 79 |
| Gemini 3.1 Pro | 6.5% | 5.6% | 37 of 79 |
| Gemini 3.7 Flash | 6.5% | 5.3% | 49 of 79 |
| GPT-5.6 Terra | 6.6% | 5.7% | 52 of 79 |
| Gemini 3.8 Flash | 6.6% | 5.4% | 48 of 79 |
| Grok 4.20 Reasoning | 6.8% | 6.1% | 56 of 79 |
| Claude Opus 5 | 6.8% | 5.8% | 60 of 79 |
| GLM-5 | 7.1% | 6.2% | 41 of 79 |
| Claude Haiku 4.5 | 7.7% | 7.5% | 42 of 79 |
| GPT-5.4 mini | 8.3% | 8.0% | 38 of 79 |
| "Same as my last race" | 8.3% | n/a | n/a |
Blind, all 11 beat "same as my last race"
"Same as last time" misses by 8.3% on the 76 legs where I had a previous race to copy. On those same legs, every model did better, though the weakest, GPT-5.4 mini, only just (8.1%). 7 of 11 also beat TabPFN with its full 1,000 rows.
Only 1 beat TabPFN on the same rows, though: GPT-6 Astra, at 6.0% against 6.3%. And it's the one model whose range held up as well as TabPFN's: 65 of 79.
One thing I didn't expect: TabPFN got better with fewer rows. Races plus the last six weeks gave it 6.3%, against 6.8% with everything. Most of my training sessions say very little about race day.
Told which race, 8 of 11 beat the model I built
This is the result that changed how I think about these models. All 11 models got closer when they knew the race's name, and 8 of them beat TabPFN on the same rows. The best was GPT-6 Astra, at 5.3%.
The Weymouth bikes show what the name can do. TabPFN was 10.8%, 6.2% and 8.3% too fast in 2018, 2019 and 2022. Here's Claude Opus 5, told the race:
| Weymouth bike | Actual | Opus 5, told the race | TabPFN |
|---|---|---|---|
| 2018 | 2:50:18 | 1.7% too fast | 10.8% too fast |
| 2019 | 2:44:36 | 0.9% too slow | 6.2% too fast |
| 2022 | 2:46:42 | 0.4% too fast | 8.3% too fast |
Each answer comes with a one-line reason, and for 2019 Opus wrote:
Mainly his 2018 Weymouth ride (31.95 km/h) on the same course, adjusted slightly upward for his modestly better 2019 form [...] tempered by flat-but-windy coastal conditions
That's what I'd do. Find the last time on the same course, adjust for form, remember the wind. Blind, the same model was 4 to 7% too fast on all three, like everything else.
Not every model used the name the same way, though. GPT-6 Astra noted my "slower past Weymouth result" and still went 7 to 8% too fast. Gemini 3.8 Flash got much better overall, but on Weymouth it was still 5 to 9% too fast. Its reasons mention "similar 90 km race performances" and never the course. Claude Sonnet 5 hardly moved at all: 6.5% blind, 6.5% told the race.
Most of them are far too sure of themselves
This is the bit I'd want anyone using these for planning to know. TabPFN's 80% range caught 63 of my 79 real times, about what it should. The language models caught between 37 and 65. GPT-5.4 mini's "80% range" was right 38 times out of 79, which is closer to a coin flip.
The exception was GPT-6 Astra: 65 of 79 blind and 69 told the race, so the model with the best middle number also had the most honest range. It's the one I'd actually use.
For the rest, the middle number is often good. The range around it mostly isn't. If you use one of these to plan pacing or nutrition for a race, the median is worth having, but I'd widen the range myself.
Small and fast was nearly as good as big
Gemini 3.7 Flash and Gemini 3.8 Flash were within a whisker of the flagships blind (6.5% and 6.6%), and right behind the best when told the race (5.3% and 5.4%). The small fast models from the other labs, Haiku and GPT-5.4 mini, were the weakest both ways, and the race name barely helped them.
What I'd Measure Next
- Someone else's legs. This is one athlete. I'd like to know if a model that knows Weymouth is as useful for someone who's never raced it.
- Ask for the range differently. Ask for the 10th and 90th percentile as times rather than speeds, or ask for five scenarios, and see whether the ranges get more honest.
- Give TabPFN the race name too, as a feature, and see how much of the gap closes.
How I Built It
- The dataset is built from Race Card's own pipeline: the same 79 legs and the same rows TabPFN saw. I re-ran TabPFN and all the simple baselines on exactly the published rows, so every comparison here is on identical data. TabPFN's score matches my Race Card post to within a tenth of a percent.
- Privacy. It's my training, so it's stripped down: no GPS, no coordinates, no activity names, IDs, gear, weather or time of day. A check script blocks the upload if anything else gets in.
-
The task is written with Kaggle's
kaggle-benchmarkslibrary. Each leg is one prompt with a structured answer (low, median and high speed, plus a one-line reason), scored by code against my real time. No LLM judge. The data is embedded in the task, so anyone can re-run it. - Every number in this post comes from one facts file generated from the downloaded runs. The whole week cost $25.70 of Kaggle's free model allowance.
Limits
- One athlete, 79 legs. Each model was run once or twice, not dozens of times.
- Told the race, a model could in principle know something about the event from its training. It can't know my result: my name isn't anywhere in the prompt.
- DeepSeek-R1 couldn't do the task. On every leg it wrote out its reasoning and never gave the answer in the required format. Gemma 4 31B kept hitting "heavy load" errors and never finished a run. Neither is in the table.
- A handful of legs failed for provider reasons (at most two per run) and are left out of that run's average.
My Benchmark
Run it yourself on Kaggle: Triathlon race splits
- Tasks: tri-race-splits (blind) and tri-race-splits-named (told the race). The leaderboard shows each model's latest run, scored as the share of legs within 5%. The numbers in this post average every complete run.
- Dataset: race-splits
- Race Card, the model it all started with: github.com/iamrobertmoore/race-card
Top comments (0)