DEV Community

Cover image for I Benchmarked My AI Carb Counter on 400 Meal Photos: 8 Models, 1 Holdout Set
Maksim Danilchenko
Maksim Danilchenko

Posted on Originally published at danilchenko.dev

I Benchmarked My AI Carb Counter on 400 Meal Photos: 8 Models, 1 Holdout Set

I build Soba, an iPhone app that looks at a photo of a meal and estimates the carbs in it. People with type 1 diabetes use numbers like that to dose insulin, so "it seems to work" wasn't a good enough answer. I ran it on 400 meal photos from two public research datasets where the carbs were weighed or recorded, and published the method, the misses and the per-meal data on the Soba accuracy page.

This post is the engineering side of that test: how the benchmark was set up, which of eight models won, and the bugs a production-path test turns up that a separate eval harness can miss. The full write-up on my blog has the per-meal-size breakdown and the comparison with published ChatGPT studies.

The setup

The photos come from Nutrition5k (Google cafeteria plates, every ingredient weighed, fixed overhead camera) and SNAPMe (95 people in the US photographing their own meals and logging them in a food record). The first is clean, controlled data. The second looks like what users actually shoot.

A few choices did most of the work:

  • A script with a fixed seed drew 25 meals from each quarter of the carb range, per dataset, so no part of the range ended up under-sampled by chance.
  • Cafeteria plates photographed within the same ten minutes often look nearly identical, so the script capped how many it took from one ten-minute window. The confidence intervals come from a bootstrap that resamples whole clusters (all photos from one SNAPMe participant, or from one ten-minute cafeteria window), because near-duplicate photos aren't independent samples.
  • There are two disjoint sets of 200. All error analysis and prompt work happened on the first set; the second stayed untouched until a result needed confirming.
  • Every photo was downscaled to 1,024 px on the long side like the app does and sent through the same server code, with the same instructions and model settings. There's no eval-only prompt.
  • The main metric is the share of meals within 10 g of the reference carbs, roughly one bread unit. Failed scans count as misses (there were none).

The results in one table

Setup Meals Within 10 g Median error
Photo only, all meals 400 68.5% 5.2 g
Photo + text hint 400 72.8% 4.5 g
Photo only, cafeteria (Nutrition5k) 200 77.5% 3.9 g
Photo + depth data, cafeteria 200 82.0% 3.1 g
Photo only, own photos (SNAPMe) 200 59.5% 6.8 g
Photo only, meals with 60 g+ carbs 42 19.0% 25.4 g

The headline 68.5% has a 95% interval of 64.0% to 72.9%. The last row shows the main failure. In 38 of the 42 largest meals the estimate was too low, with the median estimate at 64% of the real amount. In 30 of those 38, the model judged the food to weigh at least 15% less than it did, while its carbs per 100 g were usually close.

Scatter plot of 400 meals: actual carbs vs Soba's estimate. Most points sit inside the ±10 g band up to about 50 g; above 60 g most fall below it.

Every meal in the test. The shaded band is ±10 g; points below the dashed line are underestimates.

The per-meal CSV is public, so you can check the headline numbers yourself with the standard library:

import csv, io, statistics, urllib.request

URL = "https://soba-app.com/accuracy/soba-carb-benchmark-2026-10.csv"
with urllib.request.urlopen(URL) as resp:
    rows = list(csv.DictReader(io.TextIOWrapper(resp, "utf-8")))

errs = [abs(float(r["soba_total_carbs_g"]) - float(r["reference_total_carbs_g"])) for r in rows]
print(f"within 10 g: {sum(e <= 10 for e in errs) / len(errs):.1%}")
print(f"median error: {statistics.median(errs):.1f} g")
Enter fullscreen mode Exit fullscreen mode
within 10 g: 68.5%
median error: 5.2 g
Enter fullscreen mode Exit fullscreen mode

Eight models, and why the winner didn't ship

I ran eight models on the first 200 meals with Soba's prompt, one run each through OpenRouter, with the request settings Soba had before October 2026. These were separate runs, so production's numbers differ slightly from the table above:

Model Within 10 g Median time
Gemini 3.7 Flash 73.0% 3.7 s
Gemini 3.8 Flash 71.5% 4.8 s
Gemini 3.5 Flash-Lite 69.0% 2.1 s
Gemini 3.6 Flash (production) 67.5% 3.3 s
GPT-6 Luna 65.0% 6.9 s
GPT-6 Luna Pro 63.0% 11.7 s
GPT-6.1 Sol 58.0% 13.1 s
Claude Sonnet 5.5 53.0% 6.0 s

Gemini 3.7 Flash beat production by 5.5 points. That's enough of a lead to make a switch look obvious. But when you pick the best of eight on the same 200 samples, part of the winner's lead is luck, so I ran both models on the untouched second set. They tied at 70.5%, and 3.7 Flash costs more. Production stayed on Gemini 3.6 Flash.

Two more rows need a footnote. Claude Sonnet 5.5 returned an empty food item for 41 of the 200 photos through the provider OpenRouter routed it to; on the other 159 photos it scored 66.7%, right in the pack. If the harness had silently dropped empty answers, that model would have looked competitive, and the 41 empty answers would never have shown up. And GPT-6.1 Sol overestimated carbs by 10 g on average, a bias you'd only notice by looking at signed errors.

Letting Gemini 3.6 Flash reason longer raised the share within 10 g by 4.1 points over all 400 meals, but the share within 5 g and the number of errors above 30 g stayed the same, and each scan took 6.6 seconds instead of 3.5. More thinking pushed medium errors under the line and did nothing for the large ones. I kept the fast setting.

Bugs that only a full run finds

The test pushed 3,614 requests through production code. Two problems fell out that normal traffic never surfaced:

  1. The fallback model couldn't run. Soba sent a temperature parameter that the OpenAI backup didn't accept, and the request allowed only providers that support every parameter in it (OpenRouter's require_parameters option). So the OpenAI backup would have failed on every scan the moment Gemini went down. Removing the parameter moved the main model from 68.9% to 68.5%, inside run-to-run noise. The backups are now Gemini 3.7 Flash and GPT-6 Luna.
  2. Repetition loops. In 9 of 3,614 requests (0.25%), Gemini repeated one phrase over and over, and the scan took 35 to 87 seconds. A cap on answer length now cuts those off early.

A spot check of ten photos wouldn't have caught either one. The loops only appeared in the latency tail of 3,614 requests, and the dead fallback only when the backup itself was called.

What extra inputs added

Since the dominant error was portion weight, I tested what extra inputs add:

  • A text hint naming the foods (copied from the dataset records, no amounts) lifted the share within 10 g from 68.5% to 72.8%. It helped most on ambiguous photos, like a coffee with creamer that Soba had read as caramel sauce (162 g estimated without the hint, 5.5 g with it, 2.7 g in the food record). It didn't fix the main failure: large meals were still underestimated in 37 of 42 cases with a hint. The gain was also lopsided between the two sets (8.0 points vs 0.5), and real users won't type as precisely, so I treat it as an upper bound.
  • Depth data lets the app tell the model how wide the frame is in centimeters, on iPhones with LiDAR or multiple rear cameras. Replayed from Nutrition5k's depth recordings on its 200 cafeteria photos, it lifted the share within 10 g from 77.5% to 82.0%. Large meals were still underestimated in 12 of 13 cases, and angled photos haven't been measured yet.
  • The weight editor in the UI is the fallback for big plates. The app shows a weight range with each estimate and recalculates when the user enters a real weight. The real plate weight fell inside the shown range on 97 of 200 weighed plates, about half, so treat the range as a guide; typing in a weight from a kitchen scale is the reliable fix.

Takeaways for anyone benchmarking a vision LLM

  1. Keep a holdout set and don't touch it while tuning. Mine stopped a model switch that the first-round numbers made look obvious.
  2. Break results down by the size of what you're estimating. The 68.5% average hid a 19% result on the biggest meals.
  3. Count failures as misses, and look at signed errors. Empty answers and systematic bias both disappear in a careless average.
  4. Run the eval through the production code path. Mine found a dead fallback and a 0.25% latency tail that a separate eval harness could easily have missed.
  5. Find the bottleneck before you shop for models. Here it was portion weight on large plates, and the inputs that carry scale (depth, a weight the user enters) go after it more directly than a model swap. Even depth left 12 of 13 large meals underestimated.

The full method, the limits (US-only meals, overhead camera on one dataset, self-reported references on the other) and the 400-row CSV are on the Soba accuracy page. The blog version has the meal-size table and how the numbers compare with published ChatGPT-4o and ChatGPT-5 studies.

Soba isn't a medical device. Its numbers are estimates: check them against labels and a scale, and follow your diabetes care plan.

Top comments (0)