This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
In computer vision pipelines for automotive manufacturing, fleet maintenance, and tire recycling, reading tire specifications is a notorious challenge.
Tire sidewalls feature embossed, black-on-black rubber text following the wheel's circular arc. When training custom object detection models (such as YOLO) to detect individual character boxes (/, 0-9, R), traditional post-processing relies on sorting bounding boxes by horizontal X-coordinates to reconstruct the tire code (e.g., 205/55R16).
However, this naive approach fails on real-world tires:
-
Curvature and Rotation: When text curves along the side of the wheel arc or is oriented vertically / upside down, coordinate-based sorting reverses or scrambles the sequence (e.g.
225/60R18is read as81R06/522). - Low-Contrast Black-on-Black Rubber: Dirt, shadows, and tire wear make standard OCR tools choke.
The Question:
Can frontier general-purpose Vision-Language Models (VLMs) perform zero-shot OCR and structured extraction of ISO metric tire size specifications ([Width]/[Aspect]R[Rim]) directly from raw tire sidewall photos without any fine-tuning, and act as automated auditors for human-annotated YOLO datasets?
To measure this, I created a curated test suite of 51 tire sidewall images across 22 distinct tire dimensions (covering standard horizontal text, steep curved arcs, and rotated positions), backed by verified ground-truth labels.
Models Tested
The live Kaggle Benchmark evaluates 11 frontier models across 5 AI providers and open-weight architectures:
- Google: Gemini 3.8 Flash, Gemini 3.5 Flash-Lite, Gemini 2.5 Pro, and Gemma 4 31B (Open Weights).
- Anthropic: Claude Opus 5.5, Claude Sonnet 5.5, and Claude Haiku 5.5.
- OpenAI: GPT-6.1 Sol and gpt-oss-20b (open-weights).
- Alibaba: Qwen 3 Next 80B Thinking.
- Zhipu AI: GLM-5.
(Note: During task exploration, Gemini 3.7 Flash [46/51], Gemini 3 Flash Preview [45/51], and Claude Sonnet 5 [44/51] were also evaluated on the underlying task. DeepSeek-R1 was correctly rejected by the benchmark runner as a text-only reasoning model lacking a visual encoder, and Grok 4.6 encountered upstream proxy connection timeouts).
Evaluation Setup & Task Logic
Each model was prompted with an isolated context per image:
"Inspect the tire sidewall in this image carefully. Locate the standard ISO metric tire size specification (format: WWW/AARDI, e.g. 205/55R16, 225/60R18, 235/65R17). Return ONLY the exact 9-character tire size code."
The benchmark recorded:
- Exact String Match: Did the full 9-character code match the ground truth?
- Component Accuracy: Individual correctness of Section Width (mm), Aspect Ratio (%), and Rim Diameter (inches).
- Orientation Robustness: Performance on horizontal vs. curved/vertical arcs.
Findings
Leaderboard Results
The table below reflects the official leaderboard from the live Kaggle Benchmark:
| Rank | Model | Provider | Exact Matches | Exact Accuracy | Key Behavior |
|---|---|---|---|---|---|
| #1 | Google Gemini 3.8 Flash | 46 / 51 | 90.20% | Tied #1; top multimodal speed & curved text accuracy | |
| #1 | Anthropic Claude Opus 5.5 | Anthropic | 46 / 51 | 90.20% | Tied #1; flagship frontier reasoning & flawless ISO extraction |
| #3 | Google Gemini 3.5 Flash-Lite | 45 / 51 | 88.24% | Exceptional efficiency and accuracy for a lightweight model | |
| #4 | Google Gemini 2.5 Pro | 44 / 51 | 86.27% | Deep reasoning; minor slip on worn edge numerals | |
| #5 | Google Gemma 4 31B | Google (OSS) | 41 / 51 | 80.39% | Best open-weights model; matched Claude 5.5 models |
| #5 | Anthropic Claude Sonnet 5.5 | Anthropic | 41 / 51 | 80.39% | Solid reasoning; missed low-contrast rim digits |
| #5 | Anthropic Claude Haiku 5.5 | Anthropic | 41 / 51 | 80.39% | Tied with Sonnet 5.5 at a fraction of latency |
| #8 | OpenAI GPT-6.1 Sol | OpenAI | 39 / 51 | 76.47% | Tended to confuse scuffed numbers (3 vs 2, 4 vs 6) |
| #9 | Alibaba Qwen 3 Next 80B Thinking | Alibaba | 3 / 51 | 5.88% | Overthought reasoning; violated strict 9-char constraint |
| #10 | Zhipu AI GLM-5 | Zhipu AI | 1 / 51 | 1.96% | Low parsing accuracy on low-contrast embossments |
| #11 | OpenAI gpt-oss-20b | OpenAI (OSS) | 0 / 51 | 0.00% | Failed strict 9-character formatting constraint |
Note on Kaggle UI Score Display: On the Kaggle Benchmarks leaderboard UI, raw numeric scores appear with an automatic × 100% display badge (e.g.
46.00renders as4600.0%and41.00renders as4100.0%, representing 46 and 41 correct items out of 51 total samples, or 90.20% and 80.39%). The table above provides the normalized, exact percentage rates.
Cost vs. Accuracy: The Pareto Frontier
When deploying vision models in production—such as scanning thousands of tire sidewalls daily across an automotive fleet or recycling facility—raw accuracy must be balanced against inference cost. Kaggle Benchmarks plotted the Score vs. Total Cost Pareto Frontier across all evaluated models:
Highlights from the Efficiency Frontier:
- The Maximum ROI Champion — Claude Haiku 5.5: Scoring 41 / 51 (80.39%) at an ultra-low total evaluation cost of ~$0.005, Claude Haiku 5.5 anchors the steep rise of the Pareto frontier. For high-volume automated triage, its cost-per-accuracy ratio is unmatched.
- The Practical Sweet Spot — Gemini 3.5 Flash-Lite: At only ~$0.02 for the entire 51-image suite, Gemini 3.5 Flash-Lite achieved 45 / 51 (88.24%), outperforming models costing 5x to 15x more. It sits squarely in the Efficient quadrant.
- The Peak Performance Frontier — Claude Opus 5.5: Claude Opus 5.5 forms the right anchor of the Pareto frontier, delivering top-tier accuracy of 46 / 51 (90.20%) alongside Gemini 3.8 Flash, with flawless formatting precision at ~$0.15 total cost.
- The Inefficient Quadrant: Models like Qwen 3 Next 80B Thinking (~$0.14) and GLM-5 (~$0.30) fell into the bottom-right Inefficient quadrant because their internal reasoning traces generated excessive output tokens, running up API costs while failing the strict 9-character extraction constraint.
Key Insights & Surprises
1. Zero-Shot VLMs Completely Solved the "Rotation Trap"
In our YOLO dataset, human annotations on rotated or curved tire arcs were often scrambled by bounding-box sorting algorithms (e.g. producing 81R06/522 or 71R05/522).
All top-tier frontier VLMs effortlessly read these rotated, curved tires in the correct semantic human reading order, outputting 225/60R18 and 225/50R17 flawlessly. This demonstrates that multimodal LLMs perceive continuous text streams gestalt-style rather than through brittle geometric axis projections.
2. The Crown: Gemini 3.8 Flash & Claude Opus 5.5 Tied at #1 (90.20%)
Google's Gemini 3.8 Flash and Anthropic's flagship Claude Opus 5.5 shared top honors with identical scores of 46 / 51 (90.20%). While Claude Opus 5.5 demonstrated immense precision in distinguishing faint embossed characters, Gemini 3.8 Flash delivered that same accuracy with lightning-fast inference times.
Right behind them, Google's lightweight Gemini 3.5 Flash-Lite scored 45 / 51 (88.24%), outperforming heavyweights like Gemini 2.5 Pro (44 / 51) and GPT-6.1 Sol (39 / 51)—proving that specialized vision distillation can outperform pure model parameter scale.
3. Gemma 4 31B: A Triumph for Open-Weights Industrial Vision
One of the most exciting results is Gemma 4 31B scoring 80.39% (41 / 51). It matched proprietary frontier flagships Claude Sonnet 5.5 and Claude Haiku 5.5, while beating OpenAI GPT-6.1 Sol (76.47%). For manufacturing plants and tire depots requiring local, on-premise execution without cloud API dependency, Gemma 4 31B offers production-grade zero-shot OCR capability.
4. Claude Haiku 5.5 Punched Way Above Its Weight
Anthropic's Claude Haiku 5.5 tied Claude Sonnet 5.5 exactly at 41 / 51 (80.39%). Getting flagship-grade multimodal OCR accuracy at Haiku's speed and cost profile makes it an exceptional candidate for high-throughput batch auditing.
5. The "Reasoning Model Trap" on Pure Extraction Tasks
A fascinating divergence occurred with Qwen 3 Next 80B Thinking (3 / 51, 5.88%):
While reasoning and "thinking" architectures excel at multi-step mathematics and coding, they frequently failed this pure extraction benchmark because their outputs included internal reasoning preambles, chain-of-thought commentary, or Markdown wrappers (e.g.,
```text\n205/55R16\n```
) rather than adhering strictly to the required bare 9-character string.
6. The Dominant Physical Failure Mode: Low-Contrast Width Digits
Across the ~10% of cases where top models failed, the errors were not random hallucinations:
- In Sample 39 (
235/45R18), models predicted225/45R18(confusing a scuffed embossed3with a2). - In Sample 34 (
245/50R20), models predicted265/50R20(confusing4with6). - In Sample 29 (
255/60R18), models misread the width and rim on a heavily worn sidewall.
Notice that in almost every failure case, the Aspect Ratio and Construction ('R') remained 100% correct. The error was almost exclusively an off-by-ten millimeter width confusion caused by low-contrast black rubber embossments.
7. Automated Dataset Auditing Potential
During benchmark preparation, we discovered a human labeling typo in our original dataset where an annotator typed 3 instead of R (255/60318). During inference, the models naturally output 255/60R18 based on semantic domain awareness of tire codes, proving that VLMs can serve as effective automated validation auditors for industrial labeling pipelines.
My Benchmark
You can inspect the full benchmark, code, and live evaluation runs directly on Kaggle:
- Benchmark Page: TireSidewall-Bench on Kaggle
- Task & Model Leaderboard: tire-sidewall-ocr Task Details
- Model Comparison View: Interactive Model Compare
- Benchmark Dataset: Tire Sidewall Benchmark Dataset on Kaggle

Top comments (2)
The strict-format rule is doing a lot of the ranking here. Qwen 3 Next at 3/51 and GLM-5 at 1/51 are failing "return ONLY 9 characters", and your own note says Qwen wrapped answers in a code fence. That measures instruction-following, not tire reading. Rescoring with a regex that pulls the first \d{3}/\d{2}R\d{2} from the reply would tell you whether those two can read rubber at all, and the gap between strict and lenient scores is itself worth a column.
On the top of the table: 46/51 has a 95% Wilson interval of about 79-96%, 41/51 is 68-89%, 39/51 is 63-86%. Fisher exact for 46 vs 39 is p about 0.11, and 45 vs 44 (Flash-Lite over Gemini 2.5 Pro) is p = 1.0, so "outperforming heavyweights" is one image. With the same 51 images for every model, a paired test is the right one: if Opus got 7 images right that Sol missed and Sol got none that Opus missed, that is p about 0.016, but 9 vs 2 is only about 0.065. Publishing the per-image hit matrix would settle it.
Also, 51 photos of 22 sizes are not 51 independent draws. If one size has 6 photos and a model misreads its 2 vs 4 digit once, it can miss all six. Reporting accuracy per size, or resampling by size, gives a wider and more honest interval. Do the 5 misses shared by the top models fall on the same images (sample 39, 235/45R18 read as 225)? That overlap would show how much is label or photo difficulty rather than model difference.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.