This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
For Ruta Viva AI, an outdoor exploration app, I wanted to evaluate how reliably free AI models can generate structured outdoor adventures.
The benchmark measures more than whether a model responds successfully. Each adventure must contain valid JSON, all required fields, exactly three missions, and consistent XP totals.
I designed 30 scenarios across 10 environments, with three prompt variations for each environment. For this initial pilot, I tested three scenarios per model, resulting in nine requests.
Models Tested
I evaluated three free models through OpenRouter:
-
Liquid LFM 2.5 2.6B:
liquid/lfm-2.5-2.6b:free -
NVIDIA Nemotron 3.5 Lightning:
nvidia/nemotron-3.5-lightning:free -
Apodex 1.1 Mini:
apodex/apodex-1.1-mini:free
I chose these models to compare structured-output reliability and response speed using a common set of prompts.
Findings
The pilot revealed significant differences in structured-output reliability.
| Model | Valid JSON | Median latency |
|---|---|---|
| Liquid LFM 2.5 2.6B | 3/3 (100%) | 13.57 s |
| NVIDIA Nemotron 3.5 Lightning | 1/3 (33.3%) | 41.37 s |
| Apodex 1.1 Mini | 0/3 (0%) | 6.36 s |
Liquid was the most reliable model in this pilot. All three responses contained valid JSON, the required fields, exactly three missions, and consistent XP totals.
Nemotron produced one valid response; two were truncated at the output limit. Apodex was the fastest by median latency, but none of its responses produced valid JSON, and one returned no usable answer content.
My main takeaway is that speed alone does not determine usefulness. When an application depends on structured data, a fast response that cannot be parsed may be less valuable than a slower, correctly formatted response.
This benchmark has limitations: three requests per model are not enough to establish general performance. The safety check is a basic keyword heuristic, not proof of safety, and adventure quality still requires human evaluation.
Next, I would test more scenarios, repeat requests to measure consistency, and manually evaluate relevance, accessibility, variety, practicality, and safety.
My Benchmark
Public Kaggle notebook: https://www.kaggle.com/code/ezequielsalazar1/outdoor-adventure-ai-benchmark
GitHub repository: https://github.com/Ezequie1Sc/outdoor-adventure-ai-benchmark
The notebook documents the benchmark methodology and validation logic. The results above come from the pilot I ran locally.
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments. Some comments have been hidden by the post's author - find out more