DEV Community

Cover image for Benchmarking Free AI Models for Outdoor Adventure Generation

Benchmarking Free AI Models for Outdoor Adventure Generation

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge

What I Benchmarked

For Ruta Viva AI, an outdoor exploration app, I wanted to evaluate how reliably free AI models can generate structured outdoor adventures.

The benchmark measures more than whether a model responds successfully. Each adventure must contain valid JSON, all required fields, exactly three missions, and consistent XP totals.

I designed 30 scenarios across 10 environments, with three prompt variations for each environment. For this initial pilot, I tested three scenarios per model, resulting in nine requests.

Models Tested

I evaluated three free models through OpenRouter:

  • Liquid LFM 2.5 2.6B: liquid/lfm-2.5-2.6b:free
  • NVIDIA Nemotron 3.5 Lightning: nvidia/nemotron-3.5-lightning:free
  • Apodex 1.1 Mini: apodex/apodex-1.1-mini:free

I chose these models to compare structured-output reliability and response speed using a common set of prompts.

Findings

The pilot revealed significant differences in structured-output reliability.

Model Valid JSON Median latency
Liquid LFM 2.5 2.6B 3/3 (100%) 13.57 s
NVIDIA Nemotron 3.5 Lightning 1/3 (33.3%) 41.37 s
Apodex 1.1 Mini 0/3 (0%) 6.36 s

Liquid was the most reliable model in this pilot. All three responses contained valid JSON, the required fields, exactly three missions, and consistent XP totals.

Nemotron produced one valid response; two were truncated at the output limit. Apodex was the fastest by median latency, but none of its responses produced valid JSON, and one returned no usable answer content.

My main takeaway is that speed alone does not determine usefulness. When an application depends on structured data, a fast response that cannot be parsed may be less valuable than a slower, correctly formatted response.

This benchmark has limitations: three requests per model are not enough to establish general performance. The safety check is a basic keyword heuristic, not proof of safety, and adventure quality still requires human evaluation.

Next, I would test more scenarios, repeat requests to measure consistency, and manually evaluate relevance, accessibility, variety, practicality, and safety.

My Benchmark

Public Kaggle notebook: https://www.kaggle.com/code/ezequielsalazar1/outdoor-adventure-ai-benchmark

GitHub repository: https://github.com/Ezequie1Sc/outdoor-adventure-ai-benchmark

The notebook documents the benchmark methodology and validation logic. The results above come from the pilot I ran locally.

kagglechallenge

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments. Some comments have been hidden by the post's author - find out more