DEV Community

Cover image for TRAIL TUTOR
Rajab Baig
Rajab Baig

Posted on

TRAIL TUTOR

Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass Submission 🌿

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass

TrailTutor AI: Use AI for One Minute, Explore the Real World for Ten

What I built

TrailTutor AI is an outdoor learning companion designed around a simple idea:

AI should sometimes help us leave the screen, not stay on it.

A learner chooses:

  • their age
  • where they are
  • a topic
  • how much time they have

TrailTutor then generates one short outdoor mission with:

  • a real-world observation task
  • exactly two guiding questions
  • a safety instruction
  • a reflection prompt

The learner reads the mission, puts the device away, completes the activity outdoors, then comes back only to reflect.

That is why I designed TrailTutor around the principle:

The screen should be the shortest part of the experience.

Live Demo https://youtu.be/sXFEXBelZhA

RENDER APP

https://trailtutor-ai.onrender.com/
Local host:http://127.0.0.1:8000/

Source Code

https://github.com/rajab-rajab/TrailTutor-AI

How it works

The public application is a lightweight FastAPI service deployed on Render.

For the live Gemma path, TrailTutor uses:

google/diffusiongemma-26b-a4b-it

through NVIDIA's hosted API.

A typical request contains:

{
  "age": 13,
  "environment": "school playground",
  "topic": "plants",
  "duration_minutes": 10
}
Enter fullscreen mode Exit fullscreen mode

The response follows a small structured format:

{
  "mission": "...",
  "observe": "...",
  "questions": ["...", "..."],
  "safety": "...",
  "reflection": "..."
}
Enter fullscreen mode Exit fullscreen mode

This structure is intentionally restrictive. TrailTutor is not supposed to become another long conversation. Its job is to produce a useful mission quickly and then get out of the learner's way.

Why Gemma

I wanted the core AI component to use an open-weight model rather than treat the model as a completely closed black box.

Gemma gave me a strong foundation for generating short, structured educational activities while keeping the architecture flexible enough to swap or self-host models later.

The live TrailTutor application uses DiffusionGemma for mission generation.

Fine-tuning with Tinker

I also wanted to test whether a general open-weight model could be specialized specifically for TrailTutor's mission format.

I used Tinker to LoRA fine-tune:

Qwen/Qwen3.5-4B

The training setup was deliberately small and reproducible:

  • 80 supervised training examples
  • 20 held-out evaluation prompts
  • LoRA rank: 16
  • 1 training epoch
  • batch size: 4
  • learning rate: 1e-4

The dataset covers:

  • ages 7–17
  • 8 outdoor environments
  • 20 topics
  • 9 activity durations

The held-out prompts have no exact input overlap with the training examples.

What changed after fine-tuning

I evaluated the untuned and tuned versions of the same Qwen3.5-4B model on the 20 held-out prompts.

The final corrected evaluation produced:

Metric Untuned Tinker-tuned
Overall structured-output score 95% 100%
Valid JSON 19/20 20/20
All required fields 19/20 20/20
Exactly two questions 19/20 20/20
Outdoor action present 19/20 20/20
Safety present 19/20 20/20
Reflection present 19/20 20/20
Mission under 70 words 19/20 20/20
Average latency 3.355 s 2.907 s

That is a 5 percentage-point improvement in the held-out structural score.

Average latency also decreased by about 13.4%.

An evaluation mistake I found

My first evaluation reported a larger improvement: 60% to 80%.

That result turned out to be misleading.

Some model outputs contained two consecutive valid JSON objects. My original parser took everything from the first opening brace to the last closing brace, then tried to decode the whole string as one JSON object.

That caused valid generations to be scored as failures with errors such as:

Extra data

Instead of hiding that mistake, I kept the original result in the repository and added a corrected v1.1 evaluation that parses the first complete JSON object while preserving the raw output.

The corrected result is the one I report here:

95% → 100%

I think keeping both versions is important because reproducible AI evaluation also means documenting when the evaluator itself was wrong.

Why Render

TrailTutor is deployed publicly on Render using the project's Dockerfile.

Render hosts:

  • the FastAPI backend
  • the frontend
  • the public API endpoints

The computationally heavy Gemma inference happens through the model provider, so the web service itself stays lightweight.

This made it straightforward to turn the local prototype into a publicly accessible project that judges and users can try immediately.

Architecture

User
 |
 v
TrailTutor web interface
 |
 v
FastAPI application on Render
 |
 +----------------------------+
 |                            |
 v                            v
DiffusionGemma             Tinker experiment
via NVIDIA API             Qwen3.5-4B
                              |
                              +--> untuned baseline
                              |
                              +--> LoRA-tuned model
Enter fullscreen mode Exit fullscreen mode

Why open innovation matters here

TrailTutor benefits from open models because I can do more than simply send text to an opaque endpoint.

I can:

  • inspect and compare model behavior
  • fine-tune a model for a narrow task
  • change the serving provider
  • run experiments against the same base model
  • preserve training and evaluation evidence
  • potentially self-host the model later

The Tinker experiment is a good example.

Instead of claiming that fine-tuning helped, I could actually train an open-weight model, evaluate it against its own untuned baseline, inspect failures, discover a bug in my evaluator, and publish the corrected evidence.

That kind of experimentation is much easier when the model ecosystem is open enough to adapt.

What surprised me

The biggest surprise was not the training result.

It was the evaluation bug.

Several outputs I initially counted as complete failures were actually good TrailTutor missions repeated twice.

That reminded me that model evaluation is not only about testing the model. The evaluator also needs to be tested.

It changed the final result from an apparent 20-point improvement to a more defensible 5-point improvement.

I prefer the smaller number because I can explain exactly where it came from.

What I would build next

I would like to continue in three directions:

  1. Add a mobile-friendly reflection mode so learners can return after the activity and record what they observed.
  2. Expand the training set with more science, geography, and environmental topics.
  3. Test local/offline inference so TrailTutor can work in places without a reliable internet connection.

Prize Categories

I am entering TrailTutor AI for:

  • Best Use of Gemma
  • Best Use of Tinker
  • Best Use of Render

It is also automatically eligible for the overall Touch Grass challenge.

Final thought

Many AI applications are designed to increase engagement time.

TrailTutor deliberately tries to do the opposite.

It uses AI to give a learner one useful reason to close the screen, step outside, pay attention to the physical world, and come back with something they noticed for themselves.

Top comments (0)