This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
DrishtiPath: Offline AI Navigation Through What You See
GPS is great — until you're in a crowded event, surrounded by temporary roads, changing entry points, thousands of people, and unreliable connectivity.
That made me think about a different way to navigate:
What if your phone could understand where you are simply by looking at what's around you?
That idea became DrishtiPath.
Instead of keeping users staring at a map, DrishtiPath is designed around a short interaction:
Look → Capture → Identify → Get direction → Keep walking.
What I Built
DrishtiPath is an AI-assisted visual navigation system for large outdoor gatherings such as Kumbh Mela.
A user can take a photo of a nearby landmark, gate, ghat, help booth, temple area, or other recognizable location.
An open-weight AI model analyzes the image and attempts to identify the location.
Once the landmark is recognized, DrishtiPath can show useful nearby places such as:
- Medical camps
- Police help booths
- Water points
- Information centres
- Parking areas
- Important gates
- Ghats and temple areas
Instead of spending several minutes staring at a map, the interaction is intended to take only a few seconds.
The basic flow
See landmark
↓
Take photo
↓
Local Gemma model
↓
Identify possible landmark
↓
Match with known location
↓
Find nearby facilities
↓
Show direction
↓
Put phone away and continue walking
For the prototype, I focused on Nashik–Trimbakeshwar Kumbh and created a small landmark dataset containing places such as:
- Ramkund Ghat
- Panchavati Entry Gate
- Tapovan
- Kalaram Mandir area
- Sadhugram Entry
- Trimbakeshwar Temple area
- Information Centre
- Medical camps
- Police help booths
- Parking areas
- Water points
The interface also lets the user switch between:
- Prayagraj Kumbh
- Haridwar Kumbh
- Ujjain Simhastha
- Nashik–Trimbakeshwar Kumbh
The current project is an MVP, but the idea is designed around a real problem: helping people navigate environments where GPS coordinates alone may not tell the whole story.
Demo
How the demo works
- Select a Kumbh location.
- Upload or capture an image of a nearby landmark.
- DrishtiPath sends the image to the locally running Gemma model.
- Gemma analyzes the visual scene.
- The response is matched against known landmarks.
- DrishtiPath displays the detected location.
- Nearby useful facilities are calculated.
- The user gets simple navigation information.
Example:
Detected Landmark:
Ramkund Ghat
Nearby:
Medical Camp 120 m
Police Help Booth 180 m
Water Point 90 m
The goal isn't to make users continuously follow a blue dot.
The goal is to answer:
"Where am I, and where should I go next?"
…and then let them put the phone back in their pocket.
Demo
Live Demo:
https://drishtipath.onrender.com/
Code
The project is open source on GitHub:
https://github.com/dev-yashpawar/DrishtiPath
The repository contains the web application, backend integration, landmark data, and local AI integration used by the prototype.
How I Built It
The core of DrishtiPath is not a cloud AI API.
It uses an open-weight Gemma model running locally with Ollama.
For my local development setup I used:
OLLAMA_URL=http://127.0.0.1:11434
OLLAMA_MODEL=gemma4:e2b
Architecture
┌──────────────────┐
│ User │
│ Camera / Upload │
└────────┬─────────┘
│
▼
┌──────────────────┐
│ Web Client │
│ Map + Navigation │
└────────┬─────────┘
│
▼
┌──────────────────┐
│ Backend API │
│ Image Processing │
└────────┬─────────┘
│
▼
┌──────────────────┐
│ Ollama │
│ Local Inference │
└────────┬─────────┘
│
▼
┌──────────────────┐
│ Gemma 4 │
│ Visual Analysis │
└────────┬─────────┘
│
▼
┌──────────────────┐
│ Landmark Matcher │
│ + Facility Data │
└────────┬─────────┘
│
▼
┌──────────────────┐
│ Navigation Result│
└──────────────────┘
Why visual navigation?
Traditional navigation usually begins with coordinates.
DrishtiPath adds another signal:
What can the user actually see?
If someone is standing near a recognizable gate, temple, ghat, booth, or entrance, that visual information can help narrow down their location.
Long term, I want to combine:
GPS
+
Visual landmark recognition
+
Known event landmarks
+
Nearby infrastructure
=
More reliable navigation
This is especially interesting for temporary event environments because roads, barricades, sectors and entrances can change.
Local AI with Ollama
Ollama made it possible to run the model directly on my development machine.
Instead of doing:
Image
↓
Internet
↓
External AI company
↓
Remote inference
↓
Response
DrishtiPath can use:
Image
↓
Local Ollama server
↓
Gemma
↓
Response
That changes more than just the architecture.
It changes what kind of product can be built.
Landmark Matching
I don't want the AI model to invent random place names.
The model's result is therefore used as part of a controlled navigation pipeline.
Conceptually:
const detectedLandmark = await analyzeImage(image);
const landmark = landmarks.find(
place => matchesDetectedLocation(place, detectedLandmark)
);
const nearbyFacilities = findNearbyFacilities(
landmark.latitude,
landmark.longitude
);
Known locations contain information such as:
{
"name": "Ramkund Ghat",
"type": "ghat",
"latitude": 20.007,
"longitude": 73.789,
"city": "Nashik",
"event": "Nashik-Trimbakeshwar Kumbh"
}
This allows the application to separate two jobs:
AI
Understand the image.
Application logic
Decide which known landmark and facilities should be shown.
That separation is important because an AI model should not be trusted as the database of truth for emergency navigation.
Maps and nearby facilities
After identifying a landmark, the application can calculate nearby facilities using coordinates.
For two points:
Current landmark
↓
Calculate distance
↓
Known medical / police / water locations
↓
Sort nearest first
This makes the AI useful rather than simply producing a paragraph describing the photograph.
The result becomes an actual action:
Medical assistance is 120 metres in that direction.
What Is Offline Today?
I wanted to be careful with the word offline.
The AI inference can run locally through Ollama without sending every image to a proprietary cloud AI service.
The current MVP still uses web technologies and some online infrastructure for things such as development data and map resources.
A completely offline production version would additionally require:
- Downloadable offline map packages
- Locally stored landmark databases
- Cached facility information
- On-device or local-network inference
- Sync when connectivity returns
So the current project demonstrates the most important part first:
local AI landmark understanding.
The architecture can then progressively remove the remaining network dependencies.
Why Does Open Innovation Matter?
This was the most interesting part of the project for me.
I could have connected DrishtiPath to a closed vision API.
Technically, that would probably have been easier.
But it would weaken the exact scenario the project is designed for.
1. Connectivity should not decide whether the AI works
Large gatherings can put enormous pressure on mobile networks.
A navigation assistant that completely depends on sending every image to a cloud model becomes least reliable when the environment becomes most difficult.
With an open-weight model, inference can happen much closer to the user.
2. Images don't have to leave the user's environment
Navigation photographs can contain:
- Faces
- Families
- Vehicles
- Surroundings
- Location clues
For many AI applications, uploading that data to another company's server is accepted as normal.
It doesn't have to be.
Local inference gives developers another architectural option.
3. The model can be replaced
DrishtiPath isn't supposed to permanently depend on one vendor.
The AI layer can eventually be changed from:
Gemma
to:
Fine-tuned Gemma
or another capable open-weight vision model.
The rest of the navigation system does not need to be rebuilt from scratch.
That flexibility matters for open projects.
4. Specialized models become possible
A general-purpose AI model knows about the world.
But a Kumbh navigation model doesn't need to know everything about the world.
It needs to be exceptionally good at recognizing things like:
Ghats
Entry gates
Temporary sectors
Help booths
Temple surroundings
Parking signs
Information boards
Medical camps
With open models, a future version could be optimized or fine-tuned specifically for this environment.
That is much harder when the model is a completely closed black box.
5. AI should disappear into the experience
One thing I deliberately did not want to build was another chatbot.
The user doesn't need to have a ten-message conversation with an AI while trying to find a medical camp.
They should be able to:
Point
Tap
Understand
Walk
The AI is not the interface.
The AI is infrastructure underneath the experience.
That is probably my favorite part of DrishtiPath.
Challenges I Ran Into
Running multimodal models locally also reminded me why cloud AI became popular.
Local AI has real constraints.
I ran into issues including:
- Model size
- RAM limitations
- Slow inference on weaker hardware
- Ollama configuration
- Backend-to-Ollama connectivity
- Getting predictable structured model responses
- Distinguishing visually similar locations
For example, asking the model:
Where is this?
isn't good enough.
A much better approach is to give it context and constrain its job:
You are identifying landmarks inside the
Nashik-Trimbakeshwar Kumbh navigation system.
Analyze the image and return the most likely
landmark from the provided candidate locations.
Return structured JSON only.
That reduced the amount of uncontrolled output and made the response easier for the application to process.
It also taught me something important:
Using an AI model is easy. Building a dependable system around an AI model is the difficult part.
What I Would Build Next
DrishtiPath is still a prototype.
The next version would focus on:
GPS + visual fusion
Use GPS to narrow the candidates before asking the vision model.
Instead of checking 500 landmarks:
GPS radius: 500 metres
Possible landmarks:
- Ramkund
- Gate A
- Police Booth 1
↓
Gemma chooses from only these candidates.
That should make visual recognition considerably more reliable.
Fully offline maps
Download an event area's map before visiting.
Voice navigation
Instead of reading:
Police help booth is 180 metres northeast.
the user could hear it.
Multilingual navigation
For Kumbh deployments, I would prioritize:
- Marathi
- Hindi
- English
Emergency mode
One tap could prioritize:
- Police
- Medical assistance
- Lost-person centres
- Emergency exits
Volunteer mode
Event volunteers could receive additional operational information unavailable in normal visitor mode.
Real-time admin updates
Temporary roads and facilities can change quickly during large events.
An admin system could update those locations, which devices would synchronize whenever connectivity becomes available.
What "Touch Grass" Means for This Project
A lot of software is designed to increase screen time.
DrishtiPath tries to do the opposite.
The ideal successful session is probably less than ten seconds:
Open DrishtiPath
Take a photo
"Ramkund Ghat detected"
"Medical Camp: 120 m →"
Close phone
Walk
If someone has to continuously stare at DrishtiPath while walking, I haven't built the experience correctly.
The screen should be the shortest part of the journey.
The real product is everything happening after the screen turns off.
Prize Categories
I'm entering DrishtiPath for:
Best Use of Gemma 4
Gemma is not an optional chatbot added on top of the application. Local visual understanding is one of the core mechanisms that makes DrishtiPath possible.
Best Open-Source AI Project
The project is built around the idea that open-weight AI can enable local, private and adaptable intelligence for real-world navigation.
Final Thoughts
Hacktoberfest usually makes me think about code.
This challenge made me think more about where the code runs and what happens when the user closes the screen.
DrishtiPath started with a simple question:
Could AI navigation understand what I see instead of only knowing my coordinates?
The prototype isn't a replacement for GPS.
And it isn't pretending AI vision will perfectly recognize every landmark.
The interesting opportunity is combining several imperfect signals:
Visual understanding
+
GPS
+
Known landmarks
+
Local event data
+
Human-friendly directions
to create something more useful than any one of them individually.
And because the intelligence can run using an open-weight model, that system doesn't necessarily have to depend on a distant server to understand the world directly in front of it.
That is the direction I want to keep exploring.
Less screen. More surroundings.
That's DrishtiPath.
Top comments (2)
AI Disclosure: I used ChatGPT to help organize, refine, and improve the wording of this article. The project idea, implementation, testing, code, technical decisions, and final verification are my own.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.