DEV Community

Cover image for DrishtiPath: Offline AI Navigation Through What You See
Yash Pawar
Yash Pawar

Posted on

DrishtiPath: Offline AI Navigation Through What You See

Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass Submission 🌿

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass

DrishtiPath: Offline AI Navigation Through What You See

GPS is great — until you're in a crowded event, surrounded by temporary roads, changing entry points, thousands of people, and unreliable connectivity.

That made me think about a different way to navigate:

What if your phone could understand where you are simply by looking at what's around you?

That idea became DrishtiPath.

Instead of keeping users staring at a map, DrishtiPath is designed around a short interaction:

Look → Capture → Identify → Get direction → Keep walking.


What I Built

DrishtiPath is an AI-assisted visual navigation system for large outdoor gatherings such as Kumbh Mela.

A user can take a photo of a nearby landmark, gate, ghat, help booth, temple area, or other recognizable location.

An open-weight AI model analyzes the image and attempts to identify the location.

Once the landmark is recognized, DrishtiPath can show useful nearby places such as:

  • Medical camps
  • Police help booths
  • Water points
  • Information centres
  • Parking areas
  • Important gates
  • Ghats and temple areas

Instead of spending several minutes staring at a map, the interaction is intended to take only a few seconds.

The basic flow

See landmark
     ↓
Take photo
     ↓
Local Gemma model
     ↓
Identify possible landmark
     ↓
Match with known location
     ↓
Find nearby facilities
     ↓
Show direction
     ↓
Put phone away and continue walking
Enter fullscreen mode Exit fullscreen mode

For the prototype, I focused on Nashik–Trimbakeshwar Kumbh and created a small landmark dataset containing places such as:

  • Ramkund Ghat
  • Panchavati Entry Gate
  • Tapovan
  • Kalaram Mandir area
  • Sadhugram Entry
  • Trimbakeshwar Temple area
  • Information Centre
  • Medical camps
  • Police help booths
  • Parking areas
  • Water points

The interface also lets the user switch between:

  • Prayagraj Kumbh
  • Haridwar Kumbh
  • Ujjain Simhastha
  • Nashik–Trimbakeshwar Kumbh

The current project is an MVP, but the idea is designed around a real problem: helping people navigate environments where GPS coordinates alone may not tell the whole story.


Demo

How the demo works

  1. Select a Kumbh location.
  2. Upload or capture an image of a nearby landmark.
  3. DrishtiPath sends the image to the locally running Gemma model.
  4. Gemma analyzes the visual scene.
  5. The response is matched against known landmarks.
  6. DrishtiPath displays the detected location.
  7. Nearby useful facilities are calculated.
  8. The user gets simple navigation information.

Example:

Detected Landmark:
Ramkund Ghat

Nearby:

Medical Camp        120 m
Police Help Booth   180 m
Water Point          90 m
Enter fullscreen mode Exit fullscreen mode

The goal isn't to make users continuously follow a blue dot.

The goal is to answer:

"Where am I, and where should I go next?"

…and then let them put the phone back in their pocket.

Demo

Live Demo:

https://drishtipath.onrender.com/


Code

The project is open source on GitHub:

https://github.com/dev-yashpawar/DrishtiPath

The repository contains the web application, backend integration, landmark data, and local AI integration used by the prototype.


How I Built It

The core of DrishtiPath is not a cloud AI API.

It uses an open-weight Gemma model running locally with Ollama.

For my local development setup I used:

OLLAMA_URL=http://127.0.0.1:11434
OLLAMA_MODEL=gemma4:e2b
Enter fullscreen mode Exit fullscreen mode

Architecture

                ┌──────────────────┐
                │      User        │
                │ Camera / Upload  │
                └────────┬─────────┘
                         │
                         ▼
                ┌──────────────────┐
                │   Web Client     │
                │ Map + Navigation │
                └────────┬─────────┘
                         │
                         ▼
                ┌──────────────────┐
                │ Backend API      │
                │ Image Processing │
                └────────┬─────────┘
                         │
                         ▼
                ┌──────────────────┐
                │      Ollama      │
                │  Local Inference │
                └────────┬─────────┘
                         │
                         ▼
                ┌──────────────────┐
                │     Gemma 4      │
                │ Visual Analysis  │
                └────────┬─────────┘
                         │
                         ▼
                ┌──────────────────┐
                │ Landmark Matcher │
                │ + Facility Data  │
                └────────┬─────────┘
                         │
                         ▼
                ┌──────────────────┐
                │ Navigation Result│
                └──────────────────┘
Enter fullscreen mode Exit fullscreen mode

Why visual navigation?

Traditional navigation usually begins with coordinates.

DrishtiPath adds another signal:

What can the user actually see?

If someone is standing near a recognizable gate, temple, ghat, booth, or entrance, that visual information can help narrow down their location.

Long term, I want to combine:

GPS
+
Visual landmark recognition
+
Known event landmarks
+
Nearby infrastructure
=
More reliable navigation
Enter fullscreen mode Exit fullscreen mode

This is especially interesting for temporary event environments because roads, barricades, sectors and entrances can change.


Local AI with Ollama

Ollama made it possible to run the model directly on my development machine.

Instead of doing:

Image
 ↓
Internet
 ↓
External AI company
 ↓
Remote inference
 ↓
Response
Enter fullscreen mode Exit fullscreen mode

DrishtiPath can use:

Image
 ↓
Local Ollama server
 ↓
Gemma
 ↓
Response
Enter fullscreen mode Exit fullscreen mode

That changes more than just the architecture.

It changes what kind of product can be built.


Landmark Matching

I don't want the AI model to invent random place names.

The model's result is therefore used as part of a controlled navigation pipeline.

Conceptually:

const detectedLandmark = await analyzeImage(image);

const landmark = landmarks.find(
  place => matchesDetectedLocation(place, detectedLandmark)
);

const nearbyFacilities = findNearbyFacilities(
  landmark.latitude,
  landmark.longitude
);
Enter fullscreen mode Exit fullscreen mode

Known locations contain information such as:

{
  "name": "Ramkund Ghat",
  "type": "ghat",
  "latitude": 20.007,
  "longitude": 73.789,
  "city": "Nashik",
  "event": "Nashik-Trimbakeshwar Kumbh"
}
Enter fullscreen mode Exit fullscreen mode

This allows the application to separate two jobs:

AI

Understand the image.

Application logic

Decide which known landmark and facilities should be shown.

That separation is important because an AI model should not be trusted as the database of truth for emergency navigation.


Maps and nearby facilities

After identifying a landmark, the application can calculate nearby facilities using coordinates.

For two points:

Current landmark
        ↓
Calculate distance
        ↓
Known medical / police / water locations
        ↓
Sort nearest first
Enter fullscreen mode Exit fullscreen mode

This makes the AI useful rather than simply producing a paragraph describing the photograph.

The result becomes an actual action:

Medical assistance is 120 metres in that direction.


What Is Offline Today?

I wanted to be careful with the word offline.

The AI inference can run locally through Ollama without sending every image to a proprietary cloud AI service.

The current MVP still uses web technologies and some online infrastructure for things such as development data and map resources.

A completely offline production version would additionally require:

  • Downloadable offline map packages
  • Locally stored landmark databases
  • Cached facility information
  • On-device or local-network inference
  • Sync when connectivity returns

So the current project demonstrates the most important part first:

local AI landmark understanding.

The architecture can then progressively remove the remaining network dependencies.


Why Does Open Innovation Matter?

This was the most interesting part of the project for me.

I could have connected DrishtiPath to a closed vision API.

Technically, that would probably have been easier.

But it would weaken the exact scenario the project is designed for.

1. Connectivity should not decide whether the AI works

Large gatherings can put enormous pressure on mobile networks.

A navigation assistant that completely depends on sending every image to a cloud model becomes least reliable when the environment becomes most difficult.

With an open-weight model, inference can happen much closer to the user.


2. Images don't have to leave the user's environment

Navigation photographs can contain:

  • Faces
  • Families
  • Vehicles
  • Surroundings
  • Location clues

For many AI applications, uploading that data to another company's server is accepted as normal.

It doesn't have to be.

Local inference gives developers another architectural option.


3. The model can be replaced

DrishtiPath isn't supposed to permanently depend on one vendor.

The AI layer can eventually be changed from:

Gemma
Enter fullscreen mode Exit fullscreen mode

to:

Fine-tuned Gemma
Enter fullscreen mode Exit fullscreen mode

or another capable open-weight vision model.

The rest of the navigation system does not need to be rebuilt from scratch.

That flexibility matters for open projects.


4. Specialized models become possible

A general-purpose AI model knows about the world.

But a Kumbh navigation model doesn't need to know everything about the world.

It needs to be exceptionally good at recognizing things like:

Ghats
Entry gates
Temporary sectors
Help booths
Temple surroundings
Parking signs
Information boards
Medical camps
Enter fullscreen mode Exit fullscreen mode

With open models, a future version could be optimized or fine-tuned specifically for this environment.

That is much harder when the model is a completely closed black box.


5. AI should disappear into the experience

One thing I deliberately did not want to build was another chatbot.

The user doesn't need to have a ten-message conversation with an AI while trying to find a medical camp.

They should be able to:

Point
Tap
Understand
Walk
Enter fullscreen mode Exit fullscreen mode

The AI is not the interface.

The AI is infrastructure underneath the experience.

That is probably my favorite part of DrishtiPath.


Challenges I Ran Into

Running multimodal models locally also reminded me why cloud AI became popular.

Local AI has real constraints.

I ran into issues including:

  • Model size
  • RAM limitations
  • Slow inference on weaker hardware
  • Ollama configuration
  • Backend-to-Ollama connectivity
  • Getting predictable structured model responses
  • Distinguishing visually similar locations

For example, asking the model:

Where is this?
Enter fullscreen mode Exit fullscreen mode

isn't good enough.

A much better approach is to give it context and constrain its job:

You are identifying landmarks inside the
Nashik-Trimbakeshwar Kumbh navigation system.

Analyze the image and return the most likely
landmark from the provided candidate locations.

Return structured JSON only.
Enter fullscreen mode Exit fullscreen mode

That reduced the amount of uncontrolled output and made the response easier for the application to process.

It also taught me something important:

Using an AI model is easy. Building a dependable system around an AI model is the difficult part.


What I Would Build Next

DrishtiPath is still a prototype.

The next version would focus on:

GPS + visual fusion

Use GPS to narrow the candidates before asking the vision model.

Instead of checking 500 landmarks:

GPS radius: 500 metres

Possible landmarks:
- Ramkund
- Gate A
- Police Booth 1

        ↓

Gemma chooses from only these candidates.
Enter fullscreen mode Exit fullscreen mode

That should make visual recognition considerably more reliable.

Fully offline maps

Download an event area's map before visiting.

Voice navigation

Instead of reading:

Police help booth is 180 metres northeast.

the user could hear it.

Multilingual navigation

For Kumbh deployments, I would prioritize:

  • Marathi
  • Hindi
  • English

Emergency mode

One tap could prioritize:

  • Police
  • Medical assistance
  • Lost-person centres
  • Emergency exits

Volunteer mode

Event volunteers could receive additional operational information unavailable in normal visitor mode.

Real-time admin updates

Temporary roads and facilities can change quickly during large events.

An admin system could update those locations, which devices would synchronize whenever connectivity becomes available.


What "Touch Grass" Means for This Project

A lot of software is designed to increase screen time.

DrishtiPath tries to do the opposite.

The ideal successful session is probably less than ten seconds:

Open DrishtiPath

Take a photo

"Ramkund Ghat detected"

"Medical Camp: 120 m →"

Close phone

Walk
Enter fullscreen mode Exit fullscreen mode

If someone has to continuously stare at DrishtiPath while walking, I haven't built the experience correctly.

The screen should be the shortest part of the journey.

The real product is everything happening after the screen turns off.


Prize Categories

I'm entering DrishtiPath for:

Best Use of Gemma 4

Gemma is not an optional chatbot added on top of the application. Local visual understanding is one of the core mechanisms that makes DrishtiPath possible.

Best Open-Source AI Project

The project is built around the idea that open-weight AI can enable local, private and adaptable intelligence for real-world navigation.


Final Thoughts

Hacktoberfest usually makes me think about code.

This challenge made me think more about where the code runs and what happens when the user closes the screen.

DrishtiPath started with a simple question:

Could AI navigation understand what I see instead of only knowing my coordinates?

The prototype isn't a replacement for GPS.

And it isn't pretending AI vision will perfectly recognize every landmark.

The interesting opportunity is combining several imperfect signals:

Visual understanding
+
GPS
+
Known landmarks
+
Local event data
+
Human-friendly directions
Enter fullscreen mode Exit fullscreen mode

to create something more useful than any one of them individually.

And because the intelligence can run using an open-weight model, that system doesn't necessarily have to depend on a distant server to understand the world directly in front of it.

That is the direction I want to keep exploring.

Less screen. More surroundings.

That's DrishtiPath.

Top comments (2)

Collapse
 
yashpawar profile image
Yash Pawar •

AI Disclosure: I used ChatGPT to help organize, refine, and improve the wording of this article. The project idea, implementation, testing, code, technical decisions, and final verification are my own.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.