I Built a Presentation Remote That Uses Hand Gestures and Voice
This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend.
What I Built
My friend Dinesh Prajapath presents frequently. His decks are often large PDFs, and halfway through a talk he can lose track of where something is:
"The pricing part was around page 12... no, 14?"
Scrolling through a PDF in front of an audience is stressful. Traditional clickers make sequential navigation easy, but jumping directly to a specific slide usually still requires knowing its position.
So I built a presentation remote controlled using a laptop webcam and microphone.
The goal was simple:
Let the presenter navigate a presentation without touching the laptop.
What It Does
Hand Gestures
The webcam tracks the presenter's hand using MediaPipe.
- Swipe right/left → change slides
- Open palm → blank the presentation screen
- Fist → jump to the first slide
No special hardware or physical clicker is required.
Voice Navigation
The presenter can hold a key and say something like:
"Go to the pricing slide."
Instead of requiring the presenter to remember the page number, the system searches the content of the presentation and finds the most relevant slide.
It can also handle commands such as:
"page 7""first slide""last slide""go to pricing""show weather GPT"
Why Open Models Mattered
Everything runs locally on the laptop:
- MediaPipe → hand tracking
- Whisper → speech-to-text
- Gemma via Ollama → semantic understanding and slide selection
This was important for several reasons.
Privacy
The presentation, webcam feed, and microphone audio are processed locally and are not sent to a cloud service.
That matters when presenting slides that are confidential or haven't been published yet.
Cost
There is no per-request API cost.
Once the required models are downloaded, there is no cloud inference bill for each presentation.
Offline
After the initial model downloads, the system can operate without an internet connection.
That's useful in classrooms, seminar halls, and event venues where Wi-Fi can be unreliable.
Control
I could change the models, prompts, matching logic, and decision-making process myself.
This ended up being one of the biggest advantages of using an open/local approach.
Prize Category
I'm entering the Best Use of Gemma category because Gemma runs locally through Ollama and is used for semantic slide selection when deterministic matching is insufficient.
Instead of sending the presentation content to a cloud API, Gemma helps the system understand ambiguous voice commands and identify the most relevant slide locally.
Gemma is not responsible for controlling the presentation directly. It acts as the final semantic decision-maker when deterministic matching cannot confidently choose a slide.
The Problem With My First Approach
My first version was much simpler.
I extracted the text from every slide, gave all of it to Gemma 2B, and asked it to return the page number matching the user's request.
It didn't work reliably.
For example:
-
"intelligent record system"→ page 14 instead of page 5 -
"WeatherGPT"→ page 15 instead of page 18
One problem was that some slides contained numbers such as:
04 Land Records
Those numbers were section labels printed on the slides, not the actual PDF page numbers.
The small language model could confuse those numbers with the real slide positions.
So instead of making the LLM responsible for everything, I changed the architecture.
The Hybrid Approach
The final system uses multiple levels of matching.
1. Direct Commands
Simple commands don't need an LLM.
For example:
page 7
first slide
last slide
next slide
previous slide
These are handled directly by the application.
This makes them faster and more reliable.
2. Keyword + Fuzzy Matching
For content-based requests, the system first searches the extracted slide text.
This handles obvious matches without involving the LLM.
It also helps with small differences caused by speech recognition.
For example:
Weather GPT
can still match:
WeatherGPT
3. Gemma as a Final Decision Maker
If the keyword matching produces several plausible slides, Gemma receives the remaining candidates and decides which one best matches the request.
The normal path limits this to the top five candidates.
If keyword matching cannot produce useful candidates, the system can fall back to semantic matching against the slide titles.
Gemma can also return:
none
when none of the candidates are relevant.
In that case, the presentation does nothing instead of jumping to a potentially incorrect slide.
That conservative behavior is intentional.
How Voice Recognition Works
Whisper runs locally, so the microphone audio does not need to be uploaded to a speech API.
I also give Whisper the presentation's slide titles as vocabulary hints.
For example, if the deck contains terms such as:
Land Records
WeatherGPT
Intelligent Record System
Digital Governance
those terms can be provided as context to improve recognition.
I also filter out Whisper segments that are likely to be silence or low-confidence/noise.
This helps prevent silence or background noise from accidentally becoming a navigation command.
How It Works
Gesture Pipeline
Webcam
↓
MediaPipe Hand Landmarks
↓
Gesture Rules
↓
Keyboard / Presentation Action
Voice Pipeline
Microphone
↓
Whisper
↓
Command Detection
↓
Keyword + Fuzzy Matching
↓
Gemma / Ollama when needed
↓
Target Slide
Overall Architecture
┌─────────────────────┐
│ Presentation │
│ PDF/PPTX │
└──────────┬──────────┘
│
Extract Text
│
▼
┌─────────────────────┐
│ Slide Indexing │
└─────────────────────┘
Webcam ──→ MediaPipe ──→ Gesture Rules ──→ Slide Action
Microphone ──→ Whisper ──→ Keyword/Fuzzy Search
│
▼
Candidate Slides
│
▼
Gemma / Ollama
│
▼
Target Slide
What I Learned
The biggest lesson from this project was that using an LLM for everything isn't necessarily the best architecture.
My first instinct was:
User Request
↓
LLM
↓
Slide Number
But the system became much more reliable when I divided the problem into smaller parts:
Simple request
↓
Deterministic logic
Clear content match
↓
Keyword + fuzzy matching
Ambiguous request
↓
Small candidate set
↓
LLM decision
This reduced the amount of reasoning the model had to perform and also made failures safer.
The LLM became a decision-making component rather than the entire navigation system.
Why I Chose This Design
A presentation remote has an unusual requirement:
Being wrong is worse than doing nothing.
If the system fails to find a slide, the presenter can try again.
But if it confidently jumps to the wrong slide during a live presentation, it creates confusion.
So I designed the system to prefer:
No action
over:
Probably the right slide
when confidence is low.
Limitations
The project is still a prototype, so there are several limitations.
Gesture Detection
The gesture thresholds are tuned for a particular webcam setup.
Lighting, camera position, hand visibility, and background can affect detection accuracy.
Speech Recognition
Whisper can sometimes misrecognise accented or noisy speech.
A larger model such as small can improve recognition, but requires more computational resources.
Text-Based Navigation
Voice navigation depends on text extracted from the presentation.
Image-only or scanned PDFs may not contain useful extractable text.
Slides with very little text are also harder to locate using voice commands.
Slide Synchronization
The remote maintains its own slide position.
If the presenter manually changes slides using the mouse or keyboard, the internal slide counter can potentially drift from the actual presentation.
Local Model Resources
Larger language models can improve semantic matching, but they require more RAM and CPU/GPU resources.
The project therefore involves a trade-off between:
Model Size
↕
Accuracy
↕
Local Hardware Requirements
What I'd Improve Next
There are several directions I would like to explore:
- Better automatic gesture calibration
- More robust hand tracking across different lighting conditions
- OCR support for image-based/scanned slides
- Better synchronization with browser-based PDF viewers
- More advanced semantic slide indexing
- Support for multiple presentation applications
- Confidence scores and visual feedback before changing slides
- More natural voice commands such as:
"Go back to the slide about pricing""Show the architecture section""Find the slide mentioning MongoDB"
Demo
Code
The complete source code is available on GitHub:
The project includes the gesture recognition, voice navigation, PDF/PPTX slide extraction, fuzzy matching, and local Gemma/Ollama integration.
Final Thoughts
I started this project because a friend kept running into a very practical problem while presenting.
What looked like a simple "control the slides with gestures and voice" project turned into a useful lesson in system design.
The most important improvement wasn't choosing a bigger model.
It was reducing what the model had to solve.
Instead of asking an LLM to understand the entire presentation and return a page number, I combined deterministic rules, fuzzy search, and an LLM only where it actually added value.
That made the system faster, more predictable, and safer to use during a live presentation.
And most importantly, Dinesh no longer has to remember:
"Was pricing on page 12 or 14?"
Top comments (0)