<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Aman Kumar</title>
    <description>The latest articles on DEV Community by Aman Kumar (@devidevilaldevil).</description>
    <link>https://dev.to/devidevilaldevil</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4160411%2Fa42a8ea6-3d01-4fde-8cfb-1671cff062ed.png</url>
      <title>DEV Community: Aman Kumar</title>
      <link>https://dev.to/devidevilaldevil</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/devidevilaldevil"/>
    <language>en</language>
    <item>
      <title>I built a hands-free slide remote for a friend who can never find the right page</title>
      <dc:creator>Aman Kumar</dc:creator>
      <pubDate>Sun, 04 Oct 2026 14:44:08 +0000</pubDate>
      <link>https://dev.to/devidevilaldevil/i-built-a-hands-free-slide-remote-for-a-friend-who-can-never-find-the-right-page-1298</link>
      <guid>https://dev.to/devidevilaldevil/i-built-a-hands-free-slide-remote-for-a-friend-who-can-never-find-the-right-page-1298</guid>
      <description>&lt;h1&gt;
  
  
  I Built a Presentation Remote That Uses Hand Gestures and Voice
&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-weekend-2026-10-01"&gt;Hacktoberfest Weekend Challenge: Build for a Friend&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;My friend &lt;strong&gt;Dinesh Prajapath&lt;/strong&gt; presents frequently. His decks are often large PDFs, and halfway through a talk he can lose track of where something is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"The pricing part was around page 12... no, 14?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Scrolling through a PDF in front of an audience is stressful. Traditional clickers make sequential navigation easy, but jumping directly to a specific slide usually still requires knowing its position.&lt;/p&gt;

&lt;p&gt;So I built a &lt;strong&gt;presentation remote controlled using a laptop webcam and microphone&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The goal was simple:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Let the presenter navigate a presentation without touching the laptop.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What It Does
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Hand Gestures
&lt;/h3&gt;

&lt;p&gt;The webcam tracks the presenter's hand using MediaPipe.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Swipe right/left&lt;/strong&gt; → change slides&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open palm&lt;/strong&gt; → blank the presentation screen&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fist&lt;/strong&gt; → jump to the first slide&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;No special hardware or physical clicker is required.&lt;/p&gt;

&lt;h3&gt;
  
  
  Voice Navigation
&lt;/h3&gt;

&lt;p&gt;The presenter can hold a key and say something like:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Go to the pricing slide."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Instead of requiring the presenter to remember the page number, the system searches the content of the presentation and finds the most relevant slide.&lt;/p&gt;

&lt;p&gt;It can also handle commands such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;"page 7"&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;"first slide"&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;"last slide"&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;"go to pricing"&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;"show weather GPT"&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why Open Models Mattered
&lt;/h2&gt;

&lt;p&gt;Everything runs locally on the laptop:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;MediaPipe&lt;/strong&gt; → hand tracking&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Whisper&lt;/strong&gt; → speech-to-text&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemma via Ollama&lt;/strong&gt; → semantic understanding and slide selection&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This was important for several reasons.&lt;/p&gt;

&lt;h3&gt;
  
  
  Privacy
&lt;/h3&gt;

&lt;p&gt;The presentation, webcam feed, and microphone audio are processed locally and are not sent to a cloud service.&lt;/p&gt;

&lt;p&gt;That matters when presenting slides that are confidential or haven't been published yet.&lt;/p&gt;

&lt;h3&gt;
  
  
  Cost
&lt;/h3&gt;

&lt;p&gt;There is no per-request API cost.&lt;/p&gt;

&lt;p&gt;Once the required models are downloaded, there is no cloud inference bill for each presentation.&lt;/p&gt;

&lt;h3&gt;
  
  
  Offline
&lt;/h3&gt;

&lt;p&gt;After the initial model downloads, the system can operate without an internet connection.&lt;/p&gt;

&lt;p&gt;That's useful in classrooms, seminar halls, and event venues where Wi-Fi can be unreliable.&lt;/p&gt;

&lt;h3&gt;
  
  
  Control
&lt;/h3&gt;

&lt;p&gt;I could change the models, prompts, matching logic, and decision-making process myself.&lt;/p&gt;

&lt;p&gt;This ended up being one of the biggest advantages of using an open/local approach.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prize Category
&lt;/h2&gt;

&lt;p&gt;I'm entering the &lt;strong&gt;Best Use of Gemma&lt;/strong&gt; category because Gemma runs locally through Ollama and is used for semantic slide selection when deterministic matching is insufficient.&lt;/p&gt;

&lt;p&gt;Instead of sending the presentation content to a cloud API, Gemma helps the system understand ambiguous voice commands and identify the most relevant slide locally.&lt;/p&gt;

&lt;p&gt;Gemma is not responsible for controlling the presentation directly. It acts as the final semantic decision-maker when deterministic matching cannot confidently choose a slide.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem With My First Approach
&lt;/h2&gt;

&lt;p&gt;My first version was much simpler.&lt;/p&gt;

&lt;p&gt;I extracted the text from every slide, gave all of it to &lt;strong&gt;Gemma 2B&lt;/strong&gt;, and asked it to return the page number matching the user's request.&lt;/p&gt;

&lt;p&gt;It didn't work reliably.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;"intelligent record system"&lt;/code&gt; → page 14 instead of page 5&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;"WeatherGPT"&lt;/code&gt; → page 15 instead of page 18&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One problem was that some slides contained numbers such as:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;04 Land Records&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Those numbers were section labels printed on the slides, not the actual PDF page numbers.&lt;/p&gt;

&lt;p&gt;The small language model could confuse those numbers with the real slide positions.&lt;/p&gt;

&lt;p&gt;So instead of making the LLM responsible for everything, I changed the architecture.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Hybrid Approach
&lt;/h2&gt;

&lt;p&gt;The final system uses multiple levels of matching.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Direct Commands
&lt;/h3&gt;

&lt;p&gt;Simple commands don't need an LLM.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;page 7
first slide
last slide
next slide
previous slide
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These are handled directly by the application.&lt;/p&gt;

&lt;p&gt;This makes them faster and more reliable.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Keyword + Fuzzy Matching
&lt;/h3&gt;

&lt;p&gt;For content-based requests, the system first searches the extracted slide text.&lt;/p&gt;

&lt;p&gt;This handles obvious matches without involving the LLM.&lt;/p&gt;

&lt;p&gt;It also helps with small differences caused by speech recognition.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Weather GPT
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;can still match:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;WeatherGPT
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  3. Gemma as a Final Decision Maker
&lt;/h3&gt;

&lt;p&gt;If the keyword matching produces several plausible slides, Gemma receives the remaining candidates and decides which one best matches the request.&lt;/p&gt;

&lt;p&gt;The normal path limits this to the &lt;strong&gt;top five candidates&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;If keyword matching cannot produce useful candidates, the system can fall back to semantic matching against the slide titles.&lt;/p&gt;

&lt;p&gt;Gemma can also return:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;none
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;when none of the candidates are relevant.&lt;/p&gt;

&lt;p&gt;In that case, the presentation does nothing instead of jumping to a potentially incorrect slide.&lt;/p&gt;

&lt;p&gt;That conservative behavior is intentional.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Voice Recognition Works
&lt;/h2&gt;

&lt;p&gt;Whisper runs locally, so the microphone audio does not need to be uploaded to a speech API.&lt;/p&gt;

&lt;p&gt;I also give Whisper the presentation's slide titles as vocabulary hints.&lt;/p&gt;

&lt;p&gt;For example, if the deck contains terms such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Land Records
WeatherGPT
Intelligent Record System
Digital Governance
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;those terms can be provided as context to improve recognition.&lt;/p&gt;

&lt;p&gt;I also filter out Whisper segments that are likely to be silence or low-confidence/noise.&lt;/p&gt;

&lt;p&gt;This helps prevent silence or background noise from accidentally becoming a navigation command.&lt;/p&gt;

&lt;h2&gt;
  
  
  How It Works
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Gesture Pipeline
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Webcam
   ↓
MediaPipe Hand Landmarks
   ↓
Gesture Rules
   ↓
Keyboard / Presentation Action
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Voice Pipeline
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Microphone
   ↓
Whisper
   ↓
Command Detection
   ↓
Keyword + Fuzzy Matching
   ↓
Gemma / Ollama when needed
   ↓
Target Slide
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Overall Architecture
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                     ┌─────────────────────┐
                     │     Presentation    │
                     │       PDF/PPTX      │
                     └──────────┬──────────┘
                                │
                           Extract Text
                                │
                                ▼
                     ┌─────────────────────┐
                     │    Slide Indexing   │
                     └─────────────────────┘


 Webcam ──→ MediaPipe ──→ Gesture Rules ──→ Slide Action

 Microphone ──→ Whisper ──→ Keyword/Fuzzy Search
                              │
                              ▼
                       Candidate Slides
                              │
                              ▼
                         Gemma / Ollama
                              │
                              ▼
                         Target Slide
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What I Learned
&lt;/h2&gt;

&lt;p&gt;The biggest lesson from this project was that &lt;strong&gt;using an LLM for everything isn't necessarily the best architecture&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;My first instinct was:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User Request
     ↓
    LLM
     ↓
Slide Number
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But the system became much more reliable when I divided the problem into smaller parts:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Simple request
     ↓
Deterministic logic

Clear content match
     ↓
Keyword + fuzzy matching

Ambiguous request
     ↓
Small candidate set
     ↓
LLM decision
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This reduced the amount of reasoning the model had to perform and also made failures safer.&lt;/p&gt;

&lt;p&gt;The LLM became a &lt;strong&gt;decision-making component rather than the entire navigation system&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why I Chose This Design
&lt;/h2&gt;

&lt;p&gt;A presentation remote has an unusual requirement:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Being wrong is worse than doing nothing.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If the system fails to find a slide, the presenter can try again.&lt;/p&gt;

&lt;p&gt;But if it confidently jumps to the wrong slide during a live presentation, it creates confusion.&lt;/p&gt;

&lt;p&gt;So I designed the system to prefer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;No action
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;over:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Probably the right slide
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;when confidence is low.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limitations
&lt;/h2&gt;

&lt;p&gt;The project is still a prototype, so there are several limitations.&lt;/p&gt;

&lt;h3&gt;
  
  
  Gesture Detection
&lt;/h3&gt;

&lt;p&gt;The gesture thresholds are tuned for a particular webcam setup.&lt;/p&gt;

&lt;p&gt;Lighting, camera position, hand visibility, and background can affect detection accuracy.&lt;/p&gt;

&lt;h3&gt;
  
  
  Speech Recognition
&lt;/h3&gt;

&lt;p&gt;Whisper can sometimes misrecognise accented or noisy speech.&lt;/p&gt;

&lt;p&gt;A larger model such as &lt;code&gt;small&lt;/code&gt; can improve recognition, but requires more computational resources.&lt;/p&gt;

&lt;h3&gt;
  
  
  Text-Based Navigation
&lt;/h3&gt;

&lt;p&gt;Voice navigation depends on text extracted from the presentation.&lt;/p&gt;

&lt;p&gt;Image-only or scanned PDFs may not contain useful extractable text.&lt;/p&gt;

&lt;p&gt;Slides with very little text are also harder to locate using voice commands.&lt;/p&gt;

&lt;h3&gt;
  
  
  Slide Synchronization
&lt;/h3&gt;

&lt;p&gt;The remote maintains its own slide position.&lt;/p&gt;

&lt;p&gt;If the presenter manually changes slides using the mouse or keyboard, the internal slide counter can potentially drift from the actual presentation.&lt;/p&gt;

&lt;h3&gt;
  
  
  Local Model Resources
&lt;/h3&gt;

&lt;p&gt;Larger language models can improve semantic matching, but they require more RAM and CPU/GPU resources.&lt;/p&gt;

&lt;p&gt;The project therefore involves a trade-off between:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Model Size
    ↕
Accuracy
    ↕
Local Hardware Requirements
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What I'd Improve Next
&lt;/h2&gt;

&lt;p&gt;There are several directions I would like to explore:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Better automatic gesture calibration&lt;/li&gt;
&lt;li&gt;More robust hand tracking across different lighting conditions&lt;/li&gt;
&lt;li&gt;OCR support for image-based/scanned slides&lt;/li&gt;
&lt;li&gt;Better synchronization with browser-based PDF viewers&lt;/li&gt;
&lt;li&gt;More advanced semantic slide indexing&lt;/li&gt;
&lt;li&gt;Support for multiple presentation applications&lt;/li&gt;
&lt;li&gt;Confidence scores and visual feedback before changing slides&lt;/li&gt;
&lt;li&gt;More natural voice commands such as:

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;"Go back to the slide about pricing"&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;"Show the architecture section"&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;"Find the slide mentioning MongoDB"&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/72gDGmY8-7c" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;The complete source code is available on GitHub:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/AMAN240310/gesture-voice-presenter-remote" rel="noopener noreferrer"&gt;View the project on GitHub&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The project includes the gesture recognition, voice navigation, PDF/PPTX slide extraction, fuzzy matching, and local Gemma/Ollama integration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final Thoughts
&lt;/h2&gt;

&lt;p&gt;I started this project because a friend kept running into a very practical problem while presenting.&lt;/p&gt;

&lt;p&gt;What looked like a simple "control the slides with gestures and voice" project turned into a useful lesson in system design.&lt;/p&gt;

&lt;p&gt;The most important improvement wasn't choosing a bigger model.&lt;/p&gt;

&lt;p&gt;It was &lt;strong&gt;reducing what the model had to solve&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Instead of asking an LLM to understand the entire presentation and return a page number, I combined deterministic rules, fuzzy search, and an LLM only where it actually added value.&lt;/p&gt;

&lt;p&gt;That made the system faster, more predictable, and safer to use during a live presentation.&lt;/p&gt;

&lt;p&gt;And most importantly, Dinesh no longer has to remember:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Was pricing on page 12 or 14?"&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
