This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
What I Built
LocalWhisper Pro is an open-source, offline-first voice interface designed to minimize screen time. Instead of spending hours hunched over a keyboard typing emails, documentation, or field notes at 40 WPM, users dictate naturally at 150+ WPM. Local AI instantly cleans, restructures, and auto-types the text directly at the cursor—enabling users to wrap up screen work fast and get outdoors.
How it gets people off the screen and into the world:
- Makes the screen the shortest interaction: Speech is up to 4x faster than typing. By handling filler-word scrubbing, formatting, and summarization automatically, users spend a fraction of the time staring at word processors and code editors.
- 100% Offline Trail & Field Ready: Because both transcription and refinement run entirely on-device, you can take your laptop outside—onto the porch, into a garden, or down a hiking trail without cell signal—to dictate field logs, nature observations, and ideas without internet dependencies.
-
No Clipboard or Cloud Distractions: It uses low-level hardware stroke injection (
SendInput/pynput) rather than pasting through clipboard history, keeping workflows seamless and distraction-free.
Demo
-
Direct Download: Download the standalone Windows executable (
LocalWhisper_Pro_v3.1_Windows.zip) directly from our GitHub Releases page—no Python runtime required.
Code
Repository: [https://github.com/Harsh-Dev07/Local-WhisperFlow]
License: Open Source (MIT)
How I Built It
LocalWhisper Pro is architected entirely around open-source AI models and local runtime frameworks:
Local Speech Recognition (
faster-whisper/ CTranslate2):
Audio input is captured locally usingsounddeviceand processed throughfaster-whisper, a reimplementation of OpenAI's Whisper model running on the CTranslate2 inference engine. This delivers up to 4x faster-than-real-time transcription on consumer CPUs/GPUs without external network calls.-
Local Intelligence & Text Refinement (Ollama +
llama3.1:8b):
Raw transcription often suffers from stutters, repetitions, and unstructured ramblings. LocalWhisper integrates directly with local Ollama instances. Using open-weight models likellama3.1:8b, it cleans transcripts across multiple modes:- Smart Dictation: Removes filler words while maintaining personal voice.
- Field & Bullet Summary: Distills stream-of-consciousness rambling into structured, actionable checklists.
- Code & Tech: Synthesizes voice logs into camelCase, snake_case, and terminal commands.
- Hinglish/Hindi Processing: Accurately interprets code-switched multi-language dictation.
System Overlay & Direct Keystroke Emulation:
The interface is a lightweight, non-stealing glassmorphic HUD built in Python (Tkinter). Refined text bypasses the operating system's clipboard (Win+V) and is typed directly into active windows via hardware-level event hooks (pynput), ensuring zero data retention outside the user's encrypted local history.
Why Does Open Innovation Matter?
Open innovation was not an afterthought for this build—it was a prerequisite:
-
True Offline Portability: Closed-source voice-to-text platforms (like Google Speech-to-Text or OpenAI Whisper API) fail the moment you walk into a park, mountain trail, or remote campsite without Wi-Fi or cellular service. Open-weight models (
faster-whisper+ Llama 3.1) run autonomously on consumer laptops off the grid. - Complete Sensory Privacy: Voice notes recorded outdoors or in private moments often include personal thoughts, journal entries, and sensitive project ideas. Open-source local models guarantee that no spoken audio or transcribed text ever hits corporate telemetry servers or third-party training pipelines.
- Zero Ongoing Cost: Proprietary APIs charge per audio minute and per token, which discourages long, ambient voice captures. Open-weight inference costs nothing to run, whether you record a two-second note or an hour-long outdoor brainstorm.
-
Model Modularity: Users have the freedom to swap out the refiner engine for smaller models (e.g.,
phi3,gemma2:2b) on low-power devices, or larger parameter models on dedicated hardware.
Prize Categories
- Hacktoberfest Open-Source AI Challenge: Week 1 (Touch Grass)

Top comments (0)