Cloud transcription tools have a problem nobody talks about: they upload your audio to someone else's server.
For most people, that's a minor privacy tradeoff. For lawyers, doctors, and consultants, it's a dealbreaker. Attorney-client privilege, HIPAA, and corporate data policies all prohibit sending meeting audio to a third party. So these professionals are stuck taking manual notes while everyone else gets AI-powered transcripts.
I built Whisper Notes to fix that — a fully offline meeting transcriber that runs entirely on the Snapdragon X Elite NPU.
What It Does
Whisper Notes captures both your microphone and your system audio (Zoom, Teams, Google Meet) through WASAPI loopback, transcribes in real time, and generates a summary with action items. Everything runs on-device. Nothing touches the network.
- Meet Mode — one click captures system audio + mic simultaneously
- Real-time transcription — Whisper-Small-Quantized on the Hexagon NPU
- Local summarization — action items extracted on-device
- Encrypted storage — SQLite with AES-256
- Verifiable privacy — open any network monitor; zero bytes sent
Why Snapdragon
This is built on Qualcomm AI Hub's Whisper-Small-Quantized model, compiled for the Hexagon NPU with the QNN runtime.
| Component | Latency | Compute Unit |
|---|---|---|
| Decoder | ~7 ms | NPU |
| Encoder (30s audio) | ~308 ms | NPU |
The 7ms decoder latency is the difference between waiting for transcription and watching words appear as people speak. And because it runs on the NPU instead of the CPU or GPU, you get hours of meeting capture on battery — not minutes.
Architecture
Audio Capture (WASAPI loopback + mic)
↓
Whisper-Small-Quantized (NPU via Qualcomm AI Hub)
↓
Local Summarization (Qwen3-0.6B or rule-based)
↓
Encrypted Local Storage (SQLite + AES-256)
text
Every stage runs on-device.
The Interesting Code
Loading the model onto the NPU:
python
from geniex import AutoModel
class WhisperTranscriber:
def __init__(self, model_path):
self.model = AutoModel.from_pretrained(
model_path,
device_map="qairt" # this is what sends it to the NPU
)
def transcribe(self, audio):
input_features = self.preprocess_audio(audio)
output = self.model(input_features)
return self.decode_tokens(output)
Capturing system audio with WASAPI loopback:
python
import sounddevice as sd
# Find the loopback device (Stereo Mix or default render)
devices = sd.query_devices()
loopback_device = None
for i, dev in enumerate(devices):
name = dev["name"].lower()
if "stereo mix" in name or "loopback" in name:
loopback_device = i
break
stream = sd.InputStream(
device=loopback_device,
channels=1,
samplerate=16000,
callback=callback,
blocksize=1600,
)
The audio callback pushes 100ms chunks into a queue, and the transcription loop pulls 10-second windows for Whisper.
What I Learned
1. NPU inference is genuinely fast — not just a benchmark number.
The 7ms decoder latency isn't marketing. It's the reason the transcript feels live instead of laggy. On a CPU, the same model would be seconds behind.
2. Privacy is the feature, not the constraint.
I kept framing "no cloud" as a limitation while building this. Then I realized: for the target user, offline isn't a downgrade. It's the entire reason they'd use the tool. A lawyer doesn't want a faster cloud transcriber — they want one that never leaves the room.
3. System audio capture is harder than it looks.
WASAPI loopback works, but device discovery is inconsistent across Windows machines. Some laptops expose "Stereo Mix," others don't. The fallback is the default render device, but that's not always right. This is the part I'd spend more time on with more runway.
4. The model download is the biggest setup friction.
Once the AI Hub model is in place, the app is a single python app.py away. But getting the QNN-compiled binaries onto the machine is the step that trips people up. A one-click installer would massively improve adoption.
Try It
The repo is public:
GitHub: github.com/SASIDHARPRATHIPATI/Whisper-notes
Top comments (0)