DEV Community

Coder
Coder

Posted on

Voice Studio Pro — a local, private, and secure desktop voice assistant

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

Meet Voice Studio Pro — a local, private, and secure desktop voice assistant built specifically for my friends who spend half their day jumping between technical syncs, unplanned meetings, and voice notes.

They hate typing out long follow-up emails, translation messages, and meeting action items from scattered audio recordings, but deal with sensitive project notes daily. Sending raw audio files or confidential transcripts to third-party cloud transcription APIs was a security concern.

Voice Studio Pro solves this by running entirely air-gapped on a local machine. It allows users to record their voice or upload audio files, extract accurate local transcriptions via Whisper, and instantly transform, summarize, or translate those notes into polished executive emails, bullet-point action items, or team updates using an open-weights model—all without a single byte of data touching an external cloud server.

Homepage

Studio interface

How I Built It

Voice Studio Pro is engineered around two powerful open-source AI pillars running locally via Python:

1.Whisper (base model): An open-source speech recognition model used for robust, local acoustic speech-to-text extraction. Audio data captured via Streamlit's microphone recorder or file uploader is written to a secure temporary local file and transcribed entirely on CPU/GPU without internet calls.

2.Qwen (Qwen2.5-1.5B-Instruct): Powered by Hugging Face transformers, this lightweight yet highly capable open-weight instruction model processes the raw transcript. It formats, translates, or refines the text based on the chosen professional style prompt.

3.Frontend & UI: Built with Streamlit and custom glassmorphism CSS styling in a deep midnight plum and soft rose-violet palette.

Tech Stack
1.Language: Python
2.Speech-to-Text (ASR): Whisper (base open-source model)
3.Text Generation & NLP (LLM): Qwen (Qwen2.5-1.5B-Instruct)
4.Deep Learning Framework: PyTorch (torch) for local CPU/GPU tensor execution and hardware acceleration
5.UI & Frontend: Streamlit + Custom Glassmorphism CSS (Midnight Plum & Rose-Violet palette)

Code

Github: https://github.com/Coder-432/voicestudiopro

Why Does Open Innovation Matter?

Open innovation was non-negotiable for this project for two critical reasons:

1.Zero Data Leakage & Absolute Privacy: My friends handle proprietary company meetings, product roadmaps, and internal notes. Closed APIs would require streaming sensitive voice recordings and transcripts to external third-party cloud servers. Running open-source models locally guarantees 100% air-gapped data privacy.

2.Zero Cost & Offline Resilience: Because everything runs locally on a laptop, users can use the tool indefinitely during flights, in remote offices, or offline without incurring subscription fees, API token limits, or network latency.

Top comments (0)