I wasn't building another voice recorder.
I was building a meeting companion.
A small app that sits on your desktop, listens to a call, and writes down:
- every word
- who said it
- what was decided
Then it saves the conversation somewhere your team can find it later.
- Record a call with three people.
- Press stop.
- Open the web app.
- See a clean transcript, a summary, and your own name next to your own words.
That was the goal.
And for a while...
It looked done.
Then I played a YouTube video.
And the app decided the YouTube video was me.
Index
- The Goal
- The Architecture
- Bug #1 — The Transcript Froze
- Bug #2 — The YouTube Video Became Me
- Bug #3 — Two Microphones, One Truth
- Bug #4 — Computer Audio Made Me a Stranger
- Bug #5 — "Remember This Voice" Remembered Nothing
- Bug #6 — "Okafor" Became "Aquifer"
- Bug #7 — The Vectors That Were Never Stored
- The Final Architecture
- Lessons Learned
- Final Thoughts
1. The Goal
The idea was simple.
Your Microphone + Your Computer's Audio
↓
Live Transcription
↓
Who Is Speaking?
↓
Save to the Web App
↓
Summary + Search
The result should be:
- accurate words
- the right name on every line
- findable later by meaning, not just by keyword
Sounds straightforward. It wasn't.
2. The Architecture
The system had four parts.
The recorder
A desktop app built with Electron. It captures your microphone, and optionally your computer's sound as a second channel, so the other side of a call is recorded separately from you.
Live transcription
Audio streams to Deepgram's speech-to-text model while you talk. It passes through a small local relay that holds the API key, so the key never ships inside the app.
The words come back already split by voice. Deepgram calls this diarization: Speaker 0, Speaker 1, Speaker 2.
But diarization only knows that voices are different.
Not whose they are.
Speaker verification
That's the job of VoiceBio, a speaker identification engine.
You read a short paragraph once. It builds a voiceprint, a mathematical fingerprint of your voice. The voiceprint is encrypted and stays on your own computer.
During a recording, a few seconds of each voice is compared against it. The engine returns a similarity score: high means "that's you", low means "someone else".
The web app never sees your voice. It only receives the name tag.
The memory
The database is the record. A vector database (Pinecone) is the index.
Each transcript is cut into chunks, turned into embeddings, and stored, so you can later search for what was meant, not just what was said.
3. Bug #1 — The Transcript Froze
Halfway through a recording, the live text stopped.
Not slowed down. Stopped.
The logs showed one voice check that took exactly 120 seconds.
The Cause
The voice check had inherited the timeout of voice enrollment. Enrollment uploads a long recording and trains on it, so it was given two minutes.
Checking a five-second clip should take about one.
The Fix
A dedicated short timeout for checks. A slow check now fails fast, and the transcript keeps flowing.
A timeout isn't a detail. It's a decision about what the user waits for.
4. Bug #2 — The YouTube Video Became Me
I started a recording, played a video, then started talking.
The video was labeled as me.
I was labeled "Speaker 2".
And when my name finally appeared, it took about 15 seconds.
The Cause
Before any voice was verified, the app guessed: "the first voice heard is probably the user."
Reasonable in a quiet room.
Wrong the moment anything else speaks first.
And the real check needed several seconds of speech plus processing time, so the guess sat on screen long enough to look like an answer.
The Fix
Stop guessing.
If you have a voiceprint, nobody becomes "you" until the engine confirms it. Until then, the name is dimmed with "Identifying…".
Honest beats fast.
But 15 seconds still felt slow.
5. Bug #3 — Two Microphones, One Truth
Then came a better idea, borrowed from how dedicated meeting-notes tools work.
Don't ask whose voice it sounds like.
Ask which microphone it came from.
Channel 0 → Your Microphone → You
Channel 1 → Your Computer → Everyone else on the call
Deepgram transcribes both channels on one connection and tags every sentence with its channel. Your words come from your microphone, so your name can appear instantly.
Or so I thought.
On speakers instead of headphones, your microphone also hears the other side of the call. The same sentence arrived twice: once from the computer, once from my mic.
I added an echo filter: drop a microphone line when the computer channel said the same words within two seconds. But never drop short replies, because "yeah" said over someone else is a real thing people say.
Then a subtler problem.
During pauses, a bit of echo slipped through and briefly flashed my name on someone else's words. I tried a "quiet window" before trusting the microphone. Two seconds. Then four.
No window was ever safe.
The Fix
The channel tells you where audio came from. Not who said it.
With a voiceprint stored, the microphone voice is still verified.
6. Bug #4 — Computer Audio Made Me a Stranger
With computer audio on, the app stopped recognizing me completely.
The log said:
similarity: -11 → not you
My voice used to score around +23. Other people scored around -57.
-11 was somewhere strange in between.
The engine wasn't wrong.
It was being fed garbage.
The Cause
With two channels, audio is stored interleaved: one sample from the microphone, one from the computer, alternating.
The code that cut a clip of my voice still assumed one channel.
So the clip it sent was:
- from the wrong moment
- at the wrong speed
- with my microphone and the computer mixed sample by sample
The Fix
Cut each voice from its own channel only, at the right speed.
After the fix:
Me +9.98 → recognized, locked
Computer audio -57 → correctly someone else
A bonus: a check of my voice no longer carried the other side of the call along with it.
7. Bug #5 — "Remember This Voice" Remembered Nothing
There's a feature to name a speaker, say "Milana", and tick remember this voice, so future recordings recognize her automatically.
I named her. Recorded again.
Nobody was named Milana.
The Cause
Two things, stacked.
First, the name was applied to the transcript right away, but training her voiceprint needs about 20 seconds of her speech. When there wasn't enough, training failed silently. The screen showed success either way.
Second, the interesting one.
To reach 20 seconds, I stitched her short sentences together into one clip, gaps removed. The engine kept finding only about 6 seconds of usable speech out of 40.
The same audio, uploaded as one continuous stretch with the pauses left in, enrolled instantly.
The hard cuts between sentences were confusing the engine's voice detection.
The Fix
- Train on her densest continuous stretch of speech
- If training fails, say so
- Retry automatically as more of her speech arrives
8. Bug #6 — "Okafor" Became "Aquifer"
Now the words themselves.
Instead of guessing, I measured. I streamed a known script through the exact production pipeline and compared it word by word.
9.8% word error rate.
Most of the errors were names:
- Okafor → "Akhafar", "Aquifer"
- Priya Raghunathan → "Priyuragunathan"
The Fix
Tell the model which names to expect. Deepgram calls this keyterm prompting.
The list comes from your own name, the voices you've saved, and a short list you can edit yourself.
3.3% word error rate. Every name correct.
I also tried waiting longer before ending a sentence. It got worse (12%), so I left it alone.
Honest caveat: I measured on synthetic speech. A real laptop mic in a real room will gain less. But names are the first words people check.
9. Bug #7 — The Vectors That Were Never Stored
Now the search side.
After a recording is saved:
Save to Database (the record)
↓
Answer the app immediately
↓
In the background:
Chunk → Embed → Store in Vector Index
Chunking a transcript is different from chunking an article. I cut it by speaker turns, around 1,200 characters, each line written as "Name: words", and carried the previous line into the next chunk so a question and its answer stay together.
Every chunk stores metadata: whose recording, which workspace, which lines, and a link back to the exact moment.
Privacy mattered here. Each user gets their own space in the index, so one person's private conversation can't show up in someone else's search. And every hit is re-checked against the database before it's shown.
Shipped. Tested. Searched.
Nothing.
The background job's status said failed:
You've reached the max namespaces allowed in this index (100).
The index was full. 100 spaces. Ours was the 101st.
And it wasn't only our feature. Every new project and workspace in the whole app would fail the same way. 18 of those 100 belonged to things that had already been deleted.
While I was in there, I found one more.
The delete call passed a plain list of ids where the library expected them wrapped in an object. So deleting a document's vectors silently deleted nothing. Deleted content stayed searchable.
The Fix
- The deletes are fixed
- The namespace limit is a plan and design decision (upgrade, clean up, or share one space with stricter filters), and it's being decided as I write this
And this is the part I'm glad about.
Because the database is the record and the vectors are only an index, the failure lost nothing. Every recording is saved, waiting to be indexed the moment there's room.
10. The Final Architecture
Microphone ──┐
├──► Two-Channel Stream ──► Local Relay ──► Speech-to-Text
Computer ────┘ (names hinted)
│
▼
Lines, split by voice
│
┌──────────────────────────────────┤
▼ ▼
Voice check, own channel only Echo filter
vs voiceprint on this computer
│
▼
You / Named / "Identifying…"
│
▼
Upload (retry-safe, works offline)
│
▼
Database (the record) ──► Summary
│
▼
Chunk ──► Embed ──► Vector Index (per user)
│
▼
Search by meaning ──► check owner ──► open the exact line
11. Lessons Learned
1. Honest beats fast
A dimmed "Identifying…" is better than a confident wrong name.
2. Know where audio came from before deciding who it is
Channels are cheap, strong evidence. But they're not proof.
3. When a model gives a strange score, check what you fed it
-11 wasn't the voice engine's fault. It was my audio slicing.
4. Measure before you tune
One known script and a word error rate told me the problem was names, not the model.
5. Keep the record and the index separate
When the database is the truth and the vectors are just an index, a vector outage is an inconvenience. Not data loss.
6. Your vector database has limits you only meet in production
A namespace limit. A batch limit: the embedding model accepts 96 inputs per call, and we were sending 120.
7. Silent failures are the worst failures
"Remember this voice" and the deletes both reported success while doing nothing.
12. Final Thoughts
Before this project, I thought knowing who said what was a model problem.
Pick a good speech model. Pick a good voice engine. Done.
I was wrong.
It turned out to be:
- timeout problems
- guessing problems
- audio format problems
- channel layout problems
- splicing problems
- vector database limits
- silent failures
Every fix uncovered another hidden weakness.
Today the app can transcribe a call live, put the right name on your words, keep the other side of the call separate, and save everything somewhere your team can find it.
The biggest lesson wasn't picking a better model.
It was learning that knowing who said what is a pipeline problem, not just a model problem.
Top comments (1)
The interleaved-channel bug is a useful reminder that speaker verification scores depend on the audio contract. I would keep a deterministic two-channel fixture with different tones or spoken markers on each channel, then assert the extracted clip's channel, sample count and time interval before sending it to verification.
That catches the wrong-speed and wrong-moment errors without depending on a biometric model's threshold. It is especially valuable when capture devices change sample rate or channel count: the transcript can remain readable while the short verification clip is corrupted in a way the UI never exposes.