DEV Community

Kaustubh
Kaustubh

Posted on AI-assisted

I Built My Friend a Voice-First AI for Unfinished Thoughts

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

I built Thread for my friend, who often gets ideas while walking, reading or watching a series.

One afternoon, an idea surfaced while he was already concentrating on something else. He decided to write it down after finishing rather than interrupt the work in front of him. When he finally opened his notes, he could remember that the idea had felt useful, but not enough of it to recover the idea itself.

He normally uses a notes app or a notebook. Both work when he has enough time to stop and write. The difficulty is that an unfinished thought rarely arrives as a clear paragraph. Opening an app, deciding what to call the note and arranging the idea into proper sentences can create enough friction for it to disappear.

Thread gives that unfinished thought somewhere to wait.

The experience has three steps:

  1. Press record and speak naturally.

  2. Let Thread transcribe, organise and save the thought.

  3. Return later to see whether it relates to something recorded earlier.

Saving the thought is only the beginning. When a new one is captured, Thread retrieves a small set of relevant earlier thoughts. Gemma then examines whether one of them offers a useful connection, implication or question worth exploring.

A person may mention difficulty concentrating during a walk and, several days later, consider a voice-based space for capturing ideas. Neither recording needs to refer deliberately to the other. Thread can bring the two together and suggest that the later idea may address the earlier problem.

Thread keeps the person’s original words separate from the AI interpretation. The transcript remains visible beside the generated title, summary and suggestions, so he can judge whether the interpretation is faithful. Gemma can also return no connection when the relationship is too weak.

The AI does not decide what he meant or what he should pursue. It offers a possibility. He makes that decision.

Thread helps my friend capture an idea before it disappears, return to his original words and notice when something said on another day gives an unfinished thought a possible direction.

Demo

The demonstration begins with several sample thoughts captured at separate moments. None of them is a complete idea by itself, and the final recording does not deliberately refer to the earlier ones.

After Thread saves the new thought, it retrieves a relevant earlier recording. Gemma explains the possible relationship and offers a question that could help the person continue exploring it.

The screenshot below shows the result of that flow.

The original words remain visible beside Gemma’s interpretation.

Saved Thought

Several recordings later, Thread retrieved a related earlier thought and examined whether the relationship was useful.

Connection Established

Sample data showing how Thread brings back an earlier thought and presents Gemma’s connection as a possibility rather than a conclusion.

The complete journey from the original voice recordings to the connection produced by Gemma is shown in the demo below.

Code

GitHub logo RiverRover-stack / Thread

A voice-first AI thought companion

Thread

A voice-first thought companion for ideas that arrive before the words are in order.

Try the public demo

Thread helps someone capture an unfinished thought without stopping to write and organise it. A person speaks naturally, and the application preserves the transcript, creates a structured interpretation, and saves both to a timeline.

When a later thought is recorded, Thread retrieves relevant earlier thoughts and asks Gemma whether they reveal a useful relationship. It may suggest a connection, an implication, and a question worth exploring—or return no connection when the evidence is weak.

What it does

  • Records a short voice note in the browser.
  • Transcribes the recording with ElevenLabs Scribe.
  • Uses Gemma to create a title, summary, categories, possible action, and question.
  • Keeps the original transcript separate from the AI interpretation.
  • Saves thoughts in a chronological PostgreSQL timeline.
  • Uses embeddings and pgvector to retrieve up to five relevant earlier thoughts.
  • …

How I Built It

Thread follows a thought through three stages: capture what the person said, find earlier thoughts that may be related, and ask whether the relationship is useful enough to show.

I built this as a fixed pipeline rather than an autonomous agent. Each stage has one responsibility, which makes the result easier to inspect when something goes wrong.

Voice → transcript → structured thought → PostgreSQL
                    ↓
               embedding → pgvector
                    ↓
        relevant earlier thoughts
                    ↓
          Gemma connection analysis
Enter fullscreen mode Exit fullscreen mode

1. Preserve the thought before interpreting it

The browser records the voice note and sends the audio to ElevenLabs Scribe for transcription.

The transcript then goes to Gemma, which returns a title, a short summary, categories and, when appropriate, a possible action or a question to explore. Zod validates the structure before the result is saved.

That validation catches responses with missing fields or an unexpected format. It cannot prove that the model interpreted the thought correctly.

For that reason, Thread stores the original transcript separately from every AI-generated field. The person can always return to what he actually said instead of seeing only the model’s summary.

The audio preview is temporary and disappears after the page is refreshed. Thread keeps the transcript as the permanent record, rather than becoming an audio archive.

2. Find earlier thoughts by meaning

Keyword search works when two notes use the same words. Unfinished ideas often do not.

A person might record one thought about losing concentration while studying and, several days later, another about remembering an idea after explaining it aloud. The wording is different, but there may still be a useful relationship between them.

Thread represents each transcript as an embedding: a list of numbers that captures aspects of its meaning. PostgreSQL stores these values using the pgvector extension.

When a new thought is opened, Thread compares its embedding with earlier thoughts from the same workspace. It retrieves at most five candidates that pass the configured similarity threshold.

In simplified form, the retrieval behaves like this:

SELECT
  earlier.title,
  1 - (earlier.embedding <=> current.embedding) AS similarity
FROM thoughts AS current
JOIN thoughts AS earlier
  ON earlier.id <> current.id
  AND earlier.created_at < current.created_at
WHERE current.id = $1
  AND earlier.workspace_id = current.workspace_id
  AND 1 - (earlier.embedding <=> current.embedding) >= $2
ORDER BY similarity DESC
LIMIT 5;
Enter fullscreen mode Exit fullscreen mode

The real query also checks that the thoughts were created with a compatible embedding model and version. This matters because a score produced by one embedding model cannot safely be treated as equivalent to a score from another.

Retrieval narrows the search. It does not decide that two thoughts form a meaningful idea.

Thread currently uses pgvector’s exact cosine-distance search. Since each workspace contains a small number of thoughts, I kept retrieval simple; an HNSW index would become useful as the collection grows.

3. Ask whether the relationship is useful

The current thought and the retrieved candidates then go to Gemma for a second kind of reasoning.

The prompt does not ask, “Are these notes about the same subject?” Two thoughts can both mention learning, work or concentration without helping each other.

Instead, Gemma looks for a more specific relationship. One thought should help address, test, support or challenge a goal, action, constraint or uncertainty in another.

For this step, Thread sends only the structured fields needed for comparison. It does not send the complete timeline, database identifiers, embeddings or raw transcripts.

Gemma can also return no connection. That is part of the design. If the evidence is weak, displaying nothing is more useful than forcing two thoughts into a convincing-sounding story.

When a connection is found, Thread presents it as You Were Onto Something, followed by an explanation and a question the person may want to explore. The wording remains an AI interpretation. The original thoughts stay visible, and the person decides whether the suggestion is worth pursuing.

Connection explanations are generated when the detail page runs the analysis. They are not saved as permanent conclusions, so a later analysis may produce different wording or decide that there is no useful connection.

4. Support both local and hosted Gemma

Thread separates the application logic from the model provider.

For local reasoning, it can use Gemma 3 4B through Ollama. Local embeddings use EmbeddingGemma. This allows the reasoning and retrieval pipeline to run on hardware the user controls.

The public Render deployment uses Gemma 4 through Google AI Studio and Google’s hosted embedding model. I kept this path so judges and other users can try the application without installing models on their computers.

These paths share the same application interfaces, but they do not promise identical results. The local and hosted configurations use different reasoning and embedding models, so each configuration has its own similarity threshold and must be evaluated separately.

Speech transcription currently uses ElevenLabs in both configurations. The local option therefore moves the Gemma reasoning and embedding stages onto the user’s machine; it does not make the complete voice workflow offline.

5. Test the decisions between the AI calls

A successful API response only proves that a service answered. It does not prove that Thread found the right earlier thought or produced a useful connection.

I tested the stages separately:

Evaluation Observed result
Related thoughts with local embeddings Similarity 0.748; included above the local 0.70 threshold
Related thoughts in the hosted deployment Similarity 0.8545; included above the hosted 0.80 threshold
Unrelated thought Excluded from retrieval
Two passive observations sharing a topic Gemma returned no useful connection
No qualifying earlier thoughts Connection inference was skipped

The local and hosted similarity numbers should not be compared as a model ranking because they come from different embedding models. These checks verify several expected cases; they are not a measurement of general accuracy.

I also added Sentry tracing around transcription, embedding, retrieval and connection reasoning. In one production verification trace, retrieval took about 25 ms and connection inference took about 918 ms. That is one request rather than a performance benchmark, but it helped show where the time was spent.

Sentry Agent Activity showing Thread's retrieval and Gemma connection stages

One production trace of Thread's connection workflow. Sentry separates retrieval from Gemma inference and records timing and token usage without including the thought text.

The tracing export is filtered so thought content, transcripts, generated text, embeddings and record identifiers are not included.

The main engineering decision in Thread is the separation between finding something similar and deciding whether it is useful. Retrieval gives Gemma a small amount of relevant context. Gemma can then suggest a possible direction—or leave the thought alone when there is not enough evidence.

Why Does Open Innovation Matter?

A closed API could summarise a thought or return the same structured fields. What it cannot provide is a model whose weights I can run on the user’s own computer.

That matters for Thread because one recording is only a note, but many recordings gradually become a history of what someone has been thinking about. In the local configuration, Gemma 3 4B structures thoughts and examines their connections through Ollama. EmbeddingGemma creates the vectors used for retrieval, and PostgreSQL stores the timeline on the same computer.

ElevenLabs receives one recording for transcription, but it does not receive the timeline or the related thoughts. After the transcript returns, the accumulated memory, semantic search and connection analysis can remain local.

The public demo uses hosted Gemma, hosted embeddings and Render PostgreSQL so anyone can try Thread without installing models. That is a deployment choice rather than a requirement of the application. The reasoning and embedding layers sit behind provider interfaces, so the local and hosted versions follow the same flow even though their models may produce different results.

Open weights also mean that Gemma can be evaluated, replaced or eventually customised without rebuilding Thread around a different proprietary API. I have not fine-tuned the model or established a cost advantage. The benefit demonstrated here is more direct: the part of the application that reasons across a person’s growing thought history can run on hardware they control.

My Agent Session

I built Thread over the hackathon weekend with Codex. I used the session to turn the original idea into a working application, examine weak assumptions, trace failures and prepare the project for someone other than me to try.

Prize Categories

  • Best Use of Gemma - thought structuring and connection analysis, with local and hosted adapters.
  • Best Use of ElevenLabs - speech transcription for the voice capture flow.
  • Best Use of Render - the deployed Next.js application and PostgreSQL database.
  • Best Use of Sentry Agent Tracing - instrumented AI stages and connection workflow.

Top comments (0)