DEV Community

Godly K Mathews
Godly K Mathews

Posted on

An Offline RAG based Voice Assistant Built for a Friend

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

I built Fieldmate, a local voice assistant with document retrieval, for my friend who works as a nurse.

The problem I wanted to solve was simple: help them look things up in their documents, remember an ongoing conversation, and capture notes and tasks—even when an internet connection is unavailable.

Fieldmate brings those activities into one interface:

  • Talk or type: use push-to-talk or a hands-free mode that listens between responses.
  • Ask about documents: upload PDFs, Markdown, or text files and get answers with source references.
  • Continue a conversation: separate chat threads retain their own history and context.
  • Organize by subject: categories such as Fieldwork, Biology, and Physics have their own instructions and document collections.
  • Capture work: say “Save a note…” or “Add a task…” to store information locally.
  • Choose when to go online: review and approve a proposed search query before it leaves the laptop.

The interface has two main tabs: Chat and Documents. Settings allow changes to the local model and assistant instructions.

This is a working laptop prototype. Testing it with my intended recipient in their actual workflow is the next step.

Demo

Fieldmate runs locally, so I recorded a video demonstrating the application.

Code

Field Assistant

License: MIT Python

Fieldmate is a local AI voice assistant for working with your documents, keeping conversational context, and capturing notes and tasks. It runs on your laptop using Ollama and locally stored speech models. After setup, the core workflow works offline; internet search is optional and requires approval.

Built for a friend who works as a nurse, and adaptable to other document-based workflows through subject categories and custom instructions.

Watch the demo · Report an issue · MIT license

Fieldmate chat interface

Contents

Features

  • Local inference: use compatible installed Ollama models, including Qwen and Gemma.
  • Document retrieval: import text PDFs, Markdown, and UTF-8 text files; inspect sources and page references.
  • Persistent conversations: create, rename, close, and reopen separate chat threads with their own context.
  • Subject categories: organize documents and instructions by topic, with shared documents available across…

The application code uses the MIT license. Its dependencies and model weights retain their respective licenses.

How I Built It

Fieldmate uses a FastAPI backend, a lightweight HTML/CSS/JavaScript interface, and SQLite for persistent local storage. The backend coordinates inference, document retrieval, conversation memory, voice processing, and tool execution.

Component Role
Ollama Serves local language and embedding models
Gemma 4 Generates responses and makes structured routing decisions
Nomic Embed Text Embeds document passages and retrieval queries
SQLite Stores conversations, summaries, document vectors, notes, tasks, and search permissions
pypdf Extracts text and page information from PDFs
faster-whisper Transcribes microphone recordings locally
Piper Generates spoken responses locally
Laya Provides an advisory answer-versus-clarify signal
DDGS Performs optional, approved DuckDuckGo searches

Local inference and structured decisions

The backend communicates with Ollama through its local HTTP API. It uses /api/chat for generation and /api/embed for embeddings.

For routing and response generation, I pass Pydantic-generated JSON schemas to Ollama’s structured-output interface. The backend validates the returned objects before using them.

For example, the conversation planner returns this structure:

class Plan(BaseModel):
    route: Literal["answer", "clarify", "search"]
    query: str = Field(max_length=1200)
    clarification: str = Field(default="", max_length=500)
Enter fullscreen mode Exit fullscreen mode

This separates the decision about what to do from the code that performs the action. A model can propose an internet search, but that proposal cannot grant permission to execute it.

Model selection is configurable. Fieldmate checks installed Ollama models and their capabilities, allowing compatible local chat models to be swapped without discarding the conversation.

Document ingestion and retrieval

The ingestion pipeline works as follows:

  1. Extract text from a PDF, Markdown, or UTF-8 text file.
  2. Split it into passages of approximately 1,600 characters, with a 30-word overlap when splitting long paragraphs.
  3. Generate Nomic embeddings through Ollama.
  4. Store each passage, embedding, document identity, and page number in SQLite.

A SHA-256 digest detects duplicate document uploads. Scanned PDFs need OCR before import; this prototype extracts existing text.

For retrieval, I use Nomic’s search_document: prefix for passages and search_query: for questions. The backend performs an exact cosine-similarity scan over the eligible stored vectors.

Passages must pass a configurable similarity threshold. Their final ranking combines semantic similarity with keyword overlap:

score = 0.8 Ă— cosine_similarity + 0.2 Ă— keyword_overlap
Enter fullscreen mode Exit fullscreen mode

The top five passages are supplied to the answering model. Documents can belong to a category or be shared across categories.

This keeps retrieval simple enough for a laptop prototype without requiring a separate vector database service.

Grounding answers in evidence

The response schema includes the answer text, its stated basis, a document-check result, and supporting quotations.

When the model cites a passage, the backend checks that:

  • The source ID belongs to the retrieved sources.
  • The quotation occurs in that source.
  • The reference has not been invented.

The UI then shows the supporting document and page information.

These checks validate quotation provenance, not the truth of every conclusion. Invalid references produce an explicit notice. When documents are not relevant, Fieldmate can still answer from general model knowledge and label the response accordingly.

Conversation memory and contextual actions

I keep conversation memory separate from document retrieval.

Each thread has its own persistent transcript, category, model selection, and custom instructions. The backend loads history from SQLite rather than trusting conversation history submitted by the browser.

As a conversation grows, older exchanges are compacted into a rolling summary. The original transcript remains available for operations that need exact content.

That distinction became important when fixing a real bug:

“Save this as a note” initially saved the word “this.”

The capture handler was bypassing contextual resolution. I changed it so that:

  • “Save this” copies the preceding answer, skipping action confirmations.
  • “Save our study plan” searches the current thread’s stored messages for the relevant plan.
  • Whole-message captures use schema-constrained source IDs, and the backend copies the original text.
  • Requested excerpts must match the selected source.
  • Missing or invalid selections ask for clarification without writing a note.

Notes and tasks are saved in the same SQLite transaction as their request receipt. Repeating the same request ID returns the existing result instead of creating another item.

The distinction between compact context and recoverable source text is also explored in Smaller Context, Recoverable History. The separate problem of choosing an outdated version is discussed in Agent Memory’s Real Failure Is Currency, Not Retrieval.

Clarification and internet permissions

I integrated Laya to provide an advisory signal about whether a request needs clarification. Structured Ollama planning combines that signal with conversation context.

I also added deterministic checks for clear missing references. For example, a new conversation asking “Can I take it with my other medicine?” needs the medicine names before a useful answer.

Web search follows a separate backend permission flow. Fieldmate proposes an editable query and waits for approval. The permission is tied to that conversation and can be consumed only once.

Only the approved query is sent to the search provider. The application does not automatically upload the full conversation or document collection. Returned search snippets are treated as external evidence, with their limitations retained.

The local voice pipeline

Voice input passes through faster-whisper, using the multilingual Whisper base model with CPU int8 inference. The transcription can be reviewed in the composer before sending.

For speech output, Piper generates a WAV response using the locally stored Lessac voice model.

Hands-free mode uses microphone volume and silence detection to identify turns. Listening pauses while the assistant processes or speaks, so the prototype uses turn-taking rather than simultaneous listening and speaking.

I also added a speech-cleanup layer after discovering that the voice was reading Markdown asterisks aloud. It removes formatting and citation markers while converting meaningful mathematical operators into spoken words.

Testing the complete workflow

I used Codex to help implement and debug Fieldmate.

The automated suite reached 61 passing tests, covering conversation isolation, contextual captures, retrieval behavior, citation checks, search permissions, task dates, retries, and transaction rollback.

I also tested against local models. With Gemma 4 12B, the contextual-save checks verified the original study-plan content, recovery after a topic change, selection of a newer revision, and creation of a task from a referenced step.

Those live checks caught behavior that test doubles missed—including the model trying to rewrite a quotation instead of selecting the stored text.

Why Does Open Innovation Matter?

Local operation is the reason this project works for its intended setting.

After the initial model downloads, the core assistant can run without an internet connection. Document retrieval, conversation, transcription, and speech generation happen on the laptop. Optional web search is the exception and requires approval.

That matters because notes and documents do not need to be sent to a hosted inference service just to ask a question about them.

Open tooling and locally runnable models also gave me control over the assistant’s behavior. I could inspect the decision flow, change retrieval logic, improve clarification, and fix the speech formatting problem.

The model is replaceable, too. Fieldmate can switch between compatible installed Ollama models while retaining the conversation. I tested a switch from Qwen to Gemma with a contextual follow-up.

There are no per-request cloud inference charges for the local workflow, although it still uses the laptop’s storage, memory, and electricity.

A hosted API would have made this prototype dependent on connectivity and external processing. This approach lets the assistant remain useful on the machine my friend already has.

My Agent Session

Planning Fieldmate: an offline voice assistant for a field worker
You

Build for a Friend: build something with open-source AI at its core that solves a real problem for a friend or loved one. I'm about to build something; we'll discuss what to build.

Agent

Who might you build this for, and what’s one recurring task or frustration they have? We’ll look for something small enough to finish this weekend, useful enough to hand over, and where open-source AI brings a real benefit such as privacy, offline use, or customization.

You

I was thinking of a friend who does a lot of ground work, like a field worker. An offline assistant who can interact continuously like a voice assistant plus RAG, and save todos or quick notes. How to build this kind of tech? I'll have to use Ollama.

Agent

Proposed architecture: local speech recognition → coordinator model through Ollama → mandatory document retrieval for work questions or validated tools for notes and tasks → local speech synthesis. Suggested components: Qwen3 4B, whisper.cpp, Piper, local embeddings, Chroma and SQLite. Retrieve before answering, preserve source references, abstain when evidence is insufficient, and keep field observations separate from official procedures. Start with one coordinator and explicit tools, then add hands-free interaction. Device choice determines whether Ollama is practical.

You

We'll stick to laptop prototype. Also how to save DevRelay?

Agent

Save the development conversation as a DevRelay agent session, remove secrets and private paths, and embed the returned session ID in the DEV submission with the agent_session Liquid tag. This is separate from conversations with the finished assistant.

You

Everything is set. I have attached the folder to develop the application. I have signed in to DevRelay with MLH. Also save the DevRelay session and start developing.

Building Fieldmate: offline RAG, voice, notes and tasks with Ollama
You

We'll stick to a laptop prototype.

You

Everything is set. I have attached the folder to develop the application. I have signed in to DevRelay with MLH. Also save the DevRelay session and start developing.

Agent

I’ll build the laptop prototype in your Assistant folder and save the planning session to DevRelay. I’ll first check the folder, local tools, and Ollama setup, then implement document search, voice input and output, notes, and tasks.

Agent

Your planning session is saved privately on DEV, and Ollama already has Qwen3 8B and a local embedding model installed. I’ll use those for the first build. The prototype will have document-backed answers with visible source excerpts, persistent notes and tasks, push-to-talk, and a hands-free session mode.

Agent

The real-model test caught a routing error that the mocked tests couldn’t: Qwen treated a document question as a request to search personal notes. I’m tightening the router so ordinary work questions go directly to the fieldbook, and notes or tasks require clear user intent.

Agent

Document answers now pass the real-model checks, including refusal to answer questions missing from the guide. The browser also shows the source excerpts correctly. Voice testing found an audio-library compatibility issue; I’ve pinned a compatible version and am checking the full Piper-to-Whisper audio path.

Agent

The full integration check passes with outbound connections blocked: local document answers, missing-evidence refusal, note/task persistence, and Piper-to-Whisper transcription.

Fieldmate v2: contextual local assistant, categories and permission-gated web search
You

Update the laptop assistant with a simple dark chat interface, a Documents tab, multiple threads, subject categories, custom prompts, contextual conversations, permission-based web search, document cross-checking, and a Laya-like clarification engine. Support switching Ollama models, including Gemma 4.

Agent

Curated build summary: migrated existing SQLite exchanges into a legacy conversation; added isolated thread histories, rolling summaries, categories, prompts, and per-thread models. Separated conversation memory from document retrieval. The interface has Chat and Documents tabs plus settings and a notebook.

Agent

Integrated the open Laya typed-decision checkpoint as an advisory local ambiguity signal. A real test showed its answer hint could encourage an overly general response when medicine names were missing. Added deterministic missing-reference clarification for that case. Ollama handles nuanced context and question wording. No confidence score is treated as calibrated.

Agent

Added exact-source quotation checks around natural-language answers. Replies distinguish model knowledge, document sources, mixed evidence, and uncertainty. Exact quotation provenance is not a guarantee that generated claims are correct. Web search requires an editable query and single-use server-side approval bound to its conversation. Only the approved query is sent to DuckDuckGo; failures are explicit.

Agent

Validation: 31 automated tests passed, including thread isolation, migration, compaction, categories, atomic notes/tasks, fabricated citation rejection, single-use search approval, and rejection of remote model tags. Real offline checks passed with Qwen3, Gemma4 e2b, Laya, Nomic embeddings, and a Piper-to-Whisper audio round trip. Switching from Qwen to Gemma preserved a photosynthesis conversation. A public documentation query returned live search results. The installed app retains the user's existing records and a pre-migration SQLite backup.

Agent

Open innovation enables local inference and private document handling without recurring cloud inference fees, replaceable chat models, inspectable routing, and offline speech. Limitations: bounded/summarized context, heuristic retrieval and clarification, search snippets rather than full web pages, and turn-taking voice rather than full duplex. This is a curated summary, not a raw transcript. Private paths, existing user conversation content, and credentials have been excluded.

Prize Categories

Best Use of Gemma

Fieldmate runs Gemma 4 locally through Ollama for conversational assistance, structured decisions, and document-grounded answers. I used Gemma 4 E2B during development and tested the contextual capture workflow with Gemma 4 12B.

Gemma fits into a workflow where the user’s documents and conversation stay on their laptop.

Top comments (0)