This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
I built Fieldmate, a local voice assistant with document retrieval, for my friend who works as a nurse.
The problem I wanted to solve was simple: help them look things up in their documents, remember an ongoing conversation, and capture notes and tasks—even when an internet connection is unavailable.
Fieldmate brings those activities into one interface:
- Talk or type: use push-to-talk or a hands-free mode that listens between responses.
- Ask about documents: upload PDFs, Markdown, or text files and get answers with source references.
- Continue a conversation: separate chat threads retain their own history and context.
- Organize by subject: categories such as Fieldwork, Biology, and Physics have their own instructions and document collections.
- Capture work: say “Save a note…” or “Add a task…” to store information locally.
- Choose when to go online: review and approve a proposed search query before it leaves the laptop.
The interface has two main tabs: Chat and Documents. Settings allow changes to the local model and assistant instructions.
This is a working laptop prototype. Testing it with my intended recipient in their actual workflow is the next step.
Demo
Fieldmate runs locally, so I recorded a video demonstrating the application.
Code
Field Assistant
Fieldmate is a local AI voice assistant for working with your documents, keeping conversational context, and capturing notes and tasks. It runs on your laptop using Ollama and locally stored speech models. After setup, the core workflow works offline; internet search is optional and requires approval.
Built for a friend who works as a nurse, and adaptable to other document-based workflows through subject categories and custom instructions.
Watch the demo · Report an issue · MIT license
Contents
- Features
- Requirements
- Installation
- Usage
- Configuration
- Architecture
- Privacy and limitations
- Development
- Contributing
- Troubleshooting
- Documentation and acknowledgments
- License
Features
- Local inference: use compatible installed Ollama models, including Qwen and Gemma.
- Document retrieval: import text PDFs, Markdown, and UTF-8 text files; inspect sources and page references.
- Persistent conversations: create, rename, close, and reopen separate chat threads with their own context.
- Subject categories: organize documents and instructions by topic, with shared documents available across…
The application code uses the MIT license. Its dependencies and model weights retain their respective licenses.
How I Built It
Fieldmate uses a FastAPI backend, a lightweight HTML/CSS/JavaScript interface, and SQLite for persistent local storage. The backend coordinates inference, document retrieval, conversation memory, voice processing, and tool execution.
| Component | Role |
|---|---|
| Ollama | Serves local language and embedding models |
| Gemma 4 | Generates responses and makes structured routing decisions |
| Nomic Embed Text | Embeds document passages and retrieval queries |
| SQLite | Stores conversations, summaries, document vectors, notes, tasks, and search permissions |
| pypdf | Extracts text and page information from PDFs |
| faster-whisper | Transcribes microphone recordings locally |
| Piper | Generates spoken responses locally |
| Laya | Provides an advisory answer-versus-clarify signal |
| DDGS | Performs optional, approved DuckDuckGo searches |
Local inference and structured decisions
The backend communicates with Ollama through its local HTTP API. It uses /api/chat for generation and /api/embed for embeddings.
For routing and response generation, I pass Pydantic-generated JSON schemas to Ollama’s structured-output interface. The backend validates the returned objects before using them.
For example, the conversation planner returns this structure:
class Plan(BaseModel):
route: Literal["answer", "clarify", "search"]
query: str = Field(max_length=1200)
clarification: str = Field(default="", max_length=500)
This separates the decision about what to do from the code that performs the action. A model can propose an internet search, but that proposal cannot grant permission to execute it.
Model selection is configurable. Fieldmate checks installed Ollama models and their capabilities, allowing compatible local chat models to be swapped without discarding the conversation.
Document ingestion and retrieval
The ingestion pipeline works as follows:
- Extract text from a PDF, Markdown, or UTF-8 text file.
- Split it into passages of approximately 1,600 characters, with a 30-word overlap when splitting long paragraphs.
- Generate Nomic embeddings through Ollama.
- Store each passage, embedding, document identity, and page number in SQLite.
A SHA-256 digest detects duplicate document uploads. Scanned PDFs need OCR before import; this prototype extracts existing text.
For retrieval, I use Nomic’s search_document: prefix for passages and search_query: for questions. The backend performs an exact cosine-similarity scan over the eligible stored vectors.
Passages must pass a configurable similarity threshold. Their final ranking combines semantic similarity with keyword overlap:
score = 0.8 Ă— cosine_similarity + 0.2 Ă— keyword_overlap
The top five passages are supplied to the answering model. Documents can belong to a category or be shared across categories.
This keeps retrieval simple enough for a laptop prototype without requiring a separate vector database service.
Grounding answers in evidence
The response schema includes the answer text, its stated basis, a document-check result, and supporting quotations.
When the model cites a passage, the backend checks that:
- The source ID belongs to the retrieved sources.
- The quotation occurs in that source.
- The reference has not been invented.
The UI then shows the supporting document and page information.
These checks validate quotation provenance, not the truth of every conclusion. Invalid references produce an explicit notice. When documents are not relevant, Fieldmate can still answer from general model knowledge and label the response accordingly.
Conversation memory and contextual actions
I keep conversation memory separate from document retrieval.
Each thread has its own persistent transcript, category, model selection, and custom instructions. The backend loads history from SQLite rather than trusting conversation history submitted by the browser.
As a conversation grows, older exchanges are compacted into a rolling summary. The original transcript remains available for operations that need exact content.
That distinction became important when fixing a real bug:
“Save this as a note” initially saved the word “this.”
The capture handler was bypassing contextual resolution. I changed it so that:
- “Save this” copies the preceding answer, skipping action confirmations.
- “Save our study plan” searches the current thread’s stored messages for the relevant plan.
- Whole-message captures use schema-constrained source IDs, and the backend copies the original text.
- Requested excerpts must match the selected source.
- Missing or invalid selections ask for clarification without writing a note.
Notes and tasks are saved in the same SQLite transaction as their request receipt. Repeating the same request ID returns the existing result instead of creating another item.
The distinction between compact context and recoverable source text is also explored in Smaller Context, Recoverable History. The separate problem of choosing an outdated version is discussed in Agent Memory’s Real Failure Is Currency, Not Retrieval.
Clarification and internet permissions
I integrated Laya to provide an advisory signal about whether a request needs clarification. Structured Ollama planning combines that signal with conversation context.
I also added deterministic checks for clear missing references. For example, a new conversation asking “Can I take it with my other medicine?” needs the medicine names before a useful answer.
Web search follows a separate backend permission flow. Fieldmate proposes an editable query and waits for approval. The permission is tied to that conversation and can be consumed only once.
Only the approved query is sent to the search provider. The application does not automatically upload the full conversation or document collection. Returned search snippets are treated as external evidence, with their limitations retained.
The local voice pipeline
Voice input passes through faster-whisper, using the multilingual Whisper base model with CPU int8 inference. The transcription can be reviewed in the composer before sending.
For speech output, Piper generates a WAV response using the locally stored Lessac voice model.
Hands-free mode uses microphone volume and silence detection to identify turns. Listening pauses while the assistant processes or speaks, so the prototype uses turn-taking rather than simultaneous listening and speaking.
I also added a speech-cleanup layer after discovering that the voice was reading Markdown asterisks aloud. It removes formatting and citation markers while converting meaningful mathematical operators into spoken words.
Testing the complete workflow
I used Codex to help implement and debug Fieldmate.
The automated suite reached 61 passing tests, covering conversation isolation, contextual captures, retrieval behavior, citation checks, search permissions, task dates, retries, and transaction rollback.
I also tested against local models. With Gemma 4 12B, the contextual-save checks verified the original study-plan content, recovery after a topic change, selection of a newer revision, and creation of a task from a referenced step.
Those live checks caught behavior that test doubles missed—including the model trying to rewrite a quotation instead of selecting the stored text.
Why Does Open Innovation Matter?
Local operation is the reason this project works for its intended setting.
After the initial model downloads, the core assistant can run without an internet connection. Document retrieval, conversation, transcription, and speech generation happen on the laptop. Optional web search is the exception and requires approval.
That matters because notes and documents do not need to be sent to a hosted inference service just to ask a question about them.
Open tooling and locally runnable models also gave me control over the assistant’s behavior. I could inspect the decision flow, change retrieval logic, improve clarification, and fix the speech formatting problem.
The model is replaceable, too. Fieldmate can switch between compatible installed Ollama models while retaining the conversation. I tested a switch from Qwen to Gemma with a contextual follow-up.
There are no per-request cloud inference charges for the local workflow, although it still uses the laptop’s storage, memory, and electricity.
A hosted API would have made this prototype dependent on connectivity and external processing. This approach lets the assistant remain useful on the machine my friend already has.
My Agent Session
Prize Categories
Best Use of Gemma
Fieldmate runs Gemma 4 locally through Ollama for conversational assistance, structured decisions, and document-grounded answers. I used Gemma 4 E2B during development and tested the contextual capture workflow with Gemma 4 12B.
Gemma fits into a workflow where the user’s documents and conversation stay on their laptop.

Top comments (0)