DEV Community

Pranesh Nikhar
Pranesh Nikhar

Posted on

I built a voice language tutor that lives on my friend's laptop

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

A friend of mine is learning Spanish. They're not bad at it — they're scared of it. Practicing with a real person means being wrong out loud, in front of another human, and that freezes them. Language apps feel like flashcards with a timer, not a conversation. And a tutor costs money they don't have.

So I built Habla Conmigo — a patient voice tutor that lives on their own laptop. You talk to it like a person: tap the mic, say whatever you want in Spanish (or half-Spanish, half-English, or completely stuck English), and it replies out loud, gently correcting mistakes and keeping the conversation going.

It solves the exact problem my friend had: a practice partner with zero judgment, infinite patience, and a memory that follows the conversation across turns.

Demo

Live (deployed, UI + API — the model itself runs locally, by design): habla-conmigo.onrender.com

Code

github.com/praneshnikhar/hacktoberfest

How I Built It

Three layers, one conversation loop:

  1. Browser mic → ffmpeg — the frontend records a webm blob; the server converts it to mp3 with ffmpeg.
  2. ElevenLabs Scribe (STT) — transcribes the audio, auto-detecting language so a stuck student can fall back to English mid-sentence.
  3. Open-weight model via Ollama (the tutor brain) — the reply comes from a model running locally on the laptop (Llama 3.2 / Gemma 3), with a system prompt that makes it a tutor: reply in the target language, keep it to 1–3 sentences, correct one mistake at a time, always end with a follow-up question.
  4. ElevenLabs TTS — speaks the reply back in a warm multilingual voice.

Conversation memory (last 8 turns) lives in the server, so the tutor follows topics across turns instead of starting fresh every sentence.

The whole thing is ~700 lines of plain Node.js + Express + vanilla JS. No framework, no database, no cloud AI.

Why Does Open Innovation Matter?

This is the part I care about most, because the open pieces are why the app exists at all:

  • A nervous learner's words never leave their machine. The core of this app — the part that understands and replies — is an open-weight model running locally. A closed API would ship every halting, half-grammatical sentence of a shy person to a third-party server, and bill per message. Here it's free and private, and it even works offline.
  • No lock-in, swap any model in one line. Because the brain is Ollama, changing the teacher is a one-line .env change: Llama → Gemma → Qwen → whatever ships next. Try doing that with a closed model. This also means the tutor can be fine-tuned later for my friend's specific mistakes.
  • Open where it matters, closed only where it adds value. The one closed service (ElevenLabs) does one narrow thing — voice — and even that degrades gracefully: remove the API key and the app still works as a text tutor. The product doesn't collapse when the proprietary part is taken away.
  • It costs nothing to run. The model is free, the runtime is free, the voice tier has a free plan. "Practicing a language" shouldn't have a subscription.

The counterfactual is what convinced me: with a closed model, I would have needed a paid API, an account per user, a data policy, and a conversation my friend couldn't afford to have.

My Agent Session

Built end-to-end with the opencode agent harness (open-source CLI coding agent) driving the implementation, debugging, and this write-up.

Prize Categories

  • Best Use of ElevenLabs — an open-source agent given a voice: Scribe STT + multilingual TTS power the full conversational loop
  • Best Use of Render — deployed via the included render.yaml blueprint

Tags: #devchallenge #weekendchallenge #hf26challenge

Top comments (0)