DEV Community

Tamajit-005
Tamajit-005

Posted on

Amar Voice — Giving My Friend Their Voice Back (EN / हिंदी / বাংলা)

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

Amar Voice — A Voice That Stays Theirs.

Built for a friend who has lost their voice, but still has old voice recordings.

Amar Voice turns those recordings back into a voice they can use every day. Upload a few of their old voice notes → the app clones their voice → they tap a big button and the phone speaks as them again, in English, हिंदी, or বাংলা. Anything beyond the daily phrases can be typed into the custom-message box (up to 1000 chars).

Built for tired eyes and shaky hands: giant high-contrast cards, three language pills, one speaker menu, zero settings pages. It installs like an app (PWA), quick cards play instantly from cache, and family can add / rename / delete voice profiles right from the phone — no login, each device gets its own private space.

The 6 quick cards (every tap speaks in their voice)

Tap the feeling English हिंदी বাংলা
💧 Water — pleading, thirsty I need some water, please... मुझे थोड़ा पानी चाहिए... আমার একটু জল দরকার, দয়া করে...
💊 Medicine — routine, calm Time for my medicine... मेरी दवाई का समय हो गया है... আমার ওষুধ খাওয়ার সময় হয়েছে...
📞 Call — insistent, needs someone Please call Aniket right now! कृपया अनिकेत को बुलाओ! দয়া করে এখনই অনিকেতকে ডাকো!
☕ Tea — comfort, warmth I would like some warm tea. मुझे थोड़ी गर्म चाय चाहिए। আমাকে একটু গরম চা দাও।
😊 I'm fine — happy, reassuring I am doing fine, do not worry. मैं बिल्कुल ठीक हूँ, चिंता मत करो। আমি ভালো আছি, চিন্তা কোরো না।
🚨 Help — urgent, distressed I need help, please come here! मुझे मदद चाहिए, कृपया यहाँ आओ! আমার সাহায্য দরকার, দয়া করে এখানে এসো!

The emojis aren't decoration — they are the interface. 😊 means happy, 🚨 means urgent, 💧 means pleading. No reading needed: tap the feeling and the phone says the matching phrase in their own voice, with the feeling carried in the words themselves (... vs !) as rendered by the clone.

Demo

Live app: https://amar-voice.onrender.com

Video demo: https://youtu.be/DQC13FPReB4?si=tfCekrNp80R67fji

Tip for judges: open the link on your phone → Add to Homescreen → tap 💧 water → switch हिंदी / বাংলা pills → try custom text → open Upload to clone a voice from old recordings (max 8 clips, 800 KB each).

Code

https://github.com/Tamajit-005/Amar-voice

  • backend/main.py — routes, upload caps, rate limits
  • backend/engine.py — voice-ID map, clone / rename / delete
  • backend/eleven.py — ElevenLabs client
  • frontend/index.html — PWA UI
  • frontend/sw.js — offline cache
  • render.yaml + Dockerfile — deploy

How I Built It

FastAPI backend + single-file PWA frontend, deployed on Render free tier via Docker.

  • Voice: ElevenLabs v4 Instant Voice Cloning. Upload a few 10–25s clips from their old recordings → backend clones once, reuses that voice ID for all 3 languages. 18 presets pre-generated per speaker, cached on server + service worker for instant taps. Custom text synthesizes in a few seconds when warm (longer on a cold free-tier start), MP3 download capped at 350 KB.
  • Privacy without login: random amar_uid in localStorage, clips + voice-ID map + presets namespaced per device. default is shared, everything else is device-private.
  • Guardrails: 8 clips / 800 KB per clip / 2-min caps, per-IP rate limits on paid endpoints, atomic cache writes, honest 404s (never serves another voice on failure).
  • Open pieces: FastAPI, Pydantic, Uvicorn, pydub + ffmpeg, vanilla JS PWA with service worker, Docker — all open-source, and the repo is open for anyone to self-host with their own API key.
  • Problems faced: local models sounded distant from the real voice → moved to cloud, kept offline cache for presets; built for one friend but shipped per-device spaces so other families can reuse it; Render free idle wipes server voices (phone keeps ID, sees only default) — needs paid disk or external store.

Why Does Open Innovation Matter?

The parts that make this shippable and forkable are open: anyone can clone the repo, cp .env.example .env, deploy the same Dockerfile to Render, and have it running for their own family in minutes — no app-store review, no proprietary client, no accounts to manage.

That mattered practically: I swapped the entire voice backend (local → ElevenLabs) without touching the PWA, added per-device spaces with no auth service, and kept it runnable on a 512 MB free tier because ffmpeg/pydub/FastAPI are hackable. What I'd open next: a local-model fallback so custom speech works fully offline.

Prize Categories

  • Best Use of Render
  • Best Use of ElevenLabs

Top comments (0)