This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
Amar Voice — A Voice That Stays Theirs.
Built for a friend who has lost their voice, but still has old voice recordings.
Amar Voice turns those recordings back into a voice they can use every day. Upload a few of their old voice notes → the app clones their voice → they tap a big button and the phone speaks as them again, in English, हिंदी, or বাংলা. Anything beyond the daily phrases can be typed into the custom-message box (up to 1000 chars).
Built for tired eyes and shaky hands: giant high-contrast cards, three language pills, one speaker menu, zero settings pages. It installs like an app (PWA), quick cards play instantly from cache, and family can add / rename / delete voice profiles right from the phone — no login, each device gets its own private space.
The 6 quick cards (every tap speaks in their voice)
| Tap the feeling | English | हिंदी | বাংলা |
|---|---|---|---|
| 💧 Water — pleading, thirsty | I need some water, please... | मुझे थोड़ा पानी चाहिए... | আমার একটু জল দরকার, দয়া করে... |
| 💊 Medicine — routine, calm | Time for my medicine... | मेरी दवाई का समय हो गया है... | আমার ওষুধ খাওয়ার সময় হয়েছে... |
| 📞 Call — insistent, needs someone | Please call Aniket right now! | कृपया अनिकेत को बुलाओ! | দয়া করে এখনই অনিকেতকে ডাকো! |
| ☕ Tea — comfort, warmth | I would like some warm tea. | मुझे थोड़ी गर्म चाय चाहिए। | আমাকে একটু গরম চা দাও। |
| 😊 I'm fine — happy, reassuring | I am doing fine, do not worry. | मैं बिल्कुल ठीक हूँ, चिंता मत करो। | আমি ভালো আছি, চিন্তা কোরো না। |
| 🚨 Help — urgent, distressed | I need help, please come here! | मुझे मदद चाहिए, कृपया यहाँ आओ! | আমার সাহায্য দরকার, দয়া করে এখানে এসো! |
The emojis aren't decoration — they are the interface. 😊 means happy, 🚨 means urgent, 💧 means pleading. No reading needed: tap the feeling and the phone says the matching phrase in their own voice, with the feeling carried in the words themselves (... vs !) as rendered by the clone.
Demo
Live app: https://amar-voice.onrender.com
Video demo: https://youtu.be/DQC13FPReB4?si=tfCekrNp80R67fji
Tip for judges: open the link on your phone → Add to Homescreen → tap 💧 water → switch हिंदी / বাংলা pills → try custom text → open Upload to clone a voice from old recordings (max 8 clips, 800 KB each).
Code
https://github.com/Tamajit-005/Amar-voice
-
backend/main.py— routes, upload caps, rate limits -
backend/engine.py— voice-ID map, clone / rename / delete -
backend/eleven.py— ElevenLabs client -
frontend/index.html— PWA UI -
frontend/sw.js— offline cache -
render.yaml+Dockerfile— deploy
How I Built It
FastAPI backend + single-file PWA frontend, deployed on Render free tier via Docker.
- Voice: ElevenLabs v4 Instant Voice Cloning. Upload a few 10–25s clips from their old recordings → backend clones once, reuses that voice ID for all 3 languages. 18 presets pre-generated per speaker, cached on server + service worker for instant taps. Custom text synthesizes in a few seconds when warm (longer on a cold free-tier start), MP3 download capped at 350 KB.
- Privacy without login: random
amar_uidinlocalStorage, clips + voice-ID map + presets namespaced per device.defaultis shared, everything else is device-private. - Guardrails: 8 clips / 800 KB per clip / 2-min caps, per-IP rate limits on paid endpoints, atomic cache writes, honest 404s (never serves another voice on failure).
- Open pieces: FastAPI, Pydantic, Uvicorn, pydub + ffmpeg, vanilla JS PWA with service worker, Docker — all open-source, and the repo is open for anyone to self-host with their own API key.
- Problems faced: local models sounded distant from the real voice → moved to cloud, kept offline cache for presets; built for one friend but shipped per-device spaces so other families can reuse it; Render free idle wipes server voices (phone keeps ID, sees only
default) — needs paid disk or external store.
Why Does Open Innovation Matter?
The parts that make this shippable and forkable are open: anyone can clone the repo, cp .env.example .env, deploy the same Dockerfile to Render, and have it running for their own family in minutes — no app-store review, no proprietary client, no accounts to manage.
That mattered practically: I swapped the entire voice backend (local → ElevenLabs) without touching the PWA, added per-device spaces with no auth service, and kept it runnable on a 512 MB free tier because ffmpeg/pydub/FastAPI are hackable. What I'd open next: a local-model fallback so custom speech works fully offline.
Prize Categories
- Best Use of Render
- Best Use of ElevenLabs
Top comments (0)