DEV Community

Cover image for Resona - AI Voice Generation
Himanshu Verkiya
Himanshu Verkiya

Posted on

Resona - AI Voice Generation

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

I built Resona, an AI voice generation and cloning platform for a friend who wanted to create manhwa and manga recap videos on YouTube.

Recording narration for every video was time-consuming, so I built Resona to make it easier to generate consistent voiceovers from scripts, create custom voices, and manage generated audio in one place.

Demo

Live Demo: https://resonapro.vercel.app/

Brief Walkthrough: https://youtu.be/Bt2X_5rsFO0

Quick Demo: https://youtu.be/h0urRp9gXrU

Code

GitHub: https://github.com/verkiya/Resona

How I Built It

Resona is built around Chatterbox TTS, an open-source text-to-speech model that I self-host on Modal with a FastAPI inference service.

The platform uses Next.js, React, TypeScript, tRPC, Prisma, PostgreSQL, Clerk, AWS S3, Polar, Sentry, WaveSurfer.js, and RecordRTC.

It supports text-to-speech generation, custom voice cloning, private audio storage, organization workspaces, and usage-based billing.

Why Does Open Innovation Matter?

Using open-source Chatterbox TTS let me own the voice-generation layer instead of relying entirely on a closed API.

I could self-host the model, control the inference infrastructure, and build the rest of the platform around it. It also keeps the model layer flexible, so it can be changed or improved independently from the application.

Prize Categories

Best Use of ElevenLabs

I used ElevenLabs to generate the narration for the Resona demo video.

Top comments (0)