This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
My friend Denis doesn't speak much English, but there's a streamer he really enjoys watching.
He mostly discovers the streamer's clips and highlights on Instagram, but because they're in English, he can't always understand what's happening.
One day, he asked me:
"Could you make something that translates these videos for me?"
So I did.
I built Video Translater — an open-source AI-powered application that takes a video, understands the speech, translates it into another language, generates new speech, and synchronizes the translated audio back with the original video.
The idea is simple:
Take a video Denis can't understand → turn it into a video he can actually enjoy.
And that's how this project started.
Not with a business idea.
Not with a startup pitch.
Just a friend who wanted to watch his favorite streamer.
🎬 How It Works
Video translation sounds simple until you actually try to make the result feel like a real video.
The application turns the process into a pipeline:
Video
↓
Extract Audio
↓
Speech-to-Text
↓
Detect / Select Language
↓
Translate
↓
Text-to-Speech
↓
Synchronize Audio
↓
Render Final Video
↓
Translated Video 🎥
The important part is that the generated voice has to respect the original video's timing.
For example, imagine someone says something that lasts 4 seconds in English.
The translated sentence might take 6 seconds to say.
If we simply replace the audio, the translated voice will no longer match the video.
So the application compares the generated speech duration with the original timestamp and adjusts the audio when necessary.
That's where a simple "AI translation demo" turns into a real multimedia engineering problem.
🎥 Demo
Watch the demo:
👉 https://www.youtube.com/watch?v=4_827OMZHeY
The demo shows the complete workflow from the original video to the translated result.
The most important thing to see is not just the translation itself, but the final synchronized audio.
💻 Code
The entire project is open source:
👉 https://github.com/runpy21/video-translater
You can explore the complete implementation, run it yourself, open issues, and contribute improvements.
🧠 How I Built It
I built Video Translater around a modular processing pipeline.
The application separates video processing from the AI components so that individual parts can evolve independently.
The main pieces are:
- Python — application logic and orchestration
- FFmpeg — audio/video extraction, processing and rendering
- yt-dlp — video downloading
- Speech recognition — converting speech into text
- Translation — converting the transcript into the target language
- Text-to-speech — generating the translated voice
- Audio synchronization — fitting generated speech back into the original timeline
I wanted the AI components to behave like replaceable building blocks rather than creating one giant piece of tightly coupled code.
That makes it much easier to experiment with different models and approaches.
⏱️ The Interesting Engineering Problem
The hardest part wasn't actually translating the words.
It was making the translated speech fit the video.
Suppose the original speech segment is:
|--------- 4 seconds ---------|
But the translated speech takes:
|-------------- 6 seconds --------------|
Simply cutting the audio would make the sentence incomplete.
Simply letting it play would make it overlap with the next segment.
So the pipeline needs to reason about the timeline:
Original timestamps
↓
Translated text
↓
Generated speech
↓
Measure duration
↓
Compare with available time
↓
Adjust when necessary
↓
Place back on timeline
This was one of my favorite parts of the project because it required combining AI output with traditional audio/video engineering.
The AI generates the content.
The rest of the system has to make that content actually work.
🎯 Built for a Real Problem
One of the things I liked most about this project is that I didn't have to invent a use case.
Denis actually asked for it.
He didn't ask for an AI demo. He didn't care which model I used or how the pipeline was implemented.
He just wanted to understand the videos of a streamer he likes.
That gave me a very clear goal:
If Denis can watch the translated video and understand what is happening, the project works.
This also influenced how I built the application.
I focused on the complete experience rather than just getting a translation from an AI model:
video → speech → translation → generated voice → synchronized video
The AI is only one part of that pipeline.
The real challenge was turning the AI output into something that can actually be watched.
🌍 Why Open Innovation Matters
This project is a good example of why I think open innovation matters.
A closed API can make it easy to call an AI model.
But an open ecosystem gives developers much more freedom.
I can experiment with different components.
I can replace a model.
I can run parts of the pipeline locally.
I can inspect how everything works.
I can change the synchronization logic.
I can optimize the application for a specific use case.
And most importantly:
someone else can take this project and improve it.
Video Translater isn't built around one magical AI endpoint.
It's built around a pipeline of components.
That means when better open models become available, the project doesn't have to start from zero.
The building blocks can evolve.
That's what open innovation means to me:
Don't just give developers an answer. Give them building blocks they can turn into something new.
❤️ Why I Built It
This project is personal because the original feature request came from a real person.
Denis didn't ask me to build an AI project.
He didn't care which model I used.
He didn't care about the architecture.
He just wanted to understand what his favorite streamer was saying.
That changed the way I approached the project.
Instead of asking:
"What impressive AI demo can I build?"
I asked:
"What can I build that would actually make someone's life a little easier?"
And that made the project much more fun to work on.
Every feature had a reason.
Translation wasn't enough.
The audio needed to work.
The timing needed to work.
The downloaded video needed to be handled safely.
The whole pipeline needed to produce something Denis could actually watch.
🚀 What's Next?
This is still the beginning.
There are a lot of things I'd like to add:
- 🎙️ Better voice preservation
- 👥 Multiple-speaker detection
- 🗣️ Better speech prosody
- 📝 Automatic subtitles
- 🌍 Multiple target languages
- 📊 Better processing progress
- ⚡ Improved error recovery
- 🐳 Easier deployment
- 🔐 Authentication and rate limiting
- ☁️ Scalable processing
And I'd love to see other developers take the project in directions I haven't thought of.
🏗️ What I Learned
I started this project thinking I was building a video translation tool.
I ended up learning that AI is only one part of the problem.
A model can generate a translation.
But then you still have to ask:
Does the audio fit the timeline?
Can the video be processed safely?
What happens when something fails?
Can another developer understand the code?
Can the components be replaced?
Does the final result actually solve the user's problem?
That's probably the biggest lesson I took away from this Hacktoberfest project:
Getting AI to produce an output is easy compared to making that output useful.
🏆 Prize Categories
Overall Hacktoberfest Weekend Challenge
🙌 Final Thoughts
A few hours ago, this was just a request from a friend.
Now it's an open-source project.
That's one of the things I love about building software.
A small problem experienced by one person can become an idea that other people can use, improve and build on.
Denis wanted to understand his favorite streamer.
So I built a translator.
Maybe someday, someone else will use it to understand a documentary, a tutorial, a lecture, or a creator from another part of the world.
One video. Different language. Same experience. 🌍
If you find the project useful, feel free to:
⭐ Star the repository
🐛 Open an issue
💡 Suggest an improvement
🔧 Contribute
GitHub: https://github.com/runpy21/video-translater
Thanks for reading — and happy Hacktoberfest! 🚀
Top comments (0)