This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend.
What I Built
I built Whisper AI & FFmpeg Media Transcriber, a web application that turns audio and video files into timestamped transcripts and subtitles.
I built it for friend, who regularly has to deal with recorded meetings / lectures / interview sessions.
The problem was simple:
They had useful information trapped inside audio and video, but getting that information into searchable, readable text was slow and annoying.
Instead of requiring them to manually listen through recordings and take notes, I built a tool that can process the media and produce a synchronized transcript.
The application supports both audio and video files. For video, FFmpeg extracts and normalizes the audio into the 16 kHz, 16-bit mono PCM format expected by Whisper before sending it through the speech-to-text pipeline.
It supports formats including MP4, MKV, MOV, WebM, AVI, MP3, WAV, M4A, FLAC, OGG, AAC, and more.
What it can do
- 🎥 Transcribe video files
- 🎙️ Transcribe audio files
- 🎤 Record directly from the browser microphone
- 🌍 Automatically detect speech across 99+ languages
- 🔄 Translate speech into English
- ⏱️ Generate timestamped transcripts
- ▶️ Synchronize transcript segments with media playback
- 🔎 Click a transcript timestamp to seek directly to that part of the recording
- 📝 Export transcripts as
.txt - 🎬 Export subtitles as
.srtand.vtt - 📦 Export detailed transcription data as
.json - 🎧 Export extracted audio as
.wav
The interface also highlights the current transcript segment while the media is playing, making it possible to follow the transcript along with the recording.
I also added preloaded demo media so someone can try the application without first finding an audio or video file.
Demo
Demo Video: Click
The basic flow is:
Upload/record → FFmpeg processing → Whisper transcription → timestamped transcript → synchronized playback/export
Code
GitHub: repo
The project is open source under the MIT License and is free to modify.
The project is structured around a FastAPI backend, FFmpeg media processing, Whisper inference, asynchronous task management, and a browser-based interface.
The main components include:
-
ffmpeg_utils.py— media probing and audio extraction -
whisper_service.py— Whisper inference, caching, and subtitle generation -
task_manager.py— asynchronous task state and Server-Sent Events -
main.py— API routes and media streaming -
app.js— client-side media synchronization and player logic -
index.html— the web dashboard
The full project structure is documented in the repository README.
How I Built It
The core AI component is OpenAI Whisper, an open-source speech recognition model.
I wanted the AI to be an actual part of the application rather than simply calling a closed transcription API.
The processing pipeline looks like this:
User Upload / Microphone
│
▼
Media Probe
FFprobe
│
┌───────────┴───────────┐
▼ ▼
Video Audio
│ │
▼ ▼
FFmpeg Extraction FFmpeg Normalization
│ │
└───────────┬───────────┘
▼
16 kHz Mono WAV
│
▼
Whisper Model
│
▼
Timestamped Transcription
│
┌─────────────┼─────────────┐
▼ ▼ ▼
Web Player Full Text SRT / VTT / JSON
The project uses FFmpeg to normalize the input audio and Whisper to perform speech recognition.
Whisper models
The application supports multiple Whisper model sizes:
-
tiny— optimized for speed -
base— balanced option -
small— higher accuracy -
medium— larger model for more demanding transcription
This gives the user a choice between processing speed and transcription quality.
Real-time progress
Long-running transcription shouldn't feel like a frozen webpage.
So I added Server-Sent Events to stream task progress back to the browser. The UI can show the different stages of the pipeline while processing is happening.
The backend exposes endpoints for starting transcription jobs, checking task status, streaming progress events, accessing media, exporting transcripts, and listing available Whisper models.
Why Does Open Innovation Matter?
This is probably the most important part of the project for me.
A transcription application can be built by sending every recording to a closed AI API.
But that would change the nature of the tool.
Audio and video can contain extremely personal information:
- private conversations
- meetings
- interviews
- lectures
- family recordings
- voice messages
- work-related information
With an open model such as Whisper, the application can be built around local inference instead of requiring every recording to leave the user's machine.
That gives the project a very different privacy model.
The user can run the application locally with Python, FFmpeg, and the required dependencies. The documented setup runs the application locally through Uvicorn on localhost:8000.
That also means the AI layer isn't locked to a single hosted API.
Because the model is available locally, the application can expose different model sizes and let users choose how they want to trade off speed and accuracy.
That's the part of open innovation that mattered most for this project:
the AI isn't just an external service that the application talks to. The AI model is part of the application itself.
Open-source software also made it possible to combine different pieces of the stack:
Whisper + FFmpeg + FastAPI + browser APIs + SSE
and build the exact workflow needed instead of adapting the project around the limitations of one closed API.
What I Learned
Building this project also made me think differently about AI applications.
The interesting part isn't always the model.
The model is only one component.
A useful AI product needs everything around it:
Input
↓
Media processing
↓
Model inference
↓
Task management
↓
Progress reporting
↓
Result formatting
↓
User interface
↓
Export / reuse
Whisper handles the speech recognition, but FFmpeg makes the media usable, the backend manages the processing pipeline, SSE communicates progress, and the frontend turns the result into something a person can actually use.
That combination is what turned an AI model into a real application.
Final Thoughts
This project started with a simple idea:
Take something a friend struggles with and remove the annoying part.
The result is a local AI transcription application that can take audio or video, process it through an open speech-recognition model, synchronize the resulting transcript with the original media, and export the result in formats that can actually be reused.
For me, that's the interesting thing about open AI.
It's not just about having access to a model.
It's about being able to take the model, understand how it fits into a larger system, change the surrounding workflow, run it locally, and build something that solves a real person's problem.
And this time, I built it for a friend.
Built with ❤️, Whisper, FFmpeg, Python, and open-source innovation.

Top comments (0)