Speech-to-text models are very good now. The annoying part is everything around them: recordings that are too big to upload, video files, and files that live behind a Google Drive or Dropbox share link.
This post covers the do-it-yourself route first, then a one-call alternative I built for myself.
The do-it-yourself route
Short audio file
If you have a short MP3 and Python, open-source Whisper is enough:
pip install -U openai-whisper
whisper interview.mp3 --model small --output_format txt
Video file
Whisper wants audio. Pull the audio track out with ffmpeg first, and shrink it while you are there, since speech models work on 16 kHz mono anyway:
ffmpeg -i meeting.mp4 -vn -ac 1 -ar 16000 -b:a 48k meeting.mp3
A one-hour video becomes an MP3 of about 20 MB.
Long recordings and hosted APIs
Hosted speech-to-text APIs are much faster than a laptop, but they cap the upload size. OpenAI's transcription endpoint, for example, accepts files up to 25 MB. For a long recording you have to:
- split the audio into chunks,
- transcribe each chunk,
- shift every chunk's timestamps by where it started,
- join the text back together.
ffmpeg can do the split:
ffmpeg -i meeting.mp3 -f segment -segment_time 600 -c copy chunk_%03d.mp3
The rest is glue code, plus retries for when the API rate-limits you halfway through.
Share links
A Google Drive or Dropbox share link opens a preview page, not the file. You need the direct-download form:
-
Dropbox: change
dl=0todl=1at the end of the link. -
Google Drive: take the file ID from the link and use
https://drive.usercontent.google.com/download?id=FILE_ID&export=download&confirm=t.
None of this is hard. It is just a lot of small steps to maintain.
The one-call route
I packaged those steps into a tool: Audio & Video to Text on Apify. You give it file links, Drive or Dropbox share links, or an uploaded file, and it handles extraction, splitting, retries and timestamps. Transcription uses Whisper large-v3.
curl -X POST "https://api.apify.com/v2/acts/spokentext~audio-video-to-text/run-sync-get-dataset-items?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "urls": ["https://example.com/interview.mp3"], "includeSrt": true }'
The response is one JSON object per file:
{
"status": "ok",
"fileName": "interview.mp3",
"language": "English",
"durationSeconds": 1842,
"transcribedMinutes": 31,
"text": "Thanks for joining me today. Let's start with...",
"segments": [{ "start": 0, "end": 3.4, "text": "Thanks for joining me today." }],
"srt": "1\n00:00:00,000 --> 00:00:03,400\nThanks for joining me today.\n"
}
To process a batch from Python:
from apify_client import ApifyClient
client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("spokentext/audio-video-to-text").call(run_input={
"urls": [
"https://example.com/interview.mp3",
"https://www.dropbox.com/scl/fi/abc123/meeting.mp4?rlkey=xyz&dl=0",
],
})
for item in client.dataset(run.default_dataset_id).iterate_items():
print(item["fileName"], item["status"], len(item.get("text", "")))
From my own runs: a 50-minute MP3 shared through Dropbox was transcribed in about 40 seconds and cost $0.15. Pricing is $0.003 per audio minute, and files that fail are not charged.
What it does not do
- No YouTube or TikTok links. It takes files, not video pages.
- No speaker labels. You get what was said, not who said it.
- Your audio leaves your machine. It is sent to a hosted speech-to-text provider for processing. If the recording is sensitive, run Whisper locally instead.
Which route to pick
| Situation | Route |
|---|---|
| A few short files, privacy matters most | Whisper on your own machine |
| Long recordings, video, share links, or batches | The API |
| You need speaker labels | Neither; look for a tool with diarization |
Disclosure: I built the tool described in "The one-call route". This article was drafted with AI assistance and checked by me.
Top comments (0)