DEV Community

clement-melkior
clement-melkior

Posted on

How to transcribe audio and video files to text with one API call (MP3, MP4, Google Drive, Dropbox)

Speech-to-text models are very good now. The annoying part is everything around them: recordings that are too big to upload, video files, and files that live behind a Google Drive or Dropbox share link.

This post covers the do-it-yourself route first, then a one-call alternative I built for myself.

The do-it-yourself route

Short audio file

If you have a short MP3 and Python, open-source Whisper is enough:

pip install -U openai-whisper
whisper interview.mp3 --model small --output_format txt
Enter fullscreen mode Exit fullscreen mode

Video file

Whisper wants audio. Pull the audio track out with ffmpeg first, and shrink it while you are there, since speech models work on 16 kHz mono anyway:

ffmpeg -i meeting.mp4 -vn -ac 1 -ar 16000 -b:a 48k meeting.mp3
Enter fullscreen mode Exit fullscreen mode

A one-hour video becomes an MP3 of about 20 MB.

Long recordings and hosted APIs

Hosted speech-to-text APIs are much faster than a laptop, but they cap the upload size. OpenAI's transcription endpoint, for example, accepts files up to 25 MB. For a long recording you have to:

  1. split the audio into chunks,
  2. transcribe each chunk,
  3. shift every chunk's timestamps by where it started,
  4. join the text back together.

ffmpeg can do the split:

ffmpeg -i meeting.mp3 -f segment -segment_time 600 -c copy chunk_%03d.mp3
Enter fullscreen mode Exit fullscreen mode

The rest is glue code, plus retries for when the API rate-limits you halfway through.

Share links

A Google Drive or Dropbox share link opens a preview page, not the file. You need the direct-download form:

  • Dropbox: change dl=0 to dl=1 at the end of the link.
  • Google Drive: take the file ID from the link and use https://drive.usercontent.google.com/download?id=FILE_ID&export=download&confirm=t.

None of this is hard. It is just a lot of small steps to maintain.

The one-call route

I packaged those steps into a tool: Audio & Video to Text on Apify. You give it file links, Drive or Dropbox share links, or an uploaded file, and it handles extraction, splitting, retries and timestamps. Transcription uses Whisper large-v3.

curl -X POST "https://api.apify.com/v2/acts/spokentext~audio-video-to-text/run-sync-get-dataset-items?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{ "urls": ["https://example.com/interview.mp3"], "includeSrt": true }'
Enter fullscreen mode Exit fullscreen mode

The response is one JSON object per file:

{
    "status": "ok",
    "fileName": "interview.mp3",
    "language": "English",
    "durationSeconds": 1842,
    "transcribedMinutes": 31,
    "text": "Thanks for joining me today. Let's start with...",
    "segments": [{ "start": 0, "end": 3.4, "text": "Thanks for joining me today." }],
    "srt": "1\n00:00:00,000 --> 00:00:03,400\nThanks for joining me today.\n"
}
Enter fullscreen mode Exit fullscreen mode

To process a batch from Python:

from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")

run = client.actor("spokentext/audio-video-to-text").call(run_input={
    "urls": [
        "https://example.com/interview.mp3",
        "https://www.dropbox.com/scl/fi/abc123/meeting.mp4?rlkey=xyz&dl=0",
    ],
})

for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item["fileName"], item["status"], len(item.get("text", "")))
Enter fullscreen mode Exit fullscreen mode

From my own runs: a 50-minute MP3 shared through Dropbox was transcribed in about 40 seconds and cost $0.15. Pricing is $0.003 per audio minute, and files that fail are not charged.

What it does not do

  • No YouTube or TikTok links. It takes files, not video pages.
  • No speaker labels. You get what was said, not who said it.
  • Your audio leaves your machine. It is sent to a hosted speech-to-text provider for processing. If the recording is sensitive, run Whisper locally instead.

Which route to pick

Situation Route
A few short files, privacy matters most Whisper on your own machine
Long recordings, video, share links, or batches The API
You need speaker labels Neither; look for a tool with diarization

Disclosure: I built the tool described in "The one-call route". This article was drafted with AI assistance and checked by me.

Top comments (0)