The problem was mundane. On one side, meeting recordings: an mp3 of an hour or an hour and a half, out of which I needed text and subtitles. On the other, the reverse operation: I have text, I need a voice, because some clients listen rather than read. I did not want to buy a separate SaaS for this — the recordings are internal, and paying a subscription for them costs more than paying for the actual minutes.
The constraints I set for myself up front:
- one VPS, no external queues or brokers;
- no dependencies beyond the system
ffmpeg/ffprobe— so that an update means replacing one file and restarting a systemd unit; - recognition and synthesis go straight to Yandex SpeechKit v3, with no middleware;
- user files are not kept longer than a day;
- cost has to be predictable: the service must not run at a loss because someone else's script uploads hour-long recordings all night.
What came out is a single Python 3.12 file of roughly 2,200 lines: http.server with threads, sqlite3 as the job queue, urllib for everything external, ffmpeg for conversion. Below is what turned out to be non-obvious in this construction. Not a tutorial — a field notes collection of the places it bites.
Why the standard library was enough
The stock ThreadingHTTPServer cannot cap the number of concurrent connections: on a slow client it happily spawns thread after thread until file descriptors run out. I needed my own subclass — with a pool of 32 connection slots and a 120-second request read timeout. Everything else is ordinary routing: file upload via multipart, a JSON API for job status, result delivery, and the payment provider's webhook.
The queue is hand-rolled too, and it is surprisingly boring: a jobs table in SQLite, two worker threads, the statuses queued → processing → done|error, and a stage field holding a human-readable string like "converting" or "recognizing". Here sqlite3 works as an ordinary job log rather than a high-load database: one row written per job, one row read per status poll.
What that bought in practice: deploying the service is scp of one file plus systemctl restart. No pip install, no migrations, no "your library version is wrong".
The public REST API is documented at bowhard.ru/api, and since I am not the only one who wanted to drive this from a script, there are small clients for Python and Node — that is how the transcription and text-to-speech services at bowhard.ru/rasshifrovka and bowhard.ru/ozvuchka are consumed from automation.
Recognition: asynchronous, because the file is long
SpeechKit can do synchronous recognition, but that is for short utterances. An hour-long recording means recognizeFileAsync: send a job, get an operationId, poll the status until it is done. Poll every two seconds, with a three-hour deadline. Going below two seconds is pointless — the operation cannot have finished anyway, and you are just hammering the API for nothing.
The input to SpeechKit comes in two flavours, and the choice between them is not cosmetic:
-
content— audio as base64 inside the request body. Fine for small files, but base64 inflates the size by about a third, and holding an encoded hour-long recording in memory is not a great idea; -
uri— a link to an object in S3-compatible storage. We adopted a rule: files from 8 MB up go to Object Storage, anything smaller goes as base64. The threshold was picked from process memory: 8 MB in base64 is roughly an 11 MB string, which is perfectly safe for two workers.
The result of getRecognition is not JSON but a stream of JSON lines, one object per line — that is the first non-obvious thing. It contains both the plain recognized words and a finalRefinement.normalizedText.alternatives block. The words are needed for timings, but they carry no punctuation and no capital letters. The finished text with punctuation lives in the text field of that same normalizedText block. So we take the text from there, the timings from the word array, and splice the two into a single result.
Subtitles are assembled from that same result. SRT is a format that forgives a lot except two things: over-long lines and overlapping timecodes. We break the line by words, never letting it exceed 42 characters, and wrap only at a word boundary. A subtitle block looks like this:
- the block number line —
12; - the time line —
00:01:04,320 --> 00:01:07,880, with milliseconds separated by a comma, as the format requires; - one or two lines of text, split by words and no longer than 42 characters each —
and we decided not to touch that part.
The timecodes come straight from the words, so the subtitles end up denser than when you slice by sentences, and they do not drift on fast speech. The user gets three files: the text, the .srt, and a json with the words and their timings — the last one is for anyone who wants word highlighting in a player.
Text-to-speech: 240 characters that nearly broke everything
Synthesis in v3 is called through the utteranceSynthesis method, and it has a limit on text length that the documentation says approximately nothing about. It emerged empirically: 250 characters pass, 500 come back as a refusal reading "Too long text". And the refusal does not always arrive straight away: sometimes it looks like a dropped connection, and had I not been logging the response body, I would have hunted for the cause for a long time.
The fix is to cut the text into chunks of 240 characters, and to do it on sentence boundaries rather than by a character counter. It is inaudible, and the price is computed from the source character count, so the splitting does not affect the cost at all.
Then comes the pleasant part: MP3 is a sequence of independent frames, so the synthesized chunks are joined by plain byte concatenation, with no re-encoding and no ffmpeg. But if you cut the text crudely, mid-sentence, you can hear an unnatural pause at the seam — which is why the splitting happens at periods, question marks and exclamation marks.
Two more things about synthesis. First: the v1 API costs roughly twice as much as v3 for the same result, so you should go straight to v3. Second: the speed parameter is accepted in the range 0.5–2.0, and anything outside it has to be clamped on your side, otherwise you get an error in the middle of a long text. SpeechKit has fifteen voices — from the neutral "Alena" to the conversational "Zakhar" — and for internal recordings the difference between them is far more noticeable than for advertising ones.
The queue, the ceilings, and the money fuse
The nastiest class of bug in a service like this is not a crash but quietly burning money. So the ceilings are set at every level:
- no more than three jobs in flight per visitor;
- no more than 600 MB of uploads per day from one address;
- no more than 10 link imports per hour;
- the job directory cannot exceed 6 GB — past that, new jobs are refused until cleanup runs;
- and the main fuse: a daily cost ceiling. If the total recognition cost for the day exceeds a configured threshold, the service stops accepting new jobs and says so plainly, instead of running at a loss until morning.
A separate trap in the same area is balance accounting. Transcription is measured in minutes of audio, text-to-speech in characters. It is natural to keep both in one credits table, but if you compute usage as one combined sum you get nonsense: minutes get eaten by synthesis, characters by transcription, and the balance goes negative out of nowhere. Usage has to be counted separately per kind of work, and the interface has to show two independent balances. In the code this story survives as a comment — as a reminder.
Link import: SSRF is not only about the URL
A "paste a link to a file" feature saves the user time and creates a whole class of risk. Ours is narrowed to two providers: public Yandex Disk files (via cloud-api.yandex.net) and Dropbox. The list of trusted hosts is an allowlist, not a denylist, and the check is repeated after redirects: a provider may send you to another domain, and that is fine, but only if that domain is on the list too.
Timeouts are split as well: 30 seconds for the provider's response and 600 seconds for the download itself, otherwise the service turns into a free relay for other people's slow files. Plus a per-address attempt counter. A dedicated self-test command exercises these scenarios: selftest_links, selftest_sign, selftest_net — three buttons that ask "check that the outside world is still what we think it is".
Payments: we do not trust the webhook
Payment goes through YooKassa, and there is exactly one rule here that must not be broken: a webhook is not confirmation of a payment. It is only a reason to re-check. The scheme is this: the payment is created in the database up front with a "created" status; on notification the service makes a separate GET /payments/{id} request to the payment provider's API, looks at the real status, and only then credits the balance.
The crediting itself must happen exactly once, because webhooks repeat. Instead of "read the flag, write the credit", a conditional update is used: UPDATE payments SET paid_at = ? WHERE id = ? AND paid_at IS NULL, and the credit is granted only if that statement actually changed a row. This is cheaper and more reliable than any lock in the application: the race between two simultaneous webhooks is settled by the database.
Storage: signing SigV4 by hand
Files from 8 MB up go to S3-compatible Object Storage. I did not want to pull in an SDK for four operations (PUT, GET, DELETE, HEAD), so the AWS SigV4 signature is assembled by hand on urllib — and that turned out to be the most capricious part of the whole project.
Two traps, both diagnosed equally unpleasantly — just a 403 with no explanation:
-
x-amz-content-sha256has to appear both in the canonical headers and in theSignedHeaderslist. The header is sent, the hash is computed, and the signature still does not match, because it is missing from the list being signed. - Sub-resources such as
?lifecycleare written in the canonical request aslifecycle=— with an empty value. Leave them as they are and you getSignatureDoesNotMatchon a request that looks externally flawless.
Once the signature is assembled, though, everything else is trivial: put the object — hand SpeechKit the link — delete the object after cleanup. Job files live for 24 hours: that promise is written on the service page, which means the background cleanup has to honour it literally rather than "eventually". Statistics records are kept for 90 days — enough to understand load and compute cost.
What I would do differently
The SpeechKit client should have been extracted into a separate module from week one: right now the recognition, synthesis and queue logic all live in one file, and that is convenient exactly until the moment you want to run synthesis from a separate script. Second — progress for long operations: today it shows the stage and a speed estimate based on already-processed jobs, but an honest completion fraction for an hour-long recording would be more useful. Third — move the free-tier limits into config rather than code: those are the ones you want to change more often than anything else.
Otherwise the construction "one file, stdlib, SQLite and system ffmpeg" has held up better than I expected: the service handles recordings up to four hours, recognizing an hour-long file takes less time than the recording itself lasts, and updating it requires neither a build nor dependencies.
If you want to drive this from an automation platform rather than from code, there are ready-made n8n workflow templates for both transcription and synthesis. And if you build something similar, the most expensive time will go not into SpeechKit but into the S3 signature and the balance accounting — start with those.
Top comments (1)
tr.ee/dev-to