<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Bowhard</title>
    <description>The latest articles on DEV Community by Bowhard (@bowhard).</description>
    <link>https://dev.to/bowhard</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4176022%2Fd18bf5e3-e112-4a85-9f9a-b5054dcb8e2e.png</url>
      <title>DEV Community: Bowhard</title>
      <link>https://dev.to/bowhard</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/bowhard"/>
    <language>en</language>
    <item>
      <title>Publishing community templates to the n8n library: what the review actually asks for</title>
      <dc:creator>Bowhard</dc:creator>
      <pubDate>Sat, 10 Oct 2026 22:47:55 +0000</pubDate>
      <link>https://dev.to/bowhard/publishing-community-templates-to-the-n8n-library-what-the-review-actually-asks-for-4878</link>
      <guid>https://dev.to/bowhard/publishing-community-templates-to-the-n8n-library-what-the-review-actually-asks-for-4878</guid>
      <description>&lt;p&gt;The n8n template library is one of the few places where a self-hosted automation builder can publish something for free and get real traffic from it. It is also a gatekept library: every submission is reviewed by a human, and the guidelines are stricter than the form suggests. We have submitted templates from a self-hosted instance, and this is the checklist we work from — what the review actually looks at, in the order it tends to matter.&lt;/p&gt;

&lt;p&gt;The submission itself happens at &lt;a href="https://creators.n8n.io/hub" rel="noopener noreferrer"&gt;creators.n8n.io&lt;/a&gt;, not in the n8n app. You register a creator account, verify the e-mail, and submit from the dashboard. Everything below is drawn from the official &lt;a href="https://n8n.notion.site/p/Template-submission-guidelines-9959894476734da3b402c90b124b1f77" rel="noopener noreferrer"&gt;template submission guidelines&lt;/a&gt; and the separate &lt;a href="https://n8n.notion.site/Sticky-note-guidelines-for-templates-2aa5b6e0c94f8058b0aefddd02655887" rel="noopener noreferrer"&gt;sticky-note guidelines&lt;/a&gt;; if a rule here contradicts what you read there, trust them, not this post.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one-template rule that surprises everyone
&lt;/h2&gt;

&lt;p&gt;A new author (an account that has not been verified yet) may have &lt;strong&gt;exactly one template under review at a time&lt;/strong&gt;. The button for the next submission stays disabled until the previous one is approved or sent back. There is no queue, no batching, and no way to speed it up from your side.&lt;/p&gt;

&lt;p&gt;That single rule changes how you plan a set of templates. If you have three workflows ready, do not polish all three in parallel and hope to ship them in a week: you ship one, wait for review, then start the next. It also means the first submission is disproportionately expensive — a rejection costs you a full review cycle, so it is worth spending the extra hour on the sticky notes.&lt;/p&gt;

&lt;p&gt;After &lt;strong&gt;three approved templates&lt;/strong&gt; the account becomes verified. Verified authors get a badge, batches of up to four submissions instead of one, access to the closed n8n Discord, and the right to publish paid templates. Getting to three is the actual goal; the individual templates are the mechanism.&lt;/p&gt;

&lt;h2&gt;
  
  
  Titles: "action verb + object + where or from"
&lt;/h2&gt;

&lt;p&gt;The title formula is not a style preference, it is a review criterion. &lt;code&gt;Send website form leads to Google Sheets and Telegram&lt;/code&gt; is accepted shape: verb (&lt;code&gt;Send&lt;/code&gt;), object (&lt;code&gt;website form leads&lt;/code&gt;), destinations (&lt;code&gt;to Google Sheets and Telegram&lt;/code&gt;). &lt;code&gt;Ultimate AI Automation Beast 🚀&lt;/code&gt; is not, for three separate reasons.&lt;/p&gt;

&lt;p&gt;Practical consequences:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No emoji and no hype words.&lt;/strong&gt; "Powerful", "ultimate", "revolutionary", "10x" all read as marketing rather than description.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Name the services concretely.&lt;/strong&gt; "Your CRM" is vague; if the workflow talks to HubSpot, say HubSpot. Reviewers are checking that the title describes what the JSON actually does.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep it under the limit and check the rendered length.&lt;/strong&gt; The library truncates on cards, and a truncated title loses the "where" half of the formula.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Our own three titles, for reference:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;code&gt;Send website form leads to Google Sheets and Telegram&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;Send a daily sales report from your CRM to Telegram&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;Answer customer questions from your knowledge base with an AI model&lt;/code&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The description is a 200-word structured document
&lt;/h2&gt;

&lt;p&gt;The description is English markdown, roughly 200 words, and it is expected to have &lt;em&gt;sections&lt;/em&gt;, not a paragraph. The shape that reviewers look for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Who it is for&lt;/strong&gt; — the reader's job, not the workflow's features.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;How it works&lt;/strong&gt; — the steps in order, in prose.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;How to set up&lt;/strong&gt; — what to click, in what order, after import.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What you need&lt;/strong&gt; — the services, credentials and plan requirements.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;How to extend&lt;/strong&gt; — an honest pointer at the part that is designed to be edited.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;No HTML. Emoji are tolerated in the body far better than in the title, but the markdown has to render cleanly in the library's own renderer, which is stricter than a GitHub README — nested tables and raw HTML tend to come out mangled.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sticky notes are mandatory, and one of them is not a note
&lt;/h2&gt;

&lt;p&gt;This is the requirement people skip, and it is the one that gets a template sent back. Sticky notes in the workflow are &lt;strong&gt;not optional decoration&lt;/strong&gt;: they are how the reviewer (and the person who imports your workflow six months from now) understands it without opening a single node.&lt;/p&gt;

&lt;p&gt;The rule has two halves:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;One yellow note carries the whole description.&lt;/strong&gt; Not a summary — the full text that also goes into the submission form. The yellow note is the canonical copy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Every other note is a neutral, step-level annotation&lt;/strong&gt; — "Validate the e-mail before writing to the sheet", "Retry once on 429". They describe the &lt;em&gt;step&lt;/em&gt;, not the &lt;em&gt;template&lt;/em&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you keep the description in two places, they will diverge within a week. In our repository the submission description is generated by a build script that reads the description straight out of the workflow's own yellow sticky note, so the form and the JSON cannot disagree. If you write your own tooling, do it in that direction: JSON is the source, the form is the copy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Node names are read as documentation
&lt;/h2&gt;

&lt;p&gt;The default n8n names (&lt;code&gt;Set&lt;/code&gt;, &lt;code&gt;HTTP Request&lt;/code&gt;, &lt;code&gt;IF&lt;/code&gt;) tell a reader nothing. Renaming is expected, and it is checked: a workflow where the AI branch is labelled &lt;code&gt;Normalize lead fields&lt;/code&gt; and &lt;code&gt;Score against ICP rules&lt;/code&gt; is reviewable in a screenshot; one with four &lt;code&gt;Set&lt;/code&gt; nodes in a row is not.&lt;/p&gt;

&lt;p&gt;The same logic applies to sticky notes that arrange the canvas into sections — they should name the &lt;em&gt;stage&lt;/em&gt; of the pipeline, not repeat the node names.&lt;/p&gt;

&lt;h2&gt;
  
  
  Nothing secret, nothing personal, nothing live
&lt;/h2&gt;

&lt;p&gt;The exported JSON is public the moment it is approved, so review looks for anything that leaks or breaks someone else's instance:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No API keys or tokens anywhere in the workflow&lt;/strong&gt;, including inside HTTP Request node headers, query parameters and JSON bodies. This includes disabled nodes — reviewers open those too.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No personal or tenant identifiers&lt;/strong&gt;: real spreadsheet IDs, channel IDs, chat IDs, calendar IDs, e-mail addresses, internal hostnames. A real spreadsheet ID in a template is both a privacy problem and a broken template for everyone else.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Webhook paths should no longer be live.&lt;/strong&gt; Rename them if the original path was reachable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Credentials should not be embedded.&lt;/strong&gt; Our exported workflows deliberately contain no credential blocks at all — the import dialog then asks the user to attach their own credential to a named node, which is the behaviour you want.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Replace placeholders with obvious placeholders.&lt;/strong&gt; We use &lt;code&gt;REPLACE_WITH_YOUR_CHAT_ID&lt;/code&gt;, &lt;code&gt;REPLACE_WITH_YOUR_TOKEN&lt;/code&gt;, &lt;code&gt;YOUR-N8N-HOST&lt;/code&gt; and &lt;code&gt;example.com&lt;/code&gt; — strings that are unmistakably not real values, and greppable in the documentation.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A template whose HTTP node contains only a token placeholder and a documented "attach the credential here" step is also one your own account survives: the JSON in a public repo should not be the JSON that runs your production.&lt;/p&gt;

&lt;h2&gt;
  
  
  Preparing the JSON
&lt;/h2&gt;

&lt;p&gt;The template artefact is a single exported workflow file, and preparing it is a mechanical pass:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Build the workflow in a clean project.&lt;/strong&gt; Not in the instance where it already runs with live credentials.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Rename every node, describe every stage, add the sticky notes.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Strip credentials&lt;/strong&gt; and replace every environment-specific value with a placeholder.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check that it imports into a fresh instance&lt;/strong&gt;: Workflows → ⋯ → Import from File. If a node shows a red error badge before you have attached anything, the template will be sent back.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test the happy path with a real credential&lt;/strong&gt;, and keep the &lt;code&gt;curl&lt;/code&gt; invocation that exercises it in the template README.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Export, then re-import the exported file&lt;/strong&gt; and read it once more. This catches the placeholder you replaced in the wrong branch.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you used a community node&lt;/strong&gt;, the submission needs a top image — a screenshot of the workflow canvas with the sticky-note sections visible works best. Workflows that use only built-in nodes (ours do) do not need one.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Publishing the same JSON in a public repo alongside the submission gives reviewers and users the same source of truth, and gives you a licence statement to point at — ours is MIT.&lt;/p&gt;

&lt;h2&gt;
  
  
  What reviewers are actually filtering for
&lt;/h2&gt;

&lt;p&gt;Two prohibitions are worth stating explicitly because they are instant rejections rather than fixable notes: &lt;strong&gt;plagiarism&lt;/strong&gt; (a renamed copy of an existing library template) and &lt;strong&gt;low-effort templates&lt;/strong&gt; (one node wrapping one API call, or a workflow that only works after the user rewrites half of it).&lt;/p&gt;

&lt;p&gt;The practical test we apply before submitting: would this still be useful to someone who does not use our service, and does every node in it earn its place? Our own submissions are business reporting and lead-intake workflows built on that standard, and the ones that touch our &lt;a href="https://bowhard.ru/api" rel="noopener noreferrer"&gt;transcription HTTP API&lt;/a&gt; are documented so that the API call is one clearly-labelled step rather than the whole template.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The template collection, JSON and per-template READMEs: &lt;a href="https://github.com/VavilkinAlex/n8n-bowhard-templates" rel="noopener noreferrer"&gt;https://github.com/VavilkinAlex/n8n-bowhard-templates&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Official submission guidelines: &lt;a href="https://n8n.notion.site/p/Template-submission-guidelines-9959894476734da3b402c90b124b1f77" rel="noopener noreferrer"&gt;https://n8n.notion.site/p/Template-submission-guidelines-9959894476734da3b402c90b124b1f77&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Sticky-note guidelines: &lt;a href="https://n8n.notion.site/Sticky-note-guidelines-for-templates-2aa5b6e0c94f8058b0aefddd02655887" rel="noopener noreferrer"&gt;https://n8n.notion.site/Sticky-note-guidelines-for-templates-2aa5b6e0c94f8058b0aefddd02655887&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>automation</category>
      <category>opensource</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Building a zero-dependency Python and Node client for a speech-to-text API</title>
      <dc:creator>Bowhard</dc:creator>
      <pubDate>Sat, 10 Oct 2026 22:46:00 +0000</pubDate>
      <link>https://dev.to/bowhard/building-a-zero-dependency-python-and-node-client-for-a-speech-to-text-api-2ggk</link>
      <guid>https://dev.to/bowhard/building-a-zero-dependency-python-and-node-client-for-a-speech-to-text-api-2ggk</guid>
      <description>&lt;p&gt;Most speech-to-text and text-to-speech APIs ship a fat SDK: a package manager, a dependency tree, a version matrix. We went the other way. Our transcription and synthesis service speaks plain HTTP with JSON payloads and exactly one multipart upload, and we wrote the reference clients as a proof: they talk to the API with nothing but the standard library. No &lt;code&gt;requests&lt;/code&gt;, no &lt;code&gt;axios&lt;/code&gt;, no &lt;code&gt;node-fetch&lt;/code&gt; — Python's &lt;code&gt;urllib&lt;/code&gt; and Node's built-in &lt;code&gt;fetch&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;This is a walkthrough of how those clients work and which parts of the API forced the design. The code is on GitHub at &lt;a href="https://github.com/VavilkinAlex/bowhard-speech" rel="noopener noreferrer"&gt;VavilkinAlex/bowhard-speech&lt;/a&gt;; the endpoint reference is at &lt;a href="https://bowhard.ru/api" rel="noopener noreferrer"&gt;bowhard.ru/api&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  No API key means the cookie jar is the client
&lt;/h2&gt;

&lt;p&gt;The service does not issue API keys. It sets a signed cookie (&lt;code&gt;tools_id&lt;/code&gt;) on your first request and counts your daily free volume — 15 minutes of transcription and 5,000 characters of synthesis — by that cookie. Your jobs are bound to the same identity.&lt;/p&gt;

&lt;p&gt;That single decision shapes both clients: &lt;strong&gt;the client is a cookie jar plus a request wrapper.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In Python that is one line of standard library:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;urllib.request&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;http.cookiejar&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;CookieJar&lt;/span&gt;

&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_opener&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;urllib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;build_opener&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;urllib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;HTTPCookieProcessor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;CookieJar&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every call goes through &lt;code&gt;self._opener&lt;/code&gt;, so &lt;code&gt;Set-Cookie&lt;/code&gt; on the first response is replayed on all the following ones without any code from us.&lt;/p&gt;

&lt;p&gt;Node is where it gets interesting. &lt;code&gt;fetch&lt;/code&gt; is built in and good enough, but it is deliberately stateless about cookies — there is no jar. So the Node client carries its own, a plain &lt;code&gt;Map&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;cookies&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="c1"&gt;// when sending&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;cookies&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;size&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;cookie&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[...&lt;/span&gt;&lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;cookies&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(([&lt;/span&gt;&lt;span class="nx"&gt;k&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;v&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;k&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;=&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;v&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;; &lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// when receiving&lt;/span&gt;
&lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;raw&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;getSetCookie&lt;/span&gt;&lt;span class="p"&gt;?.()&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="p"&gt;[])&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;pair&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;;&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt;&lt;span class="nx"&gt;rest&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;pair&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;=&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;cookies&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="nx"&gt;rest&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;=&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two details are easy to get wrong. First, &lt;code&gt;getSetCookie()&lt;/code&gt; — the plural accessor — is what you need; the single &lt;code&gt;headers.get('set-cookie')&lt;/code&gt; flattens multiple cookies into one string and breaks parsing. Second, you must keep only the &lt;code&gt;name=value&lt;/code&gt; part: the attributes after the first semicolon (&lt;code&gt;Path&lt;/code&gt;, &lt;code&gt;HttpOnly&lt;/code&gt;, &lt;code&gt;Max-Age&lt;/code&gt;) are server-side instructions, not something you echo back.&lt;/p&gt;

&lt;p&gt;One consequence worth stating plainly: &lt;strong&gt;one client instance is one job owner.&lt;/strong&gt; Creating a new &lt;code&gt;Client&lt;/code&gt; discards the cookie, and with it the ability to read the jobs it just started. If you build a worker that fans out work, keep one instance alive.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building multipart by hand in Python
&lt;/h2&gt;

&lt;p&gt;There is exactly one place where a JSON library is not enough: &lt;code&gt;POST /api/tools/stt&lt;/code&gt; wants &lt;code&gt;multipart/form-data&lt;/code&gt; with a &lt;code&gt;file&lt;/code&gt; field. &lt;code&gt;urllib&lt;/code&gt; has no multipart encoder, so the Python client assembles the body itself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;boundary&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;uuid&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;uuid4&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nb"&gt;hex&lt;/span&gt;
&lt;span class="n"&gt;parts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;extra&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;{}).&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;parts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;--&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;boundary&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\r\n&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;
        &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Content-Disposition: form-data; name=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\r\n\r\n&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;
        &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\r\n&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;mime&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;mimetypes&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;guess_type&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;filename&lt;/span&gt;&lt;span class="p"&gt;)[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;application/octet-stream&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;
&lt;span class="n"&gt;parts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;--&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;boundary&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\r\n&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;
    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Content-Disposition: form-data; name=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;field_name&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;; &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;
    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;filename=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;basename&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;filename&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\r\n&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;
    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Content-Type: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;mime&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\r\n\r\n&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;parts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;parts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="se"&gt;\r\n&lt;/span&gt;&lt;span class="s"&gt;--&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;boundary&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;--&lt;/span&gt;&lt;span class="se"&gt;\r\n&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

&lt;span class="n"&gt;body&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;b&lt;/span&gt;&lt;span class="sh"&gt;''&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;parts&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;content_type&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;multipart/form-data; boundary=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;boundary&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The rules that actually matter:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The boundary must not appear in the payload.&lt;/strong&gt; A random 32-character hex string from &lt;code&gt;uuid4()&lt;/code&gt; makes that a non-issue.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CRLF, not LF.&lt;/strong&gt; The spec says &lt;code&gt;\r\n&lt;/code&gt; and some parsers check. Python's &lt;code&gt;email&lt;/code&gt; module would handle this, but then you are back to assembling anyway.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Extra fields come before the file part.&lt;/strong&gt; Cheap to respect, annoying to debug when a proxy-side parser assumes order.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Send the basename, not the path.&lt;/strong&gt; &lt;code&gt;filename="/home/you/secret/meeting.mp3"&lt;/code&gt; leaks your directory layout to the server and to any log in between.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Always send a &lt;code&gt;Content-Type&lt;/code&gt; per part.&lt;/strong&gt; &lt;code&gt;mimetypes.guess_type()&lt;/code&gt; covers the formats the service accepts (mp3, wav, ogg/opus, m4a, mp4, mov, mkv — anything ffmpeg reads); &lt;code&gt;application/octet-stream&lt;/code&gt; is the fallback.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The same request in Node is four lines, because Node 18+ ships &lt;code&gt;FormData&lt;/code&gt;, &lt;code&gt;Blob&lt;/code&gt; and &lt;code&gt;fetch&lt;/code&gt; together:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;form&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;FormData&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="nx"&gt;form&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;file&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Blob&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;readFile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;filePath&lt;/span&gt;&lt;span class="p"&gt;)]),&lt;/span&gt; &lt;span class="nf"&gt;basename&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;filePath&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="err"&gt;#&lt;/span&gt;&lt;span class="nf"&gt;request&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/stt&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;method&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;POST&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;form&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;fetch&lt;/code&gt; writes the boundary and the part headers itself. The interesting part is that the &lt;em&gt;server&lt;/em&gt; cannot tell the two apart — a hand-rolled body and a &lt;code&gt;FormData&lt;/code&gt; body are the same bytes on the wire.&lt;/p&gt;

&lt;h2&gt;
  
  
  Timeouts and retries
&lt;/h2&gt;

&lt;p&gt;The two clients make opposite choices here, on purpose.&lt;/p&gt;

&lt;p&gt;Python uses a single attempt with a 120-second timeout: &lt;code&gt;urllib&lt;/code&gt; raises, the client translates the exception into its own error type, and the caller decides. Retrying a request that uploads a 200 MB recording automatically is not obviously a kindness.&lt;/p&gt;

&lt;p&gt;Node retries — up to three attempts, with a linear backoff of &lt;code&gt;500 * (attempt + 1)&lt;/code&gt; milliseconds — because the failure it most often hits is a dropped connection on a long upload, and &lt;code&gt;fetch&lt;/code&gt; gives it a clean &lt;code&gt;AbortController&lt;/code&gt; to bound each attempt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;controller&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;AbortController&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;timer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;setTimeout&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;controller&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;abort&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;timeout&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The timeout is per attempt, not per operation, which is the behaviour you want: three tries of 120 seconds each, not one budget of 120 seconds split three ways.&lt;/p&gt;

&lt;h2&gt;
  
  
  Polling, because transcription is not instant
&lt;/h2&gt;

&lt;p&gt;Uploading a file does not return text. It returns a job:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"abc123"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"kind"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"stt"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"queued"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Transcription happens in a queue — the service processes two files at a time and the rest wait — so the time to completion depends on what else is running and on the length of your recording. There is no honest way to predict it, which is why the API is poll-based rather than webhook-based: a webhook would need a publicly reachable URL from you, and most client code does not have one.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;GET /api/tools/job/{id}&lt;/code&gt; returns everything you need to decide what to do next:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;field&lt;/th&gt;
&lt;th&gt;meaning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;status&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;queued&lt;/code&gt;, &lt;code&gt;processing&lt;/code&gt;, &lt;code&gt;done&lt;/code&gt;, &lt;code&gt;error&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;stage&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;what the worker is doing right now (used for progress output)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;error&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;human-readable failure reason when &lt;code&gt;status&lt;/code&gt; is &lt;code&gt;error&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;durationSec&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;length of the audio&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;chars&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;number of characters for synthesis jobs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;costRub&lt;/code&gt;, &lt;code&gt;free&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;what this job cost and whether it came out of the free daily allowance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;files&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;the result files that are ready to download&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Both clients wrap that in a &lt;code&gt;wait()&lt;/code&gt; helper that polls every 3 seconds, prints progress through a callback, and gives up after an hour by default:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;job&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;transcribe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;meeting.mp3&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;job&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;wait&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;on_progress&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;lambda&lt;/span&gt; &lt;span class="n"&gt;j&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;j&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;j&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stage&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;job&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;transcribe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;meeting.mp3&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;wait&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;onProgress&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;j&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;j&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;j&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;stage&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three seconds is a deliberate compromise. A tighter loop gains nothing — the queue position changes on the scale of seconds — and every poll is a request against the same cookie-counted account.&lt;/p&gt;

&lt;h2&gt;
  
  
  Downloading the results
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;GET /api/tools/job/{id}/file/{kind}&lt;/code&gt; serves the artefacts, where &lt;code&gt;kind&lt;/code&gt; is one of four:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;txt&lt;/code&gt; — the plain transcript, one paragraph flow, no timecodes;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;srt&lt;/code&gt; — subtitles, already wrapped to a maximum of 42 characters per line, which is the usual readability norm and the same limit our &lt;a href="https://bowhard.ru/subtitry" rel="noopener noreferrer"&gt;subtitle converter&lt;/a&gt; uses;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;json&lt;/code&gt; — the same transcript with a timecode for every single word, which is what you need if you are cutting video against the speech rather than just captioning it;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;mp3&lt;/code&gt; — the synthesised audio, for text-to-speech jobs.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In Python the helpers return strings or bytes (&lt;code&gt;client.text(job.id)&lt;/code&gt;, &lt;code&gt;client.srt(job.id)&lt;/code&gt;, &lt;code&gt;client.words(job.id)&lt;/code&gt;, &lt;code&gt;client.audio(job.id)&lt;/code&gt;); in Node they return strings or a &lt;code&gt;Buffer&lt;/code&gt;. Both have a &lt;code&gt;save()&lt;/code&gt; that writes straight to a path. One operational note: &lt;strong&gt;results live for 24 hours&lt;/strong&gt;, and the job statistics are kept for 90 days. If a transcript matters, download it in the same run that produced it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Errors that carry the server's own message
&lt;/h2&gt;

&lt;p&gt;An HTTP client that hides the response body costs you an hour of guessing. Both clients parse the error payload before throwing, and fall back to a truncated raw body when the response is not JSON:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="bp"&gt;...&lt;/span&gt;
&lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="n"&gt;urllib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;error&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;HTTPError&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;body&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;error&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;payload&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;decode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;payload&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;error&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;decode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;replace&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)[:&lt;/span&gt;&lt;span class="mi"&gt;300&lt;/span&gt;&lt;span class="p"&gt;]}&lt;/span&gt;
    &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;BowhardError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;error&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;HTTP &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;error&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;code&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                       &lt;span class="n"&gt;error&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;code&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The resulting &lt;code&gt;BowhardError&lt;/code&gt; carries &lt;code&gt;status&lt;/code&gt; and &lt;code&gt;payload&lt;/code&gt;, so &lt;code&gt;except BowhardError as e: print(e.status, e.payload)&lt;/code&gt; tells you whether you hit a rate limit (more than ten link imports per hour), a bad file type, or an expired job — without parsing strings.&lt;/p&gt;

&lt;h2&gt;
  
  
  The whole thing
&lt;/h2&gt;

&lt;p&gt;Both clients are a couple of hundred lines each, under MIT, with no dependencies to install and no lockfile to audit. Installation for the Python one is a single wheel from the repository's releases; the Node one installs from GitHub Packages. If you would rather not write HTTP by hand at all, there are CLI entry points in both:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;bowhard-speech transcribe soveschanie.mp3 &lt;span class="nt"&gt;-o&lt;/span&gt; protokol.txt &lt;span class="nt"&gt;--srt&lt;/span&gt; protokol.srt
bowhard-speech say &lt;span class="s2"&gt;"…"&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; privet.mp3 &lt;span class="nt"&gt;--voice&lt;/span&gt; alena &lt;span class="nt"&gt;--speed&lt;/span&gt; 1.1
npx bowhard-speech transcribe meeting.mp3 &lt;span class="nt"&gt;--srt&lt;/span&gt; meeting.srt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The interesting lesson is not "write your own HTTP client". It is that the API was small enough to make that a reasonable afternoon's work: one cookie, one multipart endpoint, one job resource, four file kinds. If your client library needs a dependency tree, it is usually because the API asked for one.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Client libraries (Python + Node, MIT): &lt;a href="https://github.com/VavilkinAlex/bowhard-speech" rel="noopener noreferrer"&gt;https://github.com/VavilkinAlex/bowhard-speech&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;HTTP API reference: &lt;a href="https://bowhard.ru/api" rel="noopener noreferrer"&gt;https://bowhard.ru/api&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;The transcription service itself: &lt;a href="https://bowhard.ru/rasshifrovka" rel="noopener noreferrer"&gt;https://bowhard.ru/rasshifrovka&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>python</category>
      <category>node</category>
      <category>api</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Speech-to-text and text-to-speech on Yandex SpeechKit from one Python stdlib file</title>
      <dc:creator>Bowhard</dc:creator>
      <pubDate>Sat, 10 Oct 2026 22:06:28 +0000</pubDate>
      <link>https://dev.to/bowhard/speech-to-text-and-text-to-speech-on-yandex-speechkit-from-one-python-stdlib-file-24b9</link>
      <guid>https://dev.to/bowhard/speech-to-text-and-text-to-speech-on-yandex-speechkit-from-one-python-stdlib-file-24b9</guid>
      <description>&lt;p&gt;The problem was mundane. On one side, meeting recordings: an mp3 of an hour or an hour and a half, out of which I needed text and subtitles. On the other, the reverse operation: I have text, I need a voice, because some clients listen rather than read. I did not want to buy a separate SaaS for this — the recordings are internal, and paying a subscription for them costs more than paying for the actual minutes.&lt;/p&gt;

&lt;p&gt;The constraints I set for myself up front:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;one VPS, no external queues or brokers;&lt;/li&gt;
&lt;li&gt;no dependencies beyond the system &lt;code&gt;ffmpeg&lt;/code&gt;/&lt;code&gt;ffprobe&lt;/code&gt; — so that an update means replacing one file and restarting a systemd unit;&lt;/li&gt;
&lt;li&gt;recognition and synthesis go straight to Yandex SpeechKit v3, with no middleware;&lt;/li&gt;
&lt;li&gt;user files are not kept longer than a day;&lt;/li&gt;
&lt;li&gt;cost has to be predictable: the service must not run at a loss because someone else's script uploads hour-long recordings all night.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What came out is a single Python 3.12 file of roughly 2,200 lines: &lt;code&gt;http.server&lt;/code&gt; with threads, &lt;code&gt;sqlite3&lt;/code&gt; as the job queue, &lt;code&gt;urllib&lt;/code&gt; for everything external, &lt;code&gt;ffmpeg&lt;/code&gt; for conversion. Below is what turned out to be non-obvious in this construction. Not a tutorial — a field notes collection of the places it bites.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the standard library was enough
&lt;/h2&gt;

&lt;p&gt;The stock &lt;code&gt;ThreadingHTTPServer&lt;/code&gt; cannot cap the number of concurrent connections: on a slow client it happily spawns thread after thread until file descriptors run out. I needed my own subclass — with a pool of 32 connection slots and a 120-second request read timeout. Everything else is ordinary routing: file upload via &lt;code&gt;multipart&lt;/code&gt;, a JSON API for job status, result delivery, and the payment provider's webhook.&lt;/p&gt;

&lt;p&gt;The queue is hand-rolled too, and it is surprisingly boring: a &lt;code&gt;jobs&lt;/code&gt; table in SQLite, two worker threads, the statuses &lt;code&gt;queued → processing → done|error&lt;/code&gt;, and a &lt;code&gt;stage&lt;/code&gt; field holding a human-readable string like "converting" or "recognizing". Here &lt;code&gt;sqlite3&lt;/code&gt; works as an ordinary job log rather than a high-load database: one row written per job, one row read per status poll.&lt;/p&gt;

&lt;p&gt;What that bought in practice: deploying the service is &lt;code&gt;scp&lt;/code&gt; of one file plus &lt;code&gt;systemctl restart&lt;/code&gt;. No &lt;code&gt;pip install&lt;/code&gt;, no migrations, no "your library version is wrong".&lt;/p&gt;

&lt;p&gt;The public REST API is documented at &lt;a href="https://bowhard.ru/api" rel="noopener noreferrer"&gt;bowhard.ru/api&lt;/a&gt;, and since I am not the only one who wanted to drive this from a script, there are small clients for &lt;a href="https://github.com/VavilkinAlex/bowhard-speech" rel="noopener noreferrer"&gt;Python and Node&lt;/a&gt; — that is how the transcription and text-to-speech services at &lt;a href="https://bowhard.ru/rasshifrovka" rel="noopener noreferrer"&gt;bowhard.ru/rasshifrovka&lt;/a&gt; and &lt;a href="https://bowhard.ru/ozvuchka" rel="noopener noreferrer"&gt;bowhard.ru/ozvuchka&lt;/a&gt; are consumed from automation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recognition: asynchronous, because the file is long
&lt;/h2&gt;

&lt;p&gt;SpeechKit can do synchronous recognition, but that is for short utterances. An hour-long recording means &lt;code&gt;recognizeFileAsync&lt;/code&gt;: send a job, get an &lt;code&gt;operationId&lt;/code&gt;, poll the status until it is &lt;code&gt;done&lt;/code&gt;. Poll every two seconds, with a three-hour deadline. Going below two seconds is pointless — the operation cannot have finished anyway, and you are just hammering the API for nothing.&lt;/p&gt;

&lt;p&gt;The input to SpeechKit comes in two flavours, and the choice between them is not cosmetic:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;content&lt;/code&gt; — audio as base64 inside the request body. Fine for small files, but base64 inflates the size by about a third, and holding an encoded hour-long recording in memory is not a great idea;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;uri&lt;/code&gt; — a link to an object in S3-compatible storage. We adopted a rule: files from 8 MB up go to Object Storage, anything smaller goes as base64. The threshold was picked from process memory: 8 MB in base64 is roughly an 11 MB string, which is perfectly safe for two workers.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The result of &lt;code&gt;getRecognition&lt;/code&gt; is not JSON but a stream of JSON lines, one object per line — that is the first non-obvious thing. It contains both the plain recognized words and a &lt;code&gt;finalRefinement.normalizedText.alternatives&lt;/code&gt; block. The words are needed for timings, but they carry no punctuation and no capital letters. The finished text with punctuation lives in the &lt;code&gt;text&lt;/code&gt; field of that same &lt;code&gt;normalizedText&lt;/code&gt; block. So we take the text from there, the timings from the word array, and splice the two into a single result.&lt;/p&gt;

&lt;p&gt;Subtitles are assembled from that same result. SRT is a format that forgives a lot except two things: over-long lines and overlapping timecodes. We break the line by words, never letting it exceed 42 characters, and wrap only at a word boundary. A subtitle block looks like this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the block number line — &lt;code&gt;12&lt;/code&gt;;&lt;/li&gt;
&lt;li&gt;the time line — &lt;code&gt;00:01:04,320 --&amp;gt; 00:01:07,880&lt;/code&gt;, with milliseconds separated by a comma, as the format requires;&lt;/li&gt;
&lt;li&gt;one or two lines of text, split by words and no longer than 42 characters each — &lt;code&gt;and we decided not to touch that part&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The timecodes come straight from the words, so the subtitles end up denser than when you slice by sentences, and they do not drift on fast speech. The user gets three files: the text, the &lt;code&gt;.srt&lt;/code&gt;, and a &lt;code&gt;json&lt;/code&gt; with the words and their timings — the last one is for anyone who wants word highlighting in a player.&lt;/p&gt;

&lt;h2&gt;
  
  
  Text-to-speech: 240 characters that nearly broke everything
&lt;/h2&gt;

&lt;p&gt;Synthesis in v3 is called through the &lt;code&gt;utteranceSynthesis&lt;/code&gt; method, and it has a limit on text length that the documentation says approximately nothing about. It emerged empirically: 250 characters pass, 500 come back as a refusal reading "Too long text". And the refusal does not always arrive straight away: sometimes it looks like a dropped connection, and had I not been logging the response body, I would have hunted for the cause for a long time.&lt;/p&gt;

&lt;p&gt;The fix is to cut the text into chunks of 240 characters, and to do it on sentence boundaries rather than by a character counter. It is inaudible, and the price is computed from the source character count, so the splitting does not affect the cost at all.&lt;/p&gt;

&lt;p&gt;Then comes the pleasant part: MP3 is a sequence of independent frames, so the synthesized chunks are joined by plain byte concatenation, with no re-encoding and no &lt;code&gt;ffmpeg&lt;/code&gt;. But if you cut the text crudely, mid-sentence, you can hear an unnatural pause at the seam — which is why the splitting happens at periods, question marks and exclamation marks.&lt;/p&gt;

&lt;p&gt;Two more things about synthesis. First: the v1 API costs roughly twice as much as v3 for the same result, so you should go straight to v3. Second: the speed parameter is accepted in the range 0.5–2.0, and anything outside it has to be clamped on your side, otherwise you get an error in the middle of a long text. SpeechKit has fifteen voices — from the neutral "Alena" to the conversational "Zakhar" — and for internal recordings the difference between them is far more noticeable than for advertising ones.&lt;/p&gt;

&lt;h2&gt;
  
  
  The queue, the ceilings, and the money fuse
&lt;/h2&gt;

&lt;p&gt;The nastiest class of bug in a service like this is not a crash but quietly burning money. So the ceilings are set at every level:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;no more than three jobs in flight per visitor;&lt;/li&gt;
&lt;li&gt;no more than 600 MB of uploads per day from one address;&lt;/li&gt;
&lt;li&gt;no more than 10 link imports per hour;&lt;/li&gt;
&lt;li&gt;the job directory cannot exceed 6 GB — past that, new jobs are refused until cleanup runs;&lt;/li&gt;
&lt;li&gt;and the main fuse: a daily cost ceiling. If the total recognition cost for the day exceeds a configured threshold, the service stops accepting new jobs and says so plainly, instead of running at a loss until morning.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A separate trap in the same area is balance accounting. Transcription is measured in minutes of audio, text-to-speech in characters. It is natural to keep both in one credits table, but if you compute usage as one combined sum you get nonsense: minutes get eaten by synthesis, characters by transcription, and the balance goes negative out of nowhere. Usage has to be counted separately per kind of work, and the interface has to show two independent balances. In the code this story survives as a comment — as a reminder.&lt;/p&gt;

&lt;h2&gt;
  
  
  Link import: SSRF is not only about the URL
&lt;/h2&gt;

&lt;p&gt;A "paste a link to a file" feature saves the user time and creates a whole class of risk. Ours is narrowed to two providers: public Yandex Disk files (via &lt;code&gt;cloud-api.yandex.net&lt;/code&gt;) and Dropbox. The list of trusted hosts is an allowlist, not a denylist, and the check is repeated after redirects: a provider may send you to another domain, and that is fine, but only if that domain is on the list too.&lt;/p&gt;

&lt;p&gt;Timeouts are split as well: 30 seconds for the provider's response and 600 seconds for the download itself, otherwise the service turns into a free relay for other people's slow files. Plus a per-address attempt counter. A dedicated self-test command exercises these scenarios: &lt;code&gt;selftest_links&lt;/code&gt;, &lt;code&gt;selftest_sign&lt;/code&gt;, &lt;code&gt;selftest_net&lt;/code&gt; — three buttons that ask "check that the outside world is still what we think it is".&lt;/p&gt;

&lt;h2&gt;
  
  
  Payments: we do not trust the webhook
&lt;/h2&gt;

&lt;p&gt;Payment goes through YooKassa, and there is exactly one rule here that must not be broken: a webhook is not confirmation of a payment. It is only a reason to re-check. The scheme is this: the payment is created in the database up front with a "created" status; on notification the service makes a separate &lt;code&gt;GET /payments/{id}&lt;/code&gt; request to the payment provider's API, looks at the real status, and only then credits the balance.&lt;/p&gt;

&lt;p&gt;The crediting itself must happen exactly once, because webhooks repeat. Instead of "read the flag, write the credit", a conditional update is used: &lt;code&gt;UPDATE payments SET paid_at = ? WHERE id = ? AND paid_at IS NULL&lt;/code&gt;, and the credit is granted only if that statement actually changed a row. This is cheaper and more reliable than any lock in the application: the race between two simultaneous webhooks is settled by the database.&lt;/p&gt;

&lt;h2&gt;
  
  
  Storage: signing SigV4 by hand
&lt;/h2&gt;

&lt;p&gt;Files from 8 MB up go to S3-compatible Object Storage. I did not want to pull in an SDK for four operations (&lt;code&gt;PUT&lt;/code&gt;, &lt;code&gt;GET&lt;/code&gt;, &lt;code&gt;DELETE&lt;/code&gt;, &lt;code&gt;HEAD&lt;/code&gt;), so the AWS SigV4 signature is assembled by hand on &lt;code&gt;urllib&lt;/code&gt; — and that turned out to be the most capricious part of the whole project.&lt;/p&gt;

&lt;p&gt;Two traps, both diagnosed equally unpleasantly — just a &lt;code&gt;403&lt;/code&gt; with no explanation:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;x-amz-content-sha256&lt;/code&gt; has to appear both in the canonical headers and in the &lt;code&gt;SignedHeaders&lt;/code&gt; list. The header is sent, the hash is computed, and the signature still does not match, because it is missing from the list being signed.&lt;/li&gt;
&lt;li&gt;Sub-resources such as &lt;code&gt;?lifecycle&lt;/code&gt; are written in the canonical request as &lt;code&gt;lifecycle=&lt;/code&gt; — with an empty value. Leave them as they are and you get &lt;code&gt;SignatureDoesNotMatch&lt;/code&gt; on a request that looks externally flawless.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Once the signature is assembled, though, everything else is trivial: put the object — hand SpeechKit the link — delete the object after cleanup. Job files live for 24 hours: that promise is written on the service page, which means the background cleanup has to honour it literally rather than "eventually". Statistics records are kept for 90 days — enough to understand load and compute cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would do differently
&lt;/h2&gt;

&lt;p&gt;The SpeechKit client should have been extracted into a separate module from week one: right now the recognition, synthesis and queue logic all live in one file, and that is convenient exactly until the moment you want to run synthesis from a separate script. Second — progress for long operations: today it shows the stage and a speed estimate based on already-processed jobs, but an honest completion fraction for an hour-long recording would be more useful. Third — move the free-tier limits into config rather than code: those are the ones you want to change more often than anything else.&lt;/p&gt;

&lt;p&gt;Otherwise the construction "one file, stdlib, SQLite and system &lt;code&gt;ffmpeg&lt;/code&gt;" has held up better than I expected: the service handles recordings up to four hours, recognizing an hour-long file takes less time than the recording itself lasts, and updating it requires neither a build nor dependencies.&lt;/p&gt;

&lt;p&gt;If you want to drive this from an automation platform rather than from code, there are ready-made &lt;a href="https://github.com/VavilkinAlex/n8n-bowhard-templates" rel="noopener noreferrer"&gt;n8n workflow templates&lt;/a&gt; for both transcription and synthesis. And if you build something similar, the most expensive time will go not into SpeechKit but into the S3 signature and the balance accounting — start with those.&lt;/p&gt;

</description>
      <category>python</category>
      <category>api</category>
      <category>opensource</category>
      <category>programming</category>
    </item>
  </channel>
</rss>
