<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: whisperdirect</title>
    <description>The latest articles on DEV Community by whisperdirect (@whisperdirect).</description>
    <link>https://dev.to/whisperdirect</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4174116%2Fa9be0690-257b-4c04-b253-ef767fe92e92.jpeg</url>
      <title>DEV Community: whisperdirect</title>
      <link>https://dev.to/whisperdirect</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/whisperdirect"/>
    <language>en</language>
    <item>
      <title>1 Year of Building a "Bring Your Own API Key" Transcription App — and Why It's Now an Automation Hub</title>
      <dc:creator>whisperdirect</dc:creator>
      <pubDate>Fri, 09 Oct 2026 20:48:10 +0000</pubDate>
      <link>https://dev.to/whisperdirect/1-year-of-building-a-bring-your-own-api-key-transcription-app-and-why-its-now-an-automation-hub-57jk</link>
      <guid>https://dev.to/whisperdirect/1-year-of-building-a-bring-your-own-api-key-transcription-app-and-why-its-now-an-automation-hub-57jk</guid>
      <description>&lt;p&gt;Hi, I'm a solo iOS developer.&lt;/p&gt;

&lt;p&gt;About a year ago, when I first released my transcription app, I wrote this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"The moment you go subscription, you can't stop."&lt;br&gt;
"Even if only one person has ever paid, you have to keep the servers running — even at a loss."&lt;br&gt;
"A single traffic spike can degrade your service."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;While every AI app was racing toward monthly subscriptions, I realized that for a solo developer, running a backend long-term is just too much risk. Then it hit me:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What if I just build the "vessel" — and let the user bring their own OpenAI API key?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A year later, WhisperDirect has grown into something with transcription, summarization, meeting minutes, and external automation built in.&lt;/p&gt;

&lt;p&gt;This post is the story of that year. Why an "API-direct, fully one-time-purchase" transcription app ended up here.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. WhisperDirect is API-direct. WhisText is the lab.
&lt;/h2&gt;

&lt;p&gt;WhisperDirect started with a very simple idea:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The user's own OpenAI API key, used directly against Whisper API.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The app itself is a one-time purchase. No subscription.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;API usage is paid by the user directly to OpenAI.&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In other words, WhisperDirect was, from day one, &lt;strong&gt;a vessel for using OpenAI via your own API key&lt;/strong&gt;. No backend to maintain, no middleman markup. You pay OpenAI for what you use. That part of the design has never changed.&lt;/p&gt;

&lt;p&gt;On the other hand, I had another app: &lt;strong&gt;WhisText&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Think of it like Fedora vs. Red Hat in the Linux world.&lt;/p&gt;

&lt;p&gt;Cutting-edge tech goes into &lt;strong&gt;WhisText&lt;/strong&gt; first. Whatever survives real-world use gets ported into &lt;strong&gt;WhisperDirect&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Over the past year, that cycle has accelerated dramatically.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqgcm37clrcsyj2btscs0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqgcm37clrcsyj2btscs0.png" alt="WhisperDirect API settings" width="563" height="1146"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  2. WhisText moved from GPU → CPU → on-device
&lt;/h2&gt;

&lt;p&gt;WhisText's transcription originally ran on my own GPU server. The model was Whisper large-v3-turbo. My plan was: let people use it for free, watch the usage, then figure out pricing.&lt;/p&gt;

&lt;p&gt;Reality was less generous.&lt;/p&gt;

&lt;p&gt;Usage never justified keeping a GPU running 24/7, and the bills kept coming. So I switched to CPU and started testing every ASR engine I could find.&lt;/p&gt;

&lt;p&gt;My dev server's &lt;code&gt;asr.sherpa&lt;/code&gt; directory still holds the scars:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Faster-Whisper&lt;/li&gt;
&lt;li&gt;ReazonSpeech Zipformer&lt;/li&gt;
&lt;li&gt;SenseVoice&lt;/li&gt;
&lt;li&gt;Parakeet&lt;/li&gt;
&lt;li&gt;Nemotron&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I tried different language combinations, split audio into chunks for parallel processing, tuned for speed. A lot of unglamorous trial and error.&lt;/p&gt;

&lt;p&gt;Eventually, &lt;strong&gt;Apple Speech became the center of WhisText's final version.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Apple Speech used to be bad. iOS 26 changed that.
&lt;/h2&gt;

&lt;p&gt;I had actually tried Apple's native Apple Speech pretty early on.&lt;/p&gt;

&lt;p&gt;At the time, it was unusable. The biggest blocker was the &lt;strong&gt;1-minute limit&lt;/strong&gt;. For an app that records meetings and produces minutes, cutting off every 60 seconds is fatal. So I went back to tuning my own server.&lt;/p&gt;

&lt;p&gt;Then iOS 26 changed everything.&lt;/p&gt;

&lt;p&gt;Apple Speech's accuracy jumped, and the old limits and instability were gone. In my own benchmarks, it was &lt;strong&gt;matching or beating&lt;/strong&gt; engines I'd been running on my own server — Parakeet, ReazonSpeech, all of them.&lt;/p&gt;

&lt;p&gt;So WhisText's final version put Apple Speech at the core.&lt;/p&gt;

&lt;p&gt;No server round-trip means: &lt;strong&gt;the instant you speak into the mic, the text appears.&lt;/strong&gt; I added live preview as a bonus.&lt;/p&gt;

&lt;p&gt;That said, live preview is a bonus. If you want accurate meeting minutes, batch-processing the whole recording with Whisper API afterward still produces cleaner results. Live preview is more about the &lt;em&gt;feeling&lt;/em&gt; — "it's recording," "AI is working right now."&lt;/p&gt;

&lt;p&gt;But if Apple Speech on iOS 26 is this good, there must be more we can do on-device.&lt;/p&gt;

&lt;p&gt;So the modules that moved to on-device in WhisText got ported into WhisperDirect:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;On-device transcription via Apple Speech&lt;/li&gt;
&lt;li&gt;On-device speaker diarization via NVIDIA Nemotron 3&lt;/li&gt;
&lt;li&gt;OCR for text from images&lt;/li&gt;
&lt;li&gt;Waveform / simple display toggle&lt;/li&gt;
&lt;li&gt;Live preview (bonus)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;These became the foundation that works without an API key, for free.&lt;/strong&gt; The accuracy improvements in iOS 26's Apple Speech are what made this foundation actually usable.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkbw8t5geb9ntki2p2xzq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkbw8t5geb9ntki2p2xzq.png" alt="WhisperDirect transcription engine options — Whisper API or on-device Apple Speech" width="579" height="1151"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  4. The latest WhisperDirect is no longer just a transcription app
&lt;/h2&gt;

&lt;p&gt;The latest version — currently in App Store review — packs in a lot more.&lt;/p&gt;

&lt;p&gt;WhisperDirect is a voice transcription and summarization app built around Whisper API's accuracy. Alongside Whisper API (with your own key), it ships with on-device features: Apple Speech, speaker diarization, and OCR. For summarization and minutes, you can choose your LLM from OpenAI, Gemini, or any OpenAI-compatible endpoint.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The app is a one-time purchase. No subscriptions.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Pricing &amp;amp; experience
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;7-day free trial&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;On-device features (Apple Speech / diarization / OCR): no API key, free&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Whisper API and LLM calls: paid per-use to the provider&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The app itself never charges for API usage&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Cost reference: Whisper API is about &lt;strong&gt;$0.006/min&lt;/strong&gt;.&lt;br&gt;
That's roughly &lt;strong&gt;$0.36/hour&lt;/strong&gt; — about &lt;strong&gt;8.7 hours of transcription for around $3&lt;/strong&gt;.&lt;/p&gt;
&lt;h3&gt;
  
  
  Transcription engines
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Whisper API (OpenAI)&lt;/strong&gt;: your key, high accuracy&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Apple Speech (on-device)&lt;/strong&gt;: no API key, iOS 26+&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Switchable in Settings&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;
  
  
  Summarization / minutes models
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;OpenAI&lt;/strong&gt;: GPT-6 Luna / GPT-5 nano / GPT-5 mini / GPT-4.1 nano / GPT-4.1 mini, etc.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemini&lt;/strong&gt;: available with a Google AI Studio API key&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Custom&lt;/strong&gt;: any OpenAI-compatible API (Ollama / LM Studio / your own server)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;With a local LLM, you can bring the summarization / minutes cost down to zero.&lt;/strong&gt;&lt;br&gt;
For a few thousand characters, cloud APIs cost just a few cents. That kind of freedom is only possible for a solo developer.&lt;/p&gt;
&lt;h3&gt;
  
  
  Main features
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Record from mic → text, on the spot&lt;/li&gt;
&lt;li&gt;Live transcription display while recording (live preview / bonus)&lt;/li&gt;
&lt;li&gt;Import audio files&lt;/li&gt;
&lt;li&gt;Import video files (audio extract &amp;amp; compress)&lt;/li&gt;
&lt;li&gt;Auto-highlight synced to playback position&lt;/li&gt;
&lt;li&gt;Timeline insertion at configurable intervals&lt;/li&gt;
&lt;li&gt;Summaries / minutes (prompts editable)&lt;/li&gt;
&lt;li&gt;Image OCR&lt;/li&gt;
&lt;li&gt;Speaker diarization (NVIDIA Nemotron 3, on-device)&lt;/li&gt;
&lt;li&gt;Export: audio / text / summary / minutes / subtitles (VTT / SRT)&lt;/li&gt;
&lt;li&gt;Auto-post to Slack&lt;/li&gt;
&lt;li&gt;Webhook integration (Discord / Slack / n8n, etc.)&lt;/li&gt;
&lt;li&gt;Auto-backup to Google Drive&lt;/li&gt;
&lt;li&gt;Estimated cost display&lt;/li&gt;
&lt;li&gt;LLM model selection, timeline interval, prompt customization&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9df17d4brx7rajsvg56q.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9df17d4brx7rajsvg56q.png" alt="WhisperDirect detail view with transcription, speaker diarization, and playback controls" width="567" height="1142"&gt;&lt;/a&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  5. Webhook integration — send what you want, when you want
&lt;/h2&gt;

&lt;p&gt;I built this for myself, honestly.&lt;/p&gt;

&lt;p&gt;You can POST transcription / summary / minutes results to an external endpoint, each at the moment it completes. From there, &lt;strong&gt;n8n handles the rest.&lt;/strong&gt; With n8n, you can do whatever LLM processing you want, then route to Notion, Google Sheets, Slack, or anything else — just configure the URL. Header auth is supported too.&lt;/p&gt;

&lt;p&gt;The payload looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"transcript"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"recorded_at"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-10-09T10:00:00Z"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"sent_at"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-10-09T10:05:00Z"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"source"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"whisperdirect"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"version"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"1.0"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6xfy0bp159kdphxj48ji.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6xfy0bp159kdphxj48ji.png" alt="System integration settings and payload preview JSON" width="800" height="801"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  6. What only a solo developer could do
&lt;/h2&gt;

&lt;p&gt;For a company-run app, you need a monthly subscription just to cover server costs, salaries, and ad spend.&lt;/p&gt;

&lt;p&gt;But WhisperDirect has no backend to maintain.&lt;/p&gt;

&lt;p&gt;The app is a one-time purchase.&lt;/p&gt;

&lt;p&gt;When an API is used, the user pays the provider directly — at cost.&lt;/p&gt;

&lt;p&gt;I can't spend money on ads.&lt;br&gt;
So instead of competing on ad spend, I compete on cost-performance and freedom.&lt;/p&gt;

&lt;p&gt;What I can do as a solo developer, and what I can't.&lt;br&gt;
I've been carrying both, the whole way here.&lt;/p&gt;

&lt;p&gt;The latest version is currently in App Store review (live preview, etc.).&lt;br&gt;
There's a 7-day free trial.&lt;/p&gt;

&lt;p&gt;If this sounds interesting, search for WhisperDirect on the App Store.&lt;br&gt;
Once it's approved, give the new transcription experience a try.&lt;/p&gt;




&lt;h2&gt;
  
  
  7. How I actually use it
&lt;/h2&gt;

&lt;p&gt;My recommended setup:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Set Apple Speech as the default engine.&lt;/strong&gt; For everyday use, just leave it as-is.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;For quick notes&lt;/strong&gt;, use the on-device engine. It's free and instant.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;For meetings or anything that needs accuracy&lt;/strong&gt; — you don't have to decide upfront. Record with Apple Speech, then from the Detail view, tap the top-right menu and run &lt;strong&gt;"Redo Transcription (Whisper)."&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you want to save even more&lt;/strong&gt;: from the Detail view, re-transcribe only the segments that came out wrong (Whisper, segment-level). A few seconds of audio costs a fraction of a cent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;For LLMs, use whatever fits the task.&lt;/strong&gt; Gemini's free tier works well. If you run your own LLM on Oracle Cloud or similar, tunnel in via Tailscale and point the app there. I use Gemma 4 E2B for this — I don't need summarization or minutes to be fast.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Webhook (Slack, Discord — sends to yourself, synced across devices)&lt;/strong&gt; and &lt;strong&gt;auto-backup to Google Drive&lt;/strong&gt; are also available.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The point is: &lt;strong&gt;you don't have to choose between "cheap" and "accurate." You decide per recording, per segment.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;WhisperDirect: &lt;a href="https://apps.apple.com/app/id6748595475" rel="noopener noreferrer"&gt;https://apps.apple.com/app/id6748595475&lt;/a&gt;&lt;/p&gt;

</description>
      <category>whisper</category>
      <category>ios</category>
      <category>automation</category>
      <category>showdev</category>
    </item>
  </channel>
</rss>
