DEV Community

Cover image for I Plugged Ollama Into My iPhone Keyboard. Here's the Full Self-Hosted Stack.
Ondrej Machala
Ondrej Machala

Posted on Edited on

I Plugged Ollama Into My iPhone Keyboard. Here's the Full Self-Hosted Stack.

You've got Ollama running on your home server. Your iPhone keyboard is still phoning home.

That gap bothered me enough to close it. Diction is an iOS dictation keyboard that speaks the standard OpenAI transcription API, POST /v1/audio/transcriptions. The app itself is on the App Store, but the gateway it talks to is open-source Go, configured entirely with environment variables, and you can run the whole thing on hardware you own.

Two things changed recently that make this worth revisiting. Cleanup by your own LLM is now free when it runs on your own server, because you are supplying the model and the electricity. And pairing your phone to your gateway is now a QR code instead of a key you type with your thumbs.

Here's the full stack.


What Each Piece Does

The speech engine does the transcription. Whisper is the default answer and works everywhere. If you dictate in English or another European language, NVIDIA's Parakeet model via the dictionlabs/parakeet image is a good alternative: 25 languages, roughly 10x faster than Whisper on CPU, about 2GB RAM, and models baked into the image so there is no first-run download.

If you need Asian languages, Arabic, or anything outside those 25, use Whisper. Everything below works with either.

The gateway sits in front of the speech engine. It handles WebSocket streaming, so your phone streams audio while you are still talking and the transcript is mostly ready by the time you stop. It also handles authentication and the LLM cleanup step.

Ollama cleans up the transcript after transcription. The gateway calls it with a system prompt you write, and the cleaned text is what lands in the app. Your model, your prompt.


The Stack

Three containers. No GPU required, though the cleanup model is where you will feel its absence: transcription on CPU is fine, while a 9B model cleaning a paragraph on CPU is the part that takes seconds rather than milliseconds. Under 3GB RAM without Ollama, 8-10GB with a 9B model loaded.

Start with the CPU version, then add the GPU blocks below if you have a card.

services:
  parakeet:
    image: dictionlabs/parakeet:latest-int8

  gateway:
    image: dictionlabs/gateway:v13.0
    ports:
      - "8080:8080"
    depends_on:
      - parakeet
    volumes:
      - gateway-data:/data
    environment:
      DEFAULT_MODEL: parakeet-v3
      PUBLIC_URL: https://whisper.example.com
      LLM_BASE_URL: http://ollama:11434/v1
      LLM_MODEL: gemma2:9b

  ollama:
    image: ollama/ollama
    volumes:
      - ollama-data:/root/.ollama

volumes:
  gateway-data:
  ollama-data:
Enter fullscreen mode Exit fullscreen mode
docker compose up -d
docker compose exec ollama ollama pull gemma2:9b
Enter fullscreen mode Exit fullscreen mode

Giving it a GPU

If the box has an NVIDIA card, install the NVIDIA Container Toolkit on the host and add a reservation to whichever services you want accelerated. Both benefit, and the cleanup model benefits most:

services:
  parakeet:
    image: dictionlabs/parakeet:latest-int8
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]

  ollama:
    image: ollama/ollama
    volumes:
      - ollama-data:/root/.ollama
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]
Enter fullscreen mode Exit fullscreen mode

One card can serve both. Transcription is bursty and short, cleanup is bursty and short, and they rarely collide for a single user dictating into a keyboard. If you do share a card, keep an eye on VRAM rather than utilization: a speech model and a 9B language model resident at once is the thing that runs you out of memory, not contention for compute.

With the cleanup model on a GPU you can also lower DICTION_ENHANCE_TIMEOUT_MS from its CPU-sized default, which is covered further down.

Skip the Ollama block entirely if you do not want AI cleanup. The gateway checks for LLM_BASE_URL on startup, and without it transcriptions come back raw. If Ollama already runs on another machine, point LLM_BASE_URL at it and use any model you have pulled. LLM_API_KEY is only needed for endpoints that require a bearer token, so local Ollama does not need it.

For Whisper instead of Parakeet, take the whisper-small service from the compose file in the repo and set DEFAULT_MODEL: small. Swapping the image alone is not enough: the gateway resolves backends by Docker hostname, so the service name has to match the model you asked for. Name it parakeet while DEFAULT_MODEL says small and every request returns 502. The Whisper services also pull their weights through an init container, which is why lifting them from the repo compose beats hand-rolling them.

An unrecognised model name is not an error either. The gateway falls back to DEFAULT_MODEL, so a typo transcribes with the wrong model instead of telling you. The X-Diction-Route-Model response header shows which backend actually served a request.

Note the gateway-data volume. That one matters, and the next section explains why.


Pairing: Scan a QR Code Instead of Typing a Key

Here is what the gateway prints on startup now:

Pair the Diction app with this gateway. Scan in Diction > Self-Hosted > Scan to pair:

  █▀▀▀▀▀█ █▄▀▀▄█ ▄ ▀▄▄▄███▀▄▀▄█ █ █ ▀ █  █
  █ ███ █ ▄▄ █ ▄▀▀█ █▄██▀ ▀▄▄▀██▄▄▀█▀▀███▀█
  █ ▀▀▀ █ ▀   ▄▄▀▄█▀ ████ ▀ ▄ ▄ ▄█▄  █ ██▄▀
  ▀▀▀▀▀▀▀ ▀▀   ▀▀▀  ▀   ▀▀ ▀ ▀  ▀▀▀▀ ▀  ▀▀

Or open this link on your iPhone:
  diction://pair?key=dk_qX6KAJT5aSEvqEz...
Enter fullscreen mode Exit fullscreen mode

The gateway generates its own key on first start. Open the app, go to Self-Hosted, tap Scan to Pair, point the camera at your terminal. No key gets typed anywhere.

Set PUBLIC_URL to the address your phone will actually reach, as in the compose file above. It gets embedded in the QR alongside the key, so pairing fills in both the address and the credential. Without it the QR carries only the key and you still enter the URL by hand.

If the phone is not in the same room as the terminal, send yourself the diction://pair link and tap it on the device. Same result.

Persist the keyring. The keys live at /data/gateway-keys.json, inside the container. No volume means a new key on every docker compose up, and every paired device stops working. That is what gateway-data:/data is for. Use DICTION_KEY_PATH if you would rather put the file somewhere specific.

Three auth modes

DICTION_GATEWAY_AUTH takes three values:

Value Behavior
off No key, no QR, nothing changes
optional Key generated, QR printed, key accepted, keyless requests still pass
required Keyless requests rejected with 401 invalid_key

The default is optional, deliberately. Pulling a new image should not lock you out of your own server, so on upgrade the gateway advertises pairing while continuing to serve devices you already had.

Understand that optional is a migration state, not a security setting. Pair your devices, confirm dictation works, then commit:

environment:
  DICTION_GATEWAY_AUTH: required
Enter fullscreen mode Exit fullscreen mode

Rotation without a re-pairing party

Each device holds its own token minted from the gateway's current secret, rather than every device sharing one string. So rotating invalidates the secret, not any particular device.

The retired secret keeps verifying for DICTION_KEY_GRACE, which defaults to 720h, or 30 days. A phone that was switched off during a rotation comes back, presents its old token, is accepted inside the window, and quietly re-mints against the new secret. Nobody scans anything again.

To expire tokens on a schedule:

environment:
  DICTION_TOKEN_TTL: 2160h   # 90 days
Enter fullscreen mode Exit fullscreen mode

Unset means tokens never expire, which is the default. Two things to know before enabling it. Setting a TTL invalidates every token minted while it was unset, so every device re-pairs once. And TTL only means something alongside DICTION_GATEWAY_AUTH: required, because expiring a credential accomplishes nothing while keyless requests are still accepted. The gateway prints a warning at startup for that combination, since it looks secure and is not.

Be honest with yourself about what TTL buys. It bounds passive exposure: a QR in logs you shipped somewhere, a token in an old device backup, a photographed QR from months ago. It does not stop someone who already holds a valid token and refreshes it before expiry. Stateless bearer auth cannot tell that refresh chain from the real device.


Connecting the App

  1. Install Diction from the App Store
  2. In iPhone Settings: General → Keyboard → Keyboards → Add New Keyboard → Diction
  3. Open the Diction app and switch to Self-Hosted
  4. Tap Scan to Pair and point the camera at the QR in your terminal. If you prefer to enter things by hand, Enter Manually takes an address and an API key.
  5. The screen then shows plain-language status: whether the speech model is ready, and whether the LLM is ready

The keyboard is now using your server. Audio goes from your phone to your server, the speech engine transcribes it, Ollama cleans the result, text lands in whatever app you are typing in.

Away from home, Tailscale or a Cloudflare Tunnel connects your phone without opening router ports.


Cleanup Runs on Your Server, and It Is Free There

This used to be the awkward part of this guide. The cleanup toggle needed a subscription even though the work happened on your hardware, which I could never really defend.

That is fixed. Cleanup, grammar and punctuation, and voice editing where you select text and say what to change all run against your own model at no charge. I charge when I supply the inference. On your server, I am not supplying anything.

Check that the gateway sees your model:

curl -s http://localhost:8080/v1/models | jq .capabilities
Enter fullscreen mode Exit fullscreen mode
{
  "formatting": true,
  "key_rotation": true,
  "llm": true,
  "pairing": true,
  "text_process": true,
  "text_suggest": true,
  "text_summarize": true
}
Enter fullscreen mode Exit fullscreen mode

The app reads this. When llm is false, Writing Tools shows as unavailable with a reason instead of appearing configured and silently doing nothing, which is how the previous version failed and the thing I most wanted to stop.

Why pairing is what unlocks this safely

This connection surprised me while building it.

The text endpoints, /v1/text/process and /v1/text/suggest, are closed by default. Call one on an unauthenticated gateway and you get:

{"error":"text_routes_closed","hint":"Set TEXT_ROUTES_OPEN=true or AUTH_ENABLED=true to enable /v1/text/* routes. See AGENTS.md."}
Enter fullscreen mode Exit fullscreen mode

You can force them open with TEXT_ROUTES_OPEN=true. I would rather you did not, because that leaves an unauthenticated text-generation endpoint on your network for anything to use, burning your GPU on someone else's prompts.

A paired device is authenticated, so it bypasses that guard. Pairing is therefore not only what stops strangers sending audio to your server. It is also what lets you turn cleanup on without leaving an open LLM proxy behind. Pair your devices, set required, and leave TEXT_ROUTES_OPEN alone.

Two timeouts sized for hardware you own

A local model on CPU is slower than a hosted one, and the gateway bounds the cleanup pass itself rather than letting your phone's patience decide.

environment:
  DICTION_ENHANCE_TIMEOUT_MS: 20000      # cleanup while you wait
  DICTION_LIVE_ENHANCE_TIMEOUT_MS: 8000  # cleanup after raw text is delivered
Enter fullscreen mode Exit fullscreen mode

Defaults are 20 seconds and 8. The first is generous on purpose, because a 7B model on CPU genuinely takes that long and the alternative is failing a request that would have succeeded. On timeout you get the raw transcript, so a slow model costs you the cleanup, never your words.

Do not bother raising the second much past 9 seconds. The app stops waiting for the enhanced frame at 9 and keeps the raw text, so a bigger number changes nothing. Both accept 0 for no limit.


Writing a Prompt That Works

LLM_PROMPT is a single system prompt sent with every transcription request. The transcript is the user message. You control both. Leave it unset and the gateway uses its own cleanup prompt, which is a reasonable default.

A few starting points:

# General dictation
Remove filler words (um, uh, like, you know). Fix punctuation and grammar.
Preserve meaning and tone. Return only the cleaned result.
Enter fullscreen mode Exit fullscreen mode
# Technical / developer notes
Fix transcription errors. Preserve technical terms, command names, and file paths
exactly as spoken. Remove filler words. Return cleaned text only.
Enter fullscreen mode Exit fullscreen mode
# Medical or domain-specific
Fix transcription errors. Preserve all domain-specific terminology exactly as spoken.
Fix grammar and punctuation only. Return the corrected text.
Enter fullscreen mode Exit fullscreen mode

One practical note that costs people an evening: models under 7B parameters often answer the transcript instead of cleaning it. You dictate "quick note about tomorrow's meeting" and get back "Sure, what would you like the note to say?" Gemma2 9B is reliable here. Qwen2.5 7B is borderline. Anything 9B or above behaves predictably.

If you want separate prompts for the editing paths, LLM_PROMPT_EDIT and LLM_PROMPT_EDIT_SELECTED exist and fall back to sensible built-ins.


What You Get

Audio from your phone to your server. Your speech engine transcribes it. Your model cleans it. Text inserted, in any app. No third party in the path, no word limits, and nothing to pay for the AI because it is your hardware running it.

The gateway is open source, so you can read every line that runs on your network. The image is on Docker Hub as dictionlabs/gateway.

If you already run a homelab with Ollama, the marginal effort here is one compose file, one QR scan, and about ten minutes.

Top comments (0)