Moondream recently released Parakeet Redux, a speech-to-text model that fits in 178 MB. It's NVIDIA's Parakeet TDT 0.6B v3, one of the strongest open speech models, with its weights cut down to three values each. The original is 1,256 MB, and the model card says accuracy stays close to it.
A 0.6B-parameter model in roughly the space of a 61M one would change where good speech recognition can run: on an ordinary CPU, offline, inside an app you download in seconds. So I wanted to know what it actually gives up. I tested it against its own original, a standard 8-bit copy of that original, four other open models and three cloud APIs. First on my laptop, then again on a cloud machine with an NVIDIA Tesla T4, to see how it does on a GPU.
The short version: Redux is 7× smaller than its original and was faster on every device I tried: 15× real time on six laptop CPU cores and 85× on the T4, about 15% ahead of the original there. In exchange it made 0.6 to 0.8 more errors per 100 words, mostly on accented speech and dictation. And the small file doesn't mean a small footprint: it used about 1 GB of RAM on a CPU and 2 GB of GPU memory on the T4. English only, measured on 5 and 6 October 2026.
What Parakeet Redux is
Redux doesn't change the model's design. It keeps the architecture of Parakeet TDT 0.6B v3, and its files name that model as the "teacher" it was trained against. What changes is how the weights are stored.

Read from the model's own config.json, ternary.json and weights file.
- The input: 16 kHz audio turned into 128 mel bands, then shortened eight times by a subsampling stage, so the model sees one frame for every 80 ms of sound.
- The encoder does the listening: 24 Conformer layers, each with self-attention (8 heads), a convolution (kernel 9) and two feed-forward blocks, 1,024 wide. This is where almost all of the weights live.
- The decoder is a TDT (token-and-duration transducer): a small prediction network (a two-layer LSTM, 640 wide) and a joint network that emit a token together with how many frames to skip. Skipping ahead is a big part of why Parakeet models are fast.
- A voice activity head: a tiny 0.2M-parameter side head that marks speech and silence every 80 ms, so the model can tell where the talking is on its own.
How it got so small
Normally each weight is a 16-bit number. Redux turns the encoder's 604 million weights into ternary values: every weight is −1, 0 or +1, and each group of 128 weights shares one scale. Three possible values need about 1.58 bits, and Redux packs five weights into a single byte (35 = 243 combinations, which fits in 256), so each weight costs 1.6 bits instead of 16. In the released file, 44% of those weights are exactly zero.
That shrinks the encoder to 121 MB. The parts that are small but sensitive stay in 16-bit: the decoder, the joint network, the token embeddings, the subsampling layers and the per-group scales, about 55 MB together. Add a couple of MB of normalization values and you get the 178 MB on disk.
Ternary weights also make the math simpler: multiplying by −1, 0 or +1 is just adding, subtracting or skipping. Moondream's Photon engine is built around that, with ternary matrix kernels for CPUs with AVX2, AVX-512 or AVX-VNNI, and a CUDA path documented for NVIDIA Ampere or newer GPUs. The model is released under CC-BY-4.0.
How I tested it
I wrote the method down before collecting a single result and logged every decision after that. The other nine systems:
- Parakeet TDT 0.6B v3 at 16-bit (1,256 MB) and standard 8-bit (740 MB): Redux's original.
- Parakeet Unified EN 0.6B, an English-only model that can stream (731 MB, 8-bit).
- Whisper large-v3 (1.55B parameters, 1,669 MB) and Whisper large-v3-turbo (0.8B, 886 MB).
- Moonshine base, a 61M-parameter model built for edge devices (132 MB).
- Three cloud APIs: OpenAI
gpt-transcribe, ElevenLabsscribe_v2and Google'sgemini-3.5-transcribe.
About 35 minutes of English audio, cut into 147 pieces, each with a reference transcript:
- Real-world: six YouTube clips with human captions: a TED talk, a podcast, an MIT lecture, a news segment, a TEDx talk by an Indian-English speaker, and street interviews.
- Benchmark sets: 50 utterances from LibriSpeech, Earnings-22, AMI meetings and VoxPopuli, from the Open ASR Leaderboard test sets.
- Accents: 12 speakers with 12 different first languages reading the same passage, from the Speech Accent Archive (CC BY-NC-SA 2.0).
- Noise: ten of the clean clips again, with pink noise at 5 dB.
- Dictation: me reading four paragraphs (225 words) into the laptop microphone. I didn't read the script word for word, so its reference is the script corrected to what I actually said, using where the models agreed. That can favour the models, so treat dictation numbers as a hint.
Accuracy is word error rate (WER) after the Whisper text normalizer, on the 78 public clips (5,531 words), with 95% confidence intervals from 10,000 bootstrap resamples. 5% WER means about five wrong words in every 100. Speed comes from 20 short benchmark clips (3 minutes of audio): a warm-up, then the median of three timed passes, one clip at a time.
There were two setups, with the same audio, references and scoring:
- My laptop: an AMD Ryzen 5 5600H (6 cores, 6 threads used), 15.3 GB of RAM and an NVIDIA GTX 1650 4 GB, on Windows 11. The other open models ran in transcribe-cpp with community GGUF conversions, on the GPU through Vulkan. Redux ran in Photon on the CPU only: its GPU path isn't documented for a card this old, so I didn't use it there.
- A cloud T4: a Tesla T4 16 GB with a 2-core Xeon (4 threads) and 32 GB of RAM, on Ubuntu. Here each model ran in the engine people usually deploy it with: Photon for Redux, NVIDIA NeMo for the original Parakeets, faster-whisper for Whisper and Hugging Face transformers for Moonshine. The cloud APIs weren't run again.
Accuracy

Laptop run: 78 public clips, 5,531 words. Redux in green.
On the laptop Redux scored 5.33%, behind its original (4.54%) and Whisper large-v3 (5.04%), and well ahead of Moonshine (7.02%). The top of the table is a five-way tie: ElevenLabs (4.03%), OpenAI (4.09%), Parakeet Unified (4.16%), Whisper turbo (4.34%) and Gemini (4.38%), with confidence intervals that overlap almost completely. With 78 clips, gaps under about a point aren't reliable.
On the T4 the ranking of the open models held: Parakeet Unified 4.25%, Whisper turbo 4.56%, the original Parakeet 4.72%, Whisper large-v3 4.88%, Redux 5.32% on the GPU and 5.50% on the CPU, Moonshine 7.07%.
One result about compression on its own: the 8-bit original scored 4.52% against 4.54% for the 16-bit one. In this sample, going from 16-bit to 8-bit cost nothing measurable. Going from 8-bit to ternary is where a cost shows up.

Weights on disk against WER. The grey band is the three cloud APIs, whose sizes aren't published.
Plotted against size, Redux is the odd one out: everything in its accuracy range is four to nine times bigger. Size in general predicts little here; the 1.67 GB Whisper large-v3 scored worse than three models half its size.
Redux against its own original

Same audio, same architecture. On the laptop the original ran as GGUF; on the T4, in NeMo.
This is the comparison the whole test was built around. On identical clips Redux made +0.80 more errors per 100 words than its original on the laptop (95% interval −0.08 to +1.54) and +0.60 on the T4 (−0.25 to +1.26). The two setups agree on the direction and roughly the size. But both intervals include zero, and both reach past the +0.5 points I'd set in advance as "no meaningful loss", so 78 clips can't say whether the true gap is tiny or about a point and a half.
Where the gap lands is clearer than its size. On clean benchmark clips it was small, +0.2 on the laptop and +0.4 on the T4. On the 12 accented speakers it was about +1.5 on both, and on my dictation +1.8 and +2.7. Its mistakes were phonetic: "flight lines" for "lands", "revenue grow" for "grew". The model card reports a 6.55% average on seven Open ASR Leaderboard sets and no accent breakdown; on that kind of audio the gap here was small too.
Two setups

The T4 numbers come from two runs on the same machine type.
- It runs on the T4. Photon documents its CUDA path for Ampere or newer, and the T4 is an older Turing card. It worked anyway, and fast. That's one card, though, not a promise for every Turing GPU.
- The server CPU was slower than my laptop. Redux ran at 7.7× on the 2-core Xeon against 15.1× on six laptop cores. The Xeon has AVX-512 but not VNNI, so Photon used the same AVX2 kernel as on the laptop, and its fastest CPU path is still untested here.
- Rankings shift with the machine. On the laptop the 8-bit original was faster than the 16-bit one; on the Xeon it was the other way round (4.8× against 6.6×). The machines differ in cores, clocks, engines and operating system, so I can't say which of those caused it.
- The text changes a little too. Redux's transcripts were word-for-word identical on 134 of 147 pieces between the two CPUs, and on 127 of 147 between the Xeon's CPU and the T4. Accuracy stayed about the same, but don't expect the exact same output on every device.
Where it holds up, and where it slips

Laptop run. Each column has 147 to 715 words, so read one column as a hint.
- Strong: earnings calls (5.5%, level with the best system), the podcast (3.0%), clean audiobooks (1.1%), meetings (8.2%) and noise (1.8%).
- Weak: accented speakers (3.6% against 2.2% for its original on the laptop), my dictation, and the street interviews (11.9%, among the worst).
- No system wins everywhere. ElevenLabs led on meetings, Whisper turbo on the lecture, Whisper large-v3 on accents and dictation, and Gemini in noise. Whisper large-v3 also lost points on the podcast for a reason that isn't really a mistake: it drops repeated fillers like "you know", which the captions keep.
Every one of the ten systems wrote the name "Ananya" in my dictation as "Aranya". I can't tell whether that's the models or my pronunciation, so it stays in the reference and counts as an error for all of them.
Speed

Seconds of audio per second of compute; median of 3 timed passes.
On the laptop CPU Redux ran at 15.1× real time, about four minutes for an hour of audio. That's 1.4× the 8-bit original in the same Linux container (10.7×) and twice the 16-bit one (7.6×), which makes it the fastest Parakeet on that CPU. It wasn't the fastest model overall: Moonshine ran at 20.3×, with clearly more errors. Whisper was slower than real time on this CPU.
On the T4 Redux was the fastest system I tested, at 85.0× against 73.8× for its original in NeMo and 46.1× for Parakeet Unified. At that rate an hour of audio would take about 42 seconds. That's worked out from short clips run one at a time, not a timed hour-long job.
Latency and memory

One 4 to 25 second clip at a time with the model loaded. Cloud times include the upload from my connection.
For dictation, the wait per sentence matters most. Redux took a median 0.09 s on the T4 and 0.48 s on the laptop CPU, quicker than every cloud API from my connection (1.03 s for ElevenLabs, 1.05 s for OpenAI, 3.03 s for Gemini). Whisper always processes a 30-second window, so even a short sentence took 11 to 14 seconds on the laptop CPU.
Memory is where the 178 MB figure misleads. Redux used about 980 MB of RAM on the laptop and 1.35 GB on the Xeon, not far from the 8-bit original. On the T4 it took about 2 GB of GPU memory, less than its original's 2.9 GB but nothing like the file size. The small download is real; a tiny runtime isn't, at least not with today's engine.
What it costs in real use

Two everyday jobs, with cloud list prices on 5 October 2026.
Two everyday jobs make the trade clearer. Dictating half an hour a day for a year is about 180 hours of speech: $40 with ElevenLabs, $49 with OpenAI and roughly $68 with Gemini at their list prices on 5 October, or no per-hour price with a local model. Transcribing an archive of 1,000 hours costs $220 to $370 in the cloud. With Redux the same archive would take about 66 hours on my laptop's CPU, or about 12 hours on a T4. Check current prices before you rely on these; they change.
- OpenAI gpt-transcribe: $0.27 an hour. Per OpenAI's data controls, API audio isn't used for training by default.
- ElevenLabs Scribe v2: $0.22 an hour. Audio is retained by default; zero retention is for enterprise plans.
- Gemini 3.5 Transcribe: billed per token rather than per minute, which works out to roughly $0.37 an hour by my estimate.
- Local models: no per-hour price, and the audio stays on your own machine. That isn't the same as free; I didn't measure hardware, server time or electricity.
Every model at a glance

Accuracy from the laptop run; speed in multiples of real time; the wait is the median on the laptop.
When it's worth using
- English transcription on a CPU, on Linux. This is where it's strongest: the fastest Parakeet on both CPUs I tested, from a 178 MB download. Plan for about 1 to 1.4 GB of RAM and a small accuracy cost.
- Apps that ship or download the model. 178 MB against 740 MB for the 8-bit original and 1,256 MB for the 16-bit one. I didn't measure the size of the engine itself, which you'd ship too.
- GPU batch jobs. On the T4 it was the fastest system here, with somewhat lower accuracy than its original. It worked on this card despite not being documented for it; test yours.
- Clean audio in bulk. The gap on benchmark-style clips was small. I didn't test many requests at once, so measure throughput under your own load.
Where to be careful:
- Accented speakers, names and dictation. That's where the extra errors were. Test on the people who'll actually use it; 12 accented speakers and one dictating voice is a small sample.
- Windows. On my AVX2 laptop the Windows build of Photon had no AVX2 kernel for this CPU, so Redux wouldn't run natively; it ran fine in a Linux container on the same machine. That's one version on one CPU, and may well change.
- Small-memory devices. Phones and boards with little RAM weren't tested, and the 1 GB-plus runtime is the thing to check first.
- If every word matters and you have the memory, the 8-bit original was more accurate, showed no measurable loss against 16-bit here, and runs in more places today.
What I took away
Ternary weights work. A 0.6B speech model in 178 MB, faster than its original on every device I tried, with a cost you mostly won't notice on clean audio. That's a real step for on-device speech.
Averages hide where compression hurts. The loss was small on benchmark-style audio and larger on accents and dictation, on both setups. If you ship a model like this, test it on your own speakers, not on leaderboard averages.
A small file isn't a small runtime. 178 MB of weights still needed a gigabyte of RAM on a CPU and two of GPU memory. Download size and memory are different budgets.
The engine matters as much as the weights. Most of Redux's practical limits were packaging, not the model: no native Windows kernel for my CPU, GPU support documented only for newer cards, and a runtime that uses most of a gigabyte. The T4 result shows the documented limits aren't the whole story either.
For English, local models have caught up with the cloud. On the laptop the best open models tied three commercial APIs on accuracy, and Redux sits within about a point of them while running offline with no per-hour price. The cloud still offers languages, speaker labels and other features I didn't test.
Limitations
- Small samples. 78 public clips, 12 accented speakers who read the same passage, and 225 dictated words from one speaker. The ten noisy clips are copies of clean ones, which the intervals don't account for, so treat them as rough.
- English only. The original Parakeet TDT v3 is multilingual; I didn't test other languages.
- Engines are mixed in. Redux in Photon against its original in NeMo or GGUF compares two deployments, not compression alone. Start-up times weren't measured under equal conditions, so I haven't compared them, and memory figures are sampled while running, not exact.
- Hardware I didn't have. No CPU with VNNI, no Ampere-or-newer GPU, no Apple Silicon or ARM board.
- Cloud results are from 5 October 2026, from one network, with that day's models and prices.
What I'd test next
- A CPU with VNNI and a GPU Photon officially supports, recording which kernel and precision it actually uses.
- Apple Silicon and an ARM board, where small models matter most.
- Accents at scale, 100+ speakers, and the other languages Parakeet supports.
- Native Windows on an AVX2-only CPU, once a new Photon build is out.
Thanks to Moondream for Parakeet Redux, NVIDIA for Parakeet, OpenAI for Whisper, Useful Sensors for Moonshine, the Open ASR Leaderboard maintainers for the test sets, and Steven Weinberger and George Mason University for the Speech Accent Archive.
Originally published at abhashchakraborty.tech.
Top comments (0)