The Compression slider in my app moved the output by exactly 0.00 dB. Not "a bit too little". Zero, at two of three input levels, on real presets, and none of my existing checks caught it.
It took five rounds to fix. The first four times I declared it done, the person listening in headphones heard nothing change. This post is about that bug and a few of its relatives. I hit all of them while building a virtual microphone for macOS: an app that takes your real mic, processes it in real time, and shows up in Zoom, OBS or Discord as one more input device.
The surprising part is where the difficulty sat. The driver, the bit everyone warns you about, was the easy half.
The virtual device is a loopback, and Core Audio is the transport
A virtual mic on macOS is an AudioServerPlugIn: a bundle in /Library/Audio/Plug-Ins/HAL that coreaudiod loads and publishes as a device. No kernel extension, no DriverKit.
I did not write it from scratch. Apple publishes a sample called NullAudio in Creating an Audio Server Driver Plug-in. My build script downloads the pinned zip, checks its SHA-256, applies a 134-line patch and compiles one C file with clang. The repo never commits the generated driver source. You edit the patch.
The patch does three things: renames the bundle and the device, sets stable UIDs, and adds a ring buffer of 16,384 stereo Float32 frames, indexed by sample time:
// WriteMix: what a client plays into the device's output side
ring[(sampleTime + i) % 16384] = frame;
// ReadInput: what Zoom records from the device's input side
frame = ring[(sampleTime + i) % 16384];
The IO callback does no allocation and takes no locks. The whole trick is the loopback. The device has an output side and an input side. My processing engine plays the finished voice into the output side as an ordinary Core Audio client, and Zoom records from the input side. There is no shared memory, socket or XPC between the engine and the driver. The HAL itself carries the audio.
That keeps the code running inside a system daemon tiny and dumb. The plug-in loads no models, touches no network and reads no user files. Everything interesting happens in a normal user process, where a crash is my problem and not the system audio daemon's.
The measured loopback latency, write to read, is 512 frames, or 10.667 ms. A small verifier writes an impulse at frame 4800 and a 997 Hz tone into the output scope, then reads them back from the input scope. The acceptance run fails if latency reaches 20 ms.
One thing I would tell anyone starting out: freeze the device UID on day one. When the product was renamed, the UID stayed. Change it and Zoom forgets the selected mic. Worse, if the renamed bundle installs under a new file name, the old one keeps loading too, and you publish two devices with the same UID.
Three processes, not one
The app is split into three processes:
- The GUI, a SwiftUI app that owns no audio at all.
- An engine agent under launchd (
RunAtLoad,KeepAlive). It holds the physical mic, runs the realtime DSP and writes to the virtual device. It keeps running when the window closes, because your call does not end when you close a settings window. - The HAL plug-in.
The GUI talks to the agent over a narrow XPC protocol of 13 methods: ping, configure, start, stop, apply a preset, snapshot, plus a few monitor, playback and calibration calls. Presets travel as JSON.
Realtime code is C++20. Offline code (the voice analyzer and the preset generator) is Swift: pure functions, easy to test. Inside the agent the format is 48 kHz Float32 mono, with a preferred HAL buffer of 128 frames. Every buffer is allocated up front. That includes the delay line for the headphone monitor, sized for the worst-case latency of every stage, because growing it later would mean allocating on the render thread.
The split hides a trap. A buffer bound lives in C++, while the UI that sets it lives in Swift. Raise the calibration recording from 20 to 60 seconds in Swift alone, and it compiles fine and then fails at runtime. When a limit sizes a preallocated array, the C++ side has to move with the Swift side.
Between the agent and the device sits a lock-free single-producer, single-consumer ring of 262,144 frames. The mic and the virtual device run on different clocks, so they drift. The reader keeps a reserve of 256 frames with a tolerance of ±64. When the fill drifts outside that window, it reads one frame more or one fewer and linearly stretches the block back to the requested length. A one-frame stretch now and then is cheap. A ring that slowly runs dry is not.
Bug 1: a mic that was silent, with no error anywhere
Symptom: with some mics the level meter sat at minus infinity. The engine looked like it was running. Nothing in the logs.
I first filed it as "Bluetooth headsets don't work". Wrong. The real cause: I opened the input audio unit at 48 kHz regardless of the hardware rate and let it convert. Every property call returned noErr. Then every single AudioUnitRender failed with kAudioUnitErr_CannotDoInCurrentContext.
The unit's internal converter wants an exact client frame count, but the input callback receives device frames. For 16 kHz you can scale, since it is an integer ratio. For 44.1 kHz you cannot: one frame short returned -10863, one frame long returned -10874, alternating.
The fix: open the unit at the hardware rate and resample with an AudioConverter inside the callback, skipping it when the rates already match. To prove the cause was the rate and not Bluetooth, I forced the MacBook's built-in mic to 44.1 kHz. Before the fix it was silent too. After, it read -52.0 dB.
Lesson: when a Core Audio setup call "succeeds", believe the render callback, not the setter.
Bug 2: the slider that did 0.00 dB
This is the expensive one.
A compressor has an absolute threshold in dBFS. The design assumed speech would reach the compressor at around -20.2 dBFS. I measured where it actually arrived on the built-in mic: -37.8 dBFS. That is 17.6 dB below the level the whole stage was designed around. A compressor whose input never crosses the threshold does nothing, whatever ratio you set.
Where did the level go? The per-voice EQ alone took 8.8 dB, filters 2.4 dB, room handling 1.1 dB. And the only stage that raised the signal back to the working level sat at the end of the chain, after everything that depended on it. The dynamic EQ, pop cut and a parallel density stage had the same disease.
The fix was structural, not a new constant. The working point became its own stage, right after the noise gate. Everything with a threshold relative to the raw mic sits above it. Everything with an absolute threshold sits below it. Gain became a pure volume control. Later a broadcast-style leveler replaced the static trim: it decides on speech before the gate, freezes during pauses, and moves at most 6 dB per second.
So why five rounds? Because I kept declaring it fixed on numbers, not on what the listener heard. Round one moved the output by 0.00 dB on real presets. Round two managed -0.8 dB, round three "half a notch", round four 0.01 dB. Only round five produced something a person could hear: Ratio pulled the level down by 4.0 to 6.3 dB, and sibilants dropped 12 dB.
The check that finally held is a sweep. It drives the real C++ stages at three input levels, moves the control across its range, and requires a monotonic change with a real step size. On the old code it is red. That last part became a rule: a new check has to fail on the old code before I trust it to pass on the new one. I swap the old file in with git show <sha>:path, build, run, and restore.
Bug 3: nine green tests and a de-esser that never fired
Same disease, different organ. The de-esser's output was bit-identical at every slider position. Its threshold (target minus 7, so -22 dBFS) assumed auto-level had already lifted the signal. On a real minute of speech, the 7.2 kHz band peaked at -37 dBFS. Fifteen dB short.
Nine sibilance tests were green the whole time. All nine checked what the preset generator wrote. None checked that the stage fired.
After moving the stage, a second bug showed up in the reduction curve itself:
// before: a ceiling the signal never reached
gr = std::min(depth, excess * 0.65f);
// above ~3 dB of depth, every setting gave the same 1.20 dB
// after: keeps growing with depth
gr = depth * (1.0f - std::exp(-excess / 6.0f));
// 0 / 1.20 / 2.27 / 3.20 dB at depth 0 / 3 / 6 / 9
The pop filter had the same dead shape. The threshold also became relative to a slow level envelope, so the volume knob stopped secretly driving the de-esser.
Bug 4: the numbers passed and the ear did not
I had a room-reverb reduction stage with acceptance criteria declared before measuring, which is good practice. It passed: the tail 200 ms after a word dropped by 7.5 to 9 dB, and loudness moved by no more than 0.6 dB. The person listening heard no difference at all.
They were right. On a laptop mic, what you hear as "room" is not the late tail. It is early reflections in the first 50 to 100 ms, inside the words. My metric measured the tail. The point 50 ms after a word barely moved.
The underlying physics is distance. The direct-to-reverberant ratio is around +19 dB at the mouth and around -1 dB at 50 cm. A single-channel classic DSP stage cannot put that distance back. What did help was a neural voice-isolation stage, picked from six candidates in a loudness-matched A/B test by ear.
Two rules came out of it. Numbers only screen candidates; a loudness-matched listening test is the final gate. And test on the worst mic first (the MacBook's built-in one) with at least two different recordings, because fixtures lie. A sine "vowel" has no formants, and white-noise "room tone" reads as a fricative.
The latency budget I planned and the one I chose
The plan said under 20 ms at P95, mic to virtual output. The honest status: I never measured that number end to end, and the product does not meet it.
What I do have is per-stage numbers. The mic gets a requested 128-frame HAL buffer (2.67 ms). The driver loopback measures 10.67 ms. The drift reserve adds about 5 ms. The whole default DSP chain costs 1.33 ms, and that is just the limiter's lookahead. Every shaping stage is a filter or a gain, so none of them buffer.
Then comes the voice-isolation stage that won the listening test. On its own it costs several tens of milliseconds, and I took that trade on purpose. What I did not want was a hidden cost. So the app shows each stage's latency on screen, and three screens had to be fixed because they kept charging 10 ms for a stage that was switched off.
Small bugs worth stealing
-
std::vector<float>{1024, 0.0f}is a two-element initializer list, not 1024 zeros. An envelope was being estimated from heap garbage. AddressSanitizer found it in one run, after three measurements had blamed working code. -
std::clampdoes nothing to NaN. A non-finite knob stayed on and inert. - There was no limiter at first, only
std::clampon the output. Real true peaks on four mics were 6.9 to 9.3 dB above the frame-peak figure the headroom math used. A proper limiter (1.33 ms lookahead) plus a corrected headroom formula took the worst overshoot from +8.34 dB to +0.33 dB. - A USB headset delivered a DC offset that was 36% of the recording's energy. It dragged the measured spectral centroid down to 90 Hz. Subtracting the mean fixed it. A 20 Hz one-pole high-pass would have cost a low-frequency test its margin.
- Loudness follows ITU-R BS.1770. Checking it against the EBU Tech 3341 anchor (a 1 kHz sine at -23 dBFS should read -23 LUFS) caught a 3 dB error. Published targets assume two channels, so a mono signal counted once lands every voice 3 dB loud.
- Polling a meter at 20 Hz through an
ObservableObjectre-rendered unrelated SwiftUI screens. Fixing that took idle CPU from 86% to 1%. -
launchctl bootoutreturns before launchd has actually removed the job, so an immediatebootstrapfails with "5: Input/output error" and leaves no agent. Retry every 100 ms for up to 5 seconds. - After an app update,
KeepAlivekept the old engine agent running, so users heard last version's DSP. The agent now reports its build inping, and the GUI kickstarts it on a mismatch. Read that build string at init: a lazy static reads the new Info.plist from disk and lies. - The Swift test suite takes 643 s in debug and 11 s in release with identical results. DSP tests run with
-c release.
What I would do differently
Put gain staging on paper before writing a single stage. The two worst "this control does nothing" bugs in this project were the same bug: a stage with an absolute threshold sitting where the signal never reached it.
Write the sweep before the feature. A control that cannot show a monotonic, audible step across its range on a real recording is not done, however green the unit tests are.
And let a human ear veto the metric, early. Five rounds of confident numbers lost to one listener who could not hear a difference.
Disclosure: this is the engine behind TunedMic, a Mac app I sell. You read a passage aloud for a minute, it measures your voice, mic and room, and builds the processing chain described above for that combination. If you want the non-engineering version of what a virtual mic is for, there is a plain-language explainer on the TunedMic blog. Happy to answer Core Audio questions in the comments.
Top comments (0)