DEV Community

Cover image for Building Ambient: 8 platforms with Kotlin Multiplatform
Hayami Shuhei
Hayami Shuhei

Posted on

Building Ambient: 8 platforms with Kotlin Multiplatform

I wanted to make an ambient sound app that stayed fresh every time I pressed play. After listening to the same recordings again and again, I started to remember what came next. That made it harder to relax.

So I built Ambient. It generates rain, waves, wind, and other environmental sounds in real time, with visuals that respond to the sound. The idea is simple. Press play and let the soundscape unfold while you work, rest, or settle down for the night.

Tap play to dissolve the arrow into flowing ink

You can try it in your browser or get the app for your device.

Underneath those different interfaces is a shared Kotlin engine. Ambient now has implementations for Android, iOS, macOS, watchOS, visionOS, Windows, Linux, and the web.

Ambient on phone, watch, desktop, TV, and VR

The central design decision was to share the sound engine and playback behavior while building an interface for each platform. The same Kotlin code defines how a soundscape changes over time and how playback responds to commands. Each platform app connects that code to its audio system and user interface.

Supporting visionOS required an additional step. I extended Kotlin/Native so that the shared engine could compile for the headset and its simulator. This article explains the application architecture, the compiler work, and how the two connect.

Sharing the sound engine across platforms

The shared Kotlin engine defines each scene, generates its sounds, manages playback, and supplies the data used by the visuals. A scene describes a soundscape made from several sound layers. KMP Procedural Audio handles playback, source switching, and connections to each platform's audio system.

Android runs the shared code on the Java Virtual Machine (JVM). Kotlin/Native compiles it into native code for Apple platforms, Windows, and Linux. Kotlin/JS compiles it to JavaScript for the browser.

Shared Kotlin core with platform audio output and visual renderers

Platform Shared engine bridge Interface, audio, visuals
iOS / iPadOS Kotlin/Native framework SwiftUI, AVAudioEngine, Metal
Android / Android TV Kotlin/JVM module Android Views, AudioTrack, Vulkan or OpenGL ES
watchOS Kotlin/Native framework SwiftUI, AVAudioEngine, Canvas particles
visionOS Custom Kotlin/Native target SwiftUI, AVAudioEngine, RealityKit and Metal particles
macOS Kotlin/Native C bridge SwiftUI, AVAudioEngine, Metal
Windows Kotlin/Native C bridge Win32, WASAPI, Vulkan
Linux Kotlin/Native C bridge GTK4, ALSA, Vulkan
Web Kotlin/JS in AudioWorklet HTML controls, Web Audio, WebGPU

The iPad uses the iOS app, and Android TV uses the Android APK with remote controls. Web playback controls remain available without WebGPU.

This boundary keeps sound generation and playback behavior in one place. The library sends audio samples to the output device. Each platform app handles interruptions, background playback, and system controls. A change to a sound generator does not require a separate implementation for every operating system.

The desktop apps call the Kotlin/Native engine through a small C interface. This lets their Swift or C++ code send playback commands and read data for the visuals. The shared engine uses the library to send samples to the platform audio system. macOS uses the same C interface as Windows and Linux, while the iOS, watchOS, and visionOS apps import a Kotlin framework directly from Swift.

Generating sound in real time

The engine generates stereo pulse-code modulation (PCM) audio at 48 kHz, or 48,000 samples per second for each channel. A scene combines continuous sounds, such as wind, with shorter events, such as bird calls. Noise generators and oscillators produce the signal. Filters shape its tone, while envelopes control how each sound starts and fades. Slowly changing parameters keep the soundscape evolving.

That approach removes the need to download or loop recordings. It also means the engine must finish each block of samples before the audio system needs the next one. Missing that deadline can cause clicks or gaps in playback.

The synthesis loop reuses sample buffers and the state of each active sound. It does not wait for a graphics frame. A known random seed lets tests reproduce the same sound sequence, while listening sessions can start with different seeds.

During a scene change, two audio renderers run together. One produces the outgoing scene and the other produces the incoming scene. An equal-power crossfade reduces one as it brings in the other, helping avoid a dip in perceived volume when the sounds are unrelated.

Replacing an audio source is a separate operation. It uses a short linear crossfade, where one source fades out as the other fades in at a constant rate. This smooths the switch to a sound preview without forcing it to use the longer transition between scenes.

Letting audio drive the visuals

The synthesizer already knows which sounds are active and how their energy changes. It publishes that information directly instead of making every renderer infer it from the final waveform.

The shared visual data identifies the active sounds, their relative contributions to the mix, current energy, and progress through a scene change. The audio engine writes these values into preallocated numeric storage. Each native app reads a snapshot for its visual renderer. In the browser, the audio code sends snapshots to the page less often than it produces audio blocks.

This gives the visual renderer useful information while keeping audio generation independent of UI objects and drawing speed.

The flowing ink is based on Jos Stam's Stable Fluids, using semi-Lagrangian velocity advection and pressure projection. Slow background currents keep it moving between sound events. I added conservative forward transport for pigment to preserve its amount during advection and keep thin ribbons sharper. Fading and absorption at the outer edge gradually remove the ink.

My custom color model stores RGB weighted by pigment density, giving overlapping colors a weighted average. An extra field carries density-weighted squared color values, allowing the renderer to detect color variance. Where different colors mix, it can apply a small brightness lift. A single pigment color receives no lift. Density controls background coverage rather than directly darkening the pigment color.

The graphics implementation varies by platform. iOS and macOS use Metal, while visionOS combines Metal with RealityKit. Android uses Vulkan with an OpenGL ES fallback, and Windows and Linux share the Vulkan renderer. The browser uses WebGPU. Shared C++ code and GPU compute shaders simulate the flowing ink. Each graphics backend manages its own GPU resources and draws the result.

The watch generates its own audio and draws a smaller particle scene with SwiftUI Canvas. When the app is visible, its animation targets 15 updates per second. Reducing visual work lets it reuse the sound engine while limiting battery use.

Running the engine in a browser

The browser has a different execution model, so the location of the synthesizer matters.

The Kotlin/JS engine and playback controller run inside an AudioWorklet, which processes audio separately from the browser's main UI thread. The page sends playback commands and receives playback state and data for the visuals. The worklet generates the samples directly, so the page does not have to produce and transfer every audio block.

The worklet uses the browser's audio clock to time scene changes. Hiding the page stops visual updates, while audio generation and scene changes can continue without waiting for the page to draw.

A small C++ module is compiled to WebAssembly, which lets the browser run it alongside JavaScript. It schedules the GPU work for the ink simulation. Audio generation remains in Kotlin/JS.

Adding visionOS to Kotlin/Native

Kotlin/Native already supported iOS and watchOS, which made adding visionOS relatively straightforward. I could reuse the Apple runtime and the code that exposes Kotlin to Swift through a framework. Most changes extended that support for the new platform.

The Kotlin fork adds targets for the headset and the simulator on Apple Silicon Macs. Each uses its own Apple SDK, which supplies the headers and libraries needed to compile system APIs. The target triple identifies the processor architecture, Apple platform, and whether the binary runs on a device or in the simulator.

Kotlin/Native target Apple SDK Target triple
visionos_arm64 XROS arm64-apple-xros
visionos_simulator_arm64 XRSimulator arm64-apple-xros-simulator

I updated runtime platform checks, linker settings, and framework metadata for visionOS. I also extended Gradle so the targets can share Apple source sets and package device and simulator frameworks together. The API compatibility tools now recognize both targets.

After Ambient's first visionOS release, I rebuilt the Kotlin fork with reorganized commits and Xcode 27 support.

Making Apple APIs available to Kotlin

Kotlin/Native needs Kotlin declarations for the Apple APIs that the engine calls. These declarations are generated from SDK headers and packaged as platform libraries.

I started with Foundation and UIKit, together with their dependencies. I also extended SDK scanning so the platform libraries can be regenerated from the device and simulator SDKs.

For audio playback, the fork includes AVFAudio, AudioToolbox, CoreAudio, CoreAudioTypes, and CoreMedia. These bindings let KMP Procedural Audio's Apple backend use AVAudioEngine on visionOS.

Linking Premium across devices

An iOS or Google Play Android purchase can unlock Premium on Windows, Linux, and the web. The receiving device shows a QR code. The user scans it in the mobile app, checks the device name, and approves the link. There is no separate account to create.

This flow has three parts. Shared Kotlin code controls the linking process. Each app connects that code to its operating system. A server verifies the purchase and grants access to the receiving device.

Sharing the linking logic

The Kotlin Multiplatform core handles both sides of the exchange. On the receiving device, it requests a pairing code, waits for approval, saves the resulting access, and checks when it expires. In the mobile app, it prepares purchase verification requests, manages approval, and keeps track of linked devices so users can remove them.

The core formats API requests and validates the JSON responses. It does not send network traffic itself. A small platform adapter sends each request and passes the result back. The shared code then decides what should happen next. This keeps the sequence of steps and the rules for expiry consistent across apps, even though their networking libraries differ.

The same Kotlin source reaches each platform in a different form. Android uses it directly. Kotlin/Native produces a framework that Swift can call in the iOS app, and a C interface that the Windows and Linux apps can call. Kotlin/JS produces JavaScript that runs in the browser. These bridges let each interface use the shared workflow without reproducing it in another language.

Connecting to the operating system and the server

Platform adapters handle the screen, QR scanning, HTTP communication, local credential storage, and obtaining purchase proof from StoreKit or Google Play. Credentials use Apple Keychain, AndroidKeyStore encryption, Windows DPAPI, a private Linux settings file, or the browser's localStorage, depending on the platform.

A Cloudflare Workers service verifies the purchase proof and checks that it matches the purchase behind an active Premium entitlement in RevenueCat. D1 stores device registrations and limits each eligible purchase to three linked devices or browser profiles. The server is responsible for granting access. The client is responsible for using that access within its allowed lifetime.

The QR code contains a pairing code that is valid for five minutes. It does not contain the device's access token. A separate token lets the receiving device check whether approval has arrived and refresh its access afterward.

Handling delays and expired access

The shared core checks for approval every three seconds, then refreshes linked access every minute. If the network becomes unavailable, previously verified Premium lasts only until the server's access deadline, at most 24 hours. Once access expires or a check reports that Premium is no longer active, playback returns to the free 60 minute session limit.

Late responses from an earlier pairing cannot overwrite a newer one. If credentials cannot be saved, the core keeps the previous pairing and asks the server to clean up the unsaved registration. Tests use simulated responses and timestamps to cover these cases across the shared logic.

Store checks and network requests run separately from sound generation. In the browser, linking runs on the page while the sound engine runs in an AudioWorklet. The playback engine receives the resulting access state, so a slow purchase check does not interrupt the audio.

Demo QR code for linking Premium to another device

Sharing the playback layer as open source

I extracted the reusable PCM playback layer into KMP Procedural Audio. It is a lightweight Kotlin Multiplatform audio playback library released under the MIT license.

The library exposes AudioPlayer for playback and a PcmSource interface for supplying audio. A PcmSource fills a reusable buffer with 48 kHz stereo floating-point samples. AudioPlayer sends those samples to the platform audio system and applies a short crossfade when the source changes.

Its targets cover Android, iOS, macOS, watchOS, Windows, Linux, and the web. Desktop playback uses Kotlin/Native. The browser implementation runs synthesis inside an AudioWorklet.

The library supports seven platform families. Ambient also compiles its common and Apple sources for visionOS with the custom toolchain described above.

Another app can use this playback layer with its own sound generators. It does not need Ambient's soundscape model or visual renderer.

Top comments (0)