<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: zhiyi guo</title>
    <description>The latest articles on DEV Community by zhiyi guo (@fawinell).</description>
    <link>https://dev.to/fawinell</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4122754%2F06d68f93-ce27-4fac-9927-6282bef5d6b9.png</url>
      <title>DEV Community: zhiyi guo</title>
      <link>https://dev.to/fawinell</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/fawinell"/>
    <language>en</language>
    <item>
      <title>Inside murmur: prefetching, interruptions, and audio scheduling</title>
      <dc:creator>zhiyi guo</dc:creator>
      <pubDate>Wed, 16 Sep 2026 10:09:04 +0000</pubDate>
      <link>https://dev.to/fawinell/inside-murmur-prefetching-interruptions-and-audio-scheduling-1h1i</link>
      <guid>https://dev.to/fawinell/inside-murmur-prefetching-interruptions-and-audio-scheduling-1h1i</guid>
      <description>&lt;p&gt;murmur is a companion radio that runs in a terminal. It talks and picks songs without waiting for a question. The listener can type a line, get a spoken reply from the host, and then let the show continue.&lt;/p&gt;

&lt;p&gt;Writing a segment, synthesizing speech, and finding music all take time. If preparation starts only after the previous segment ends, the listener hears a gap. An interruption can make already prepared content obsolete. Music adds another uncertainty: finding a song does not mean its audio source will play.&lt;/p&gt;

&lt;p&gt;A local scheduler, the Director, handles these cases. It prepares content during playback, updates the queue when the listener interrupts, and sends playable audio to the mixer. The model writes segments and selects content; it does not directly control the speakers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the waiting happens
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;Generate text → synthesize speech → play it → generate the next segment.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Running these steps serially means waiting for the model and the speech service between every pair of segments.&lt;/p&gt;

&lt;p&gt;In two historical full runs, preparing the first batch of two segments took 24.5 and 33.9 seconds, from text generation through completed speech synthesis. In another log, generating one refill segment took 9 to 14 seconds for the model call alone, before speech synthesis.[1]&lt;/p&gt;

&lt;p&gt;What the listener notices depends on when that work happens:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Moment&lt;/th&gt;
&lt;th&gt;What the listener notices&lt;/th&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Opening the radio&lt;/td&gt;
&lt;td&gt;Why hasn't anyone started talking?&lt;/td&gt;
&lt;td&gt;Startup to first audible voice&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Moving between segments&lt;/td&gt;
&lt;td&gt;Why did it stop?&lt;/td&gt;
&lt;td&gt;Extra wait beyond the configured pause&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typing a message&lt;/td&gt;
&lt;td&gt;When will it answer me?&lt;/td&gt;
&lt;td&gt;Input to the start of reply playback&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Preparing ahead can reduce the wait between segments. The first batch still has to be generated after startup. An interruption also needs a fresh generation request: the old program can keep playing while that request runs, but the reply still takes time. These are three separate latency measurements.&lt;/p&gt;

&lt;h2&gt;
  
  
  The local program owns the schedule
&lt;/h2&gt;

&lt;p&gt;The main application is TypeScript running on Node.js. Its terminal UI runs in a separate process using Bun and OpenTUI. Model inference and production speech synthesis depend on external services.&lt;/p&gt;

&lt;p&gt;Figure 1 groups the components by responsibility: blue for local orchestration, yellow for content preparation, purple for context, and green for audio. The arrows show interactions; preparation tasks can run concurrently.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8cphu26h3lm2tazqqiro.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8cphu26h3lm2tazqqiro.png" alt="murmur architecture: the Director connects the terminal UI, context, model tasks, speech synthesis, music preparation, and audio engine" width="800" height="491"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 1. The Director controls when content airs; AudioEngine controls audio state. Music preparation also calls the model to select songs. That work is grouped inside Music here.&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The model submits results through tools
&lt;/h3&gt;

&lt;p&gt;Brain uses the Claude Agent SDK. Each task gets its own tool set. The harness disables the &lt;code&gt;CLAUDE.md&lt;/code&gt; files, skills, and arbitrary MCP tools from the listener's everyday coding environment so those settings do not affect the radio.[2]&lt;/p&gt;

&lt;p&gt;For example, the model submits a batch of segments through &lt;code&gt;emit_talk_beats&lt;/code&gt;. The program validates the tool arguments against a schema and reads the text, without having to extract a result from free-form prose. When the listener asks for another song, the model calls a music-switching tool. A callback owned by the Director performs the action.&lt;/p&gt;

&lt;p&gt;Song introductions follow the same separation of responsibilities. The Director starts the source and waits for the engine to confirm that actual audio has been queued. Only then does it update "now playing," record the song, and play the introduction. This avoids introducing a song whose source never started. The check reaches as far as audio scheduling; it cannot prove that sound came out of the speakers.[3]&lt;/p&gt;

&lt;h3&gt;
  
  
  What goes into each generation request
&lt;/h3&gt;

&lt;p&gt;Before each request, the program assembles a bounded context package:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The persona describes the host's identity. Once created, everyday conversations do not automatically rewrite it.&lt;/li&gt;
&lt;li&gt;The listener profile holds selected, consolidated long-term information. Executed one-off control instructions are not directly learned as preferences.&lt;/li&gt;
&lt;li&gt;Recent conversation keeps the current thread going. Earlier material can be retrieved when needed.&lt;/li&gt;
&lt;li&gt;Broadcast history records songs, topics, and time-based anchors to help avoid repetition.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These live in local Markdown and JSONL files. Profile consolidation runs in the background without blocking playback. If it fails, the previous profile and processing checkpoint remain intact.[4]&lt;/p&gt;

&lt;p&gt;Selected context is sent to the inference service, and text to be spoken is sent to the speech service. Data still leaves the machine.&lt;/p&gt;

&lt;p&gt;The terminal UI only displays state and receives input. It exchanges newline-delimited JSON with the engine over a Unix socket, with schema validation at both ends. Without Bun, the application can fall back to a plain-text interface using the same scheduler.&lt;/p&gt;

&lt;h2&gt;
  
  
  Preparing the next segment during playback
&lt;/h2&gt;

&lt;p&gt;Suppose preparation begins with &lt;code&gt;R&lt;/code&gt; seconds left until the planned handoff, and the work takes &lt;code&gt;P&lt;/code&gt; seconds. Ignoring playback startup overhead and failure retries, the extra wait at the handoff is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Extra wait at the handoff: W = max(0, P - R)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Reducing &lt;code&gt;P&lt;/code&gt; helps. Starting earlier also helps, because preparation can overlap the current segment's playback.&lt;/p&gt;

&lt;p&gt;Figure 2 assumes a 30-second segment and 12 seconds of preparation for the next one. Both numbers are illustrative.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpzh811tmpdq28c3id0mc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpzh811tmpdq28c3id0mc.png" alt="Prefetch sequence: prepare B during A, then play B at the next boundary" width="800" height="605"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 2. Prepare B while A plays, then play B after A ends. Diagram lengths do not represent elapsed time.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The Director maintains two kinds of prefetch.&lt;/p&gt;

&lt;p&gt;The talk buffer has a target depth of two segments. Each entry holds text and a speech-synthesis Promise that has already started. After consuming an entry, the Director refills the buffer in the background, with at most one refill task in flight. Being in the queue does not mean synthesis is complete; playback may still have to wait for the Promise.&lt;/p&gt;

&lt;p&gt;Music prefetch has one slot. Search, selection, and source resolution run in the background. If the pick is not ready at a planned music boundary, the Director plays another talk segment and checks again at the next boundary.&lt;/p&gt;

&lt;p&gt;A deeper queue could absorb larger variations in preparation time. It would also spend more on generation, and content farther back in the queue could become stale before airing. The current depth has not been established as a global optimum.&lt;/p&gt;

&lt;h3&gt;
  
  
  Prefetched content can become stale
&lt;/h3&gt;

&lt;p&gt;When refilling the talk buffer, the model needs to know what is already queued. Otherwise, it may write about the same topic again. We include queued text in the generation context so new segments follow it. Broadcast history is still updated only when the content airs.&lt;/p&gt;

&lt;p&gt;The transition out of a song needs separate handling. If the next talk segment was written before the song started, its context does not include that song. Playing it several minutes later can abruptly return the listener to an earlier topic.&lt;/p&gt;

&lt;p&gt;After a song starts, the system generates a short closing link, or coda, with the current song in context. It can play over the song's ending or go to the front of the queue after the song finishes.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happens to the queue when the listener interrupts
&lt;/h2&gt;

&lt;p&gt;Suppose the radio is playing A, with B and C buffered. The listener types, "Let's stop talking about this."&lt;/p&gt;

&lt;p&gt;B and C no longer fit. Appending the reply to the queue would make the listener sit through two more segments on the subject they just asked to leave.&lt;/p&gt;

&lt;p&gt;murmur handles it in this order:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Clear the old talk queue and invalidate its in-flight refill work.&lt;/li&gt;
&lt;li&gt;Keep the current audio playing while generating and synthesizing the reply.&lt;/li&gt;
&lt;li&gt;Once the reply audio is ready, stop any remaining old voice playback and start the reply.&lt;/li&gt;
&lt;li&gt;Refill the queue using the updated conversation context.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If another line arrives while the reply is being prepared, it is merged into the reply, and the superseded preparation task is invalidated. An ordinary interruption leaves the song playing. The mixer lowers its volume when the voice comes in.&lt;/p&gt;

&lt;p&gt;Clearing the queue alone is not enough. An old request could return a few seconds later and put obsolete text back into it.&lt;/p&gt;

&lt;p&gt;The code uses an incrementing version number, &lt;code&gt;epoch&lt;/code&gt;, to decide whether a result is still valid. The talk-refill path can be simplified to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;startedIn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;talkEpoch&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;beats&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;generateTalks&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;startedIn&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="nx"&gt;talkEpoch&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="c1"&gt;// Context changed; discard this batch.&lt;/span&gt;
&lt;span class="nf"&gt;enqueue&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;beats&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;An interruption increments &lt;code&gt;talkEpoch&lt;/code&gt;. An old task may keep running, but its result is discarded when the versions no longer match. Interactive tasks use similar checks to prevent later tool calls from a superseded task from changing state.&lt;/p&gt;

&lt;p&gt;This only blocks results that have not yet taken effect. Completed actions are not rolled back, and model requests already sent may continue consuming resources.&lt;/p&gt;

&lt;p&gt;murmur tries to wait until the reply audio is ready before switching, so the old voice may continue for a while after the listener presses Enter. That reduces silence, but the time until the reply is heard still depends on generation and synthesis. Tests need to measure silence and reply latency separately.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mixing voice and music
&lt;/h2&gt;

&lt;p&gt;murmur uses &lt;code&gt;node-web-audio-api&lt;/code&gt; to build one audio graph for voice, the main song, and the background bed. The engine schedules gain changes ahead of time on the audio clock, without relying on JavaScript timers to change volume during playback.[5]&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl18vdk8wlhb3ktvp72jr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl18vdk8wlhb3ktvp72jr.png" alt="Audio graph: voice, song gain, and background bed share one output" width="799" height="219"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 3. Solid arrows carry audio; the dotted arrow shows control. Voice changes the song channel's gain without pausing the song.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The current settings lower the main song to a linear gain of 0.3 over about 0.3 seconds, then restore it over 2.5 seconds after the voice ends. A gain of 0.3 is an amplitude ratio. It does not mean "30% as loud."&lt;/p&gt;

&lt;p&gt;The drop needs to be quick enough to keep the song from covering the voice. A slower recovery avoids a sudden jump in volume immediately after a sentence ends. These values were adjusted by listening; they are not universal.&lt;/p&gt;

&lt;p&gt;The background bed stays steady during speech. It only crossfades when the main song enters or leaves. Making it rise and fall with every sentence would produce audible pumping.&lt;/p&gt;

&lt;p&gt;Speech synthesis currently waits for the service to return a complete audio clip before playback starts. Knowing the clip's duration makes it easier to schedule music recovery, handle interruptions, and join clips. It also means the first line and each reply must wait for the whole clip to be ready.&lt;/p&gt;

&lt;p&gt;Long sources such as songs are decoded and queued onto the playback timeline in chunks. The whole song does not have to be loaded first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measurements
&lt;/h2&gt;

&lt;p&gt;Across two historical runs using the real model, real music-selection tools, and the production speech service, all 13 boundaries that consumed prefetched talk entered playback in the same logged second as &lt;code&gt;talk.buffer warm&lt;/code&gt;. These included two transitions from music back to talk. The program still kept its default two-second pause between segments.[1]&lt;/p&gt;

&lt;p&gt;Prefetch covered the generation wait at those boundaries. The logs only have one-second resolution, so they do not establish zero latency. There are also too few samples to calculate a meaningful long-run P95.&lt;/p&gt;

&lt;p&gt;There is a separate before-and-after record for the first song:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Cold-start run&lt;/th&gt;
&lt;th&gt;Subsequent run with memory from the previous session&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Before optimization: startup to first song&lt;/td&gt;
&lt;td&gt;136 s&lt;/td&gt;
&lt;td&gt;195 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;After optimization: startup to first song&lt;/td&gt;
&lt;td&gt;71 s&lt;/td&gt;
&lt;td&gt;78 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;After optimization: music preparation itself&lt;/td&gt;
&lt;td&gt;40.2 s&lt;/td&gt;
&lt;td&gt;54.7 s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The experiments used separate data directories, preset personas, cached background beds, and a fixed "listener present" signal. The changes included starting music selection earlier, simplifying search, and limiting the selection context. The improvement in the table reflects those changes together. These numbers come from small-sample engineering records in the repository; I did not rerun the measurements for this article.&lt;/p&gt;

&lt;p&gt;After optimization, the song was ready roughly 30 seconds before it played in both runs. It still had to wait for the third segment boundary. The scheduling rule at the time required two talk segments first, so making search faster would not have brought the song forward in those runs.&lt;/p&gt;

&lt;p&gt;Cold-start waiting remains. Historical measurements put the first audible voice at roughly 29 to 39 seconds. Prefetch did not cover that first batch.&lt;/p&gt;

&lt;p&gt;Music discovery also varies considerably. In another real log, five selections took roughly 82 to 192 seconds each. Talk can continue during selection, but the listener may wait a long time for the next song.&lt;/p&gt;

&lt;p&gt;Natural transitions still need listening tests. Even with a buffer hit, ready audio, and correctly scheduled gain changes, a sentence can fail to follow the previous content. An interruption can land at an uncomfortable moment.&lt;/p&gt;

&lt;p&gt;We use model-free tests for scheduling and invalidation rules, and offline audio rendering for gain changes and handoffs. Latency, source failures, and continuity need runs against real services, followed by actually listening to the result.&lt;/p&gt;

&lt;h2&gt;
  
  
  Implementation and measurement sources
&lt;/h2&gt;

&lt;p&gt;The implementation described here was checked against &lt;code&gt;d6c3619&lt;/code&gt;. Performance figures come from existing engineering records and were not remeasured for this article.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;a href="https://github.com/wine-fall/murmur/blob/d6c361948f88158b04004de35e3b6f718efe80df/docs/blog/2026-09-no-dead-air.md" rel="noopener noreferrer"&gt;Historical measurements, run conditions, and reproduction entry points&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/wine-fall/murmur/blob/d6c361948f88158b04004de35e3b6f718efe80df/src/brain/brain.ts" rel="noopener noreferrer"&gt;Model isolation and task tools&lt;/a&gt;, &lt;a href="https://github.com/wine-fall/murmur/blob/d6c361948f88158b04004de35e3b6f718efe80df/src/brain/steer-responder.ts" rel="noopener noreferrer"&gt;interruption tasks&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/wine-fall/murmur/blob/d6c361948f88158b04004de35e3b6f718efe80df/src/director/director.ts" rel="noopener noreferrer"&gt;Scheduling, prefetch, version invalidation, and audio-start confirmation&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/wine-fall/murmur/blob/d6c361948f88158b04004de35e3b6f718efe80df/src/memory/memory.ts" rel="noopener noreferrer"&gt;Local memory storage&lt;/a&gt;, &lt;a href="https://github.com/wine-fall/murmur/blob/d6c361948f88158b04004de35e3b6f718efe80df/src/memory/compaction.ts" rel="noopener noreferrer"&gt;background profile consolidation&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/wine-fall/murmur/blob/d6c361948f88158b04004de35e3b6f718efe80df/src/audio/engine.ts" rel="noopener noreferrer"&gt;Audio graph and mixing parameters&lt;/a&gt;, &lt;a href="https://github.com/wine-fall/murmur/blob/d6c361948f88158b04004de35e3b6f718efe80df/src/voice/hosted-voice.ts" rel="noopener noreferrer"&gt;complete-clip speech synthesis&lt;/a&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Project repository: &lt;a href="https://github.com/wine-fall/murmur" rel="noopener noreferrer"&gt;wine-fall/murmur&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>typescript</category>
      <category>ai</category>
      <category>architecture</category>
    </item>
    <item>
      <title>I built a radio station with one listener, and it runs in a terminal</title>
      <dc:creator>zhiyi guo</dc:creator>
      <pubDate>Sun, 13 Sep 2026 03:23:31 +0000</pubDate>
      <link>https://dev.to/fawinell/i-built-a-radio-station-with-one-listener-and-it-runs-in-a-terminal-4g7h</link>
      <guid>https://dev.to/fawinell/i-built-a-radio-station-with-one-listener-and-it-runs-in-a-terminal-4g7h</guid>
      <description>&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/DQ4YbSJF2XM" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;The film above is 86 seconds, and the middle of it is a real on-air moment. There is a longer unedited recording at the top of the &lt;a href="https://github.com/wine-fall/murmur" rel="noopener noreferrer"&gt;repo README&lt;/a&gt;. Watch either first, with the sound on. The rest of this post is why it exists and how it is put together.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it is
&lt;/h2&gt;

&lt;p&gt;murmur runs in a terminal window and behaves like a radio station whose only listener is you. It picks a topic on its own and talks about it in a voice that sounds like a person. Then it plays a song, comes back, and keeps going. At the right hours it says good morning and good night. You do not have to answer any of it.&lt;/p&gt;

&lt;p&gt;When you do type something, the host replies in character. You chat for a bit. Then it eases back into the program.&lt;/p&gt;

&lt;p&gt;There is one host. A few questions on the first run write its character to a file. After that, nothing rewrites that file except you. What grows over time is a separate thing: what it knows about you.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why I built it
&lt;/h2&gt;

&lt;p&gt;Every AI tool on my machine wants me to go faster. That is fine for work, and it is exhausting as company. I wanted something that sits closer to the person than to the task. A late-night radio host does that. They talk whether or not you are listening, they never ask you to respond, and when you do call in, they are glad and then they get back to the show.&lt;/p&gt;

&lt;p&gt;Nothing I found did the proactive half. Chat assistants wait for a prompt. Voice tools want to drive my editor. So the bet was simple: what if the AI speaks first and needs nothing back?&lt;/p&gt;

&lt;h2&gt;
  
  
  How it is built
&lt;/h2&gt;

&lt;p&gt;It is TypeScript on Node 24, no build step, run straight from source.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The brain is Claude&lt;/strong&gt;, opened through the Claude Agent SDK. It reuses your local Claude Code login, so there is no API key and no separate bill. It is a harnessed agent rather than a one-shot call. murmur gives it a small set of its own tools: search for a song, judge the candidates, commit a pick, update memory, run the setup guide. It is sealed off from your own Claude Code setup. None of your CLAUDE.md files, skills, MCP servers, or hooks reach it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The voice is fish-speech&lt;/strong&gt;, reached over a hosted HTTP endpoint you point it at. That is the reason it sounds like a person and not a screen reader. It is also the reason murmur is not a local tool, and I want to say that plainly: the brain and the voice are network services. What stays on your machine is the program logic, the keyboard, the mixer, the persona, and the memory.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Music comes through yt-dlp.&lt;/strong&gt; The rules it picks by live in a markdown file under &lt;code&gt;~/.murmur&lt;/code&gt;. The pick task re-reads that file before every song, so if you write "more Cantonese, no covers" while a track is playing, the change reaches the next pick or the one after.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The mixer is a Web Audio graph&lt;/strong&gt; on &lt;code&gt;node-web-audio-api&lt;/code&gt;. Voice and music share one output stream. A gain envelope ducks the song under the host, so the lead-in is spoken over the head of the track and the back-announce over its tail. The song never stops for the voice.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Two processes.&lt;/strong&gt; The engine owns the program and the audio. The TUI is a separate process on OpenTUI under Bun, attached over a unix socket carrying newline-delimited JSON. Without Bun it falls back to a plain text host in the same process, and every command still works typed in full.&lt;/p&gt;

&lt;h2&gt;
  
  
  The decision people ask about
&lt;/h2&gt;

&lt;p&gt;The host's character does not evolve. I built the evolving version first. A background loop watched the conversation and rewrote the persona to fit the listener better. Within days it had drifted into the same generic assistant voice every chatbot has. Language-model rewrite loops mean-revert, and there was no checkpoint where a person could catch the drift.&lt;/p&gt;

&lt;p&gt;So the persona is frozen. It is a text file. You can open it and rewrite it whenever you like, and nothing changes it behind your back. The memory of you is a separate tier that does grow, with dates the code owns and citations back to the conversation that produced each fact.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is rough
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;It needs a Claude Code subscription for the real brain. &lt;code&gt;--brain stub&lt;/code&gt; gives you a canned stand-in brain for a look around, and that is all it is.&lt;/li&gt;
&lt;li&gt;The voice is hosted. If you have no endpoint configured, it prints its lines instead of speaking. A local TTS is on the list, not in the code.&lt;/li&gt;
&lt;li&gt;Music picks sometimes stall or run long. There is an open issue with measurements.&lt;/li&gt;
&lt;li&gt;It is developed on macOS. Linux should be fine. Windows is untested.&lt;/li&gt;
&lt;li&gt;Onboarding is still being tuned by ear, and one path after &lt;code&gt;/quit&lt;/code&gt; misbehaves.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;Node 24 or newer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; murmur-radio
murmur
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It launches even with pieces missing and offers to walk you through installing ffmpeg, yt-dlp, and a voice endpoint by talking to you. Say no once and it stops asking.&lt;/p&gt;

&lt;p&gt;Repo: &lt;a href="https://github.com/wine-fall/murmur" rel="noopener noreferrer"&gt;https://github.com/wine-fall/murmur&lt;/a&gt; (MIT). Landing page: &lt;a href="https://wine-fall.github.io/murmur/" rel="noopener noreferrer"&gt;https://wine-fall.github.io/murmur/&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If you leave it on for an afternoon, I would like to know one thing: did it feel like radio, or like a chatbot with a timer?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>typescript</category>
      <category>cli</category>
      <category>showdev</category>
    </item>
  </channel>
</rss>
