DEV Community

Ivan Mikheev
Ivan Mikheev

Posted on AI-assisted

I built a music-looping tool for TTRPGs and game prototypes with Python and Web Audio

The boss finishes their speech. Your players reach for their dice. This is where the music should kick in…

…and nothing. Instead, the track enters a quiet interlude.

Or the opposite happens: the big musical climax arrives while the party is still arguing about whether to open the door.

I've been running tabletop RPGs for five years, and this mismatch kept bothering me. My music library had the right atmosphere, but the tracks were never long enough to follow the pace of the game. A session follows a certain order — arrival, exploration, tension, combat, aftermath — but every stage has a different duration each time. The same rule of musical development applies to almost every video game.

A fixed track timeline compared with a game session whose stages have varying durations

The track has one timeline. The session has the same order of stages, but each one lasts as long as it lasts.

So I built musslop: a free, open-source tool that lets you loop sections of an existing track, manually cue transitions, and export loops for game audio workflows.

Musslop demo: select a section, loop it, and cue the next section with the Next button

A 30-second silent demo: select → loop → cue → transition. All transitions are manually cued.

You decide when the music moves on. It doesn't listen to your session or automatically react to gameplay.

The stack is Python, FastAPI, librosa, React, and the Web Audio API. Here's what I learned while building it.

The problem: a good track isn't necessarily a good loop

Game music often uses two complementary techniques:

  • Horizontal re-sequencing: move between musical sections, usually at a beat, bar, or phrase boundary.
  • Vertical layering: keep the same musical passage playing while adding or removing instrumental layers.

With purpose-built interactive music, composers can prepare compatible sections, stems, and transition points.

An ordinary stereo recording gives you none of those guarantees.

To make one more controllable, I needed to solve three separate problems:

  1. Find useful musical sections.
  2. Find boundaries that work when those sections repeat.
  3. Schedule playback and transitions accurately.

Those sound like variations of the same problem. They aren't.

1. Finding structure without a neural model

My first analysis pipeline used classic music information retrieval techniques through librosa.

The simplified version looks like this:

Analysis pipeline: beat tracking, beat-synchronous features, self-similarity matrix, novelty curve, section boundaries

Heuristic pipeline: sections show up as blocks in the self-similarity matrix; novelty peaks become candidate boundaries.

For each beat, the analyzer extracts information about:

  • Harmony, using chroma features.
  • Timbre, using MFCCs.
  • Energy, using loudness-related features.

A self-similarity matrix compares positions in the track with one another. Repeated or internally consistent passages appear as blocks.

A checkerboard-shaped kernel along the matrix diagonal produces a novelty curve: peaks suggest places where the musical material changes. This approach goes back to Jonathan Foote's work on audio segmentation.

The next step is to move candidate boundaries onto a musical grid.

That is already an approximation. The heuristic downbeat estimation assumes 4/4, and real recordings can have pickups, tempo changes, or ambiguous accents.

Still, it provides a useful baseline without requiring a large model download.

2. A section boundary and a loop boundary are different things

This was the most useful lesson in the project:

Detecting where the music changes doesn't tell you where it will repeat cleanly.

A section boundary from the novelty curve vs a loop boundary refined for loop closure

The novelty peak says where the chorus starts. The loop-closure refinement moves the cut to a nearby downbeat where the section repeats cleanly.

A boundary might correctly identify the start of a chorus but produce an awkward jump when the preceding section loops.

I added a refinement step that tries nearby downbeat positions and balances several signals:

  • Loop closure: how compatible the end and beginning of the section are.
  • Phrase length: a preference for common phrase lengths, such as four or eight bars.
  • Transition evidence: whether there is an onset or energy change at the boundary.
  • Structural novelty: whether the position is still close to the detected section change.

That last signal matters.

Without a structural anchor, an optimizer can find a locally neat loop that no longer corresponds to the musical section you intended to use.

A better local score can produce a worse arrangement.

The approach was partly inspired by Paul Lamere's Infinite Jukebox, which explores extending music through jumps between compatible beats.

Some passages work better as one-shots

Build-ups are another interesting case.

If a passage keeps increasing in energy, a repeat throws the listener off a cliff: tension, more tension, even more tension — and suddenly back to the start.

A normal section repeats naturally; a build-up climbs in energy so every repeat drops off a cliff

A steady section can repeat. A build-up keeps climbing, so every repeat is a cliff — it should play once.

Occasionally that's a usable effect. Usually it just sounds broken.

musslop uses energy trends and spectral brightness as signals for identifying potential build-ups. Those sections can play once instead of repeating.

It's a suggestion, not a musical law. The user can change the loop behaviour.

3. Scheduling playback in the browser

The frontend uses React for the interface and Web Audio for playback.

The important distinction is between UI timing and audio timing.

A JavaScript timer can help schedule upcoming work, but it shouldn't be the clock that determines the exact moment of an audible transition. Playback needs to be scheduled against the audio context's clock.

Each loop pass uses a new AudioBufferSourceNode with a scheduled start time. When the user presses Next, the player queues the following section at the selected musical boundary.

Next is a cue, not an immediate seek.

The core of the scheduler is small. Every chunk is a fresh source node with its own gain envelope, started at an absolute time on the audio clock:

scheduleChunk(from, to, when, fadeIn, fadeOut) {
  const src = ctx.createBufferSource();
  src.buffer = this.buffer;
  const g = ctx.createGain();
  g.gain.setValueAtTime(0, when);
  g.gain.linearRampToValueAtTime(1, when + fadeIn);
  g.gain.setValueAtTime(1, when + (to - from) - fadeOut);
  g.gain.linearRampToValueAtTime(0, when + (to - from));
  src.connect(g).connect(this.master);
  src.start(when, from, to - from);       // sample-accurate
}
Enter fullscreen mode Exit fullscreen mode

A setInterval tick runs every 60 ms and only makes sure the next chunk is scheduled ~350 ms ahead. When a cue is pending, that next chunk simply comes from the target section instead of the current one — so the transition lands exactly on the boundary without any timer jitter.

Web Audio timeline: each loop pass is its own source node; a cue schedules the next section at a boundary; natural, soon and now cue modes

Each loop pass is a separately scheduled source node. A cue is scheduled for a boundary: natural (loop end), soon (phrase) or now (crossfade).

This is useful during a session: I can request a change and let it land at a suitable point instead of abruptly cutting the audio.

Accurate timing still needs good transitions

Scheduling alone doesn't make arbitrary cuts sound natural.

The player combines several techniques:

  • Short fades to reduce clicks at cuts.
  • Equal-power crossfades to blend outgoing and incoming material.
  • Outgoing tails so the previous section doesn't stop abruptly.
  • Bass-swap handling to reduce overlapping low-frequency content.
  • Stingers such as a hit, cymbal, or riser around the transition.

For an equal-power crossfade, the outgoing and incoming gains follow cos(t·π/2) and sin(t·π/2) for t from 0 to 1:

Linear crossfade vs equal-power cosine/sine crossfade curves

Linear fades dip in power in the middle; cos/sin curves keep the mix at constant power. Loop repeats are correlated, so they get a linear fade instead.

Their squared gains sum to one. For uncorrelated signals, that helps avoid the power dip of a simple linear crossfade.

It doesn't guarantee constant perceived loudness for every pair of musical passages. Correlation, arrangement, and frequency content still matter.

There's also a small but important distinction with stingers: a hit should generally start on the transition, while a riser may need to end there.

4. Adding neural analysis didn't eliminate the editor

The heuristic pipeline was useful, but complex arrangements exposed its limits.

I added two optional structure-analysis engines:

These can identify section boundaries and provide labels such as intro, verse, and chorus.

On my small, manually annotated set of tracks, the newer model wasn't uniformly better. SongFormer worked well on some material but struggled with an orchestral example that All-In-One handled better.

That's a practical observation from a limited personal evaluation, not a general benchmark.

The takeaway was to keep multiple analysis options and make the output editable.

You can:

  • Drag boundaries with bar snapping.
  • Split and merge sections.
  • Adjust the loop repeat start.
  • Change whether a section loops.
  • Undo edits.

The analyzer proposes an arrangement. The user gets the final say.

5. Experimenting with intensity layers

For vertical layering, musslop can run Demucs to split a track into drums, bass, vocals and everything else, and lets you mix those layers live. The same recording can then play as a sparse "exploration" arrangement or a full "combat" one.

There are limits: separated stems carry artifacts, and muting the drums doesn't magically turn a battle track into ambience. But it's a cheap way to test how an existing recording might behave as interactive music before anyone writes stems for real.

Where it fits

At a tabletop session

Prepare the track before the game, check the loops, and correct any boundaries that need attention.

During play:

  1. Select a section.
  2. Keep it looping while the scene unfolds.
  3. Press Next when you want to move on.

The purpose is to reduce the time spent searching and scrubbing through music while also running the session.

During game development

Use it to audition loops, section transitions, and intensity changes before implementing the playback behaviour in your game.

You can export sections as WAV loops for your audio workflow.

It's currently a standalone tool, not an FMOD/Wwise replacement or a drop-in game-engine integration. The in-game logic is still yours to implement.

Try it locally

The project requires Python 3.10+ and Git for the commands below.

Linux / macOS:

git clone https://github.com/Siziff/musslop.git
cd musslop
./setup.sh
./run.sh
Enter fullscreen mode Exit fullscreen mode

Windows:

git clone https://github.com/Siziff/musslop.git
cd musslop
setup.bat
run.bat
Enter fullscreen mode Exit fullscreen mode

Then open http://localhost:8801.

The base setup includes heuristic analysis, editing, playback, and loop export. Neural analysis and stem separation use optional components; installation details are in the README.

Once the required dependencies and model weights have been downloaded, processing works locally and offline.

Free, open source, and personal

musslop is free and open source under the MIT license.

No subscriptions, accounts, registration, email collection, or SMS verification. There's no paid tier or planned subscription.

This is a personal project I'd wanted to build for a long time. I finally made it for myself, and I'm sharing it because other people might find it useful too.

For transparency, I used AI assistance for the UI. Separately, the optional audio-processing features use pretrained models for analysis and stem separation. The tool works with existing music rather than generating tracks.

What I'd like feedback on

I'm especially interested in hearing from developers working with audio and people running tabletop sessions:

  • How do you balance responsive transitions against musical phrasing?
  • What makes a long-running loop feel less repetitive?
  • For a game prototype, would WAV export be enough, or would you also need section and transition metadata?
  • What would stop you from using a tool like this?

Repository: github.com/Siziff/musslop

If it looks useful for your table or game, a star helps other people discover it. Reports about awkward loops or confusing controls are just as welcome.


Adapted from my original Russian-language article on Habr.

References

Top comments (1)

Some comments have been hidden by the post's author - find out more