DEV Community

Anup Jayant Dharangutti
Anup Jayant Dharangutti

Posted on

An exploration of media as code: deterministic audio, synchronized visuals, runtime-bound media, transcription, and developer tooling in .NET.

Media as Code: Building Deterministic Audio, Visuals, and Adaptive Experiences with SoundScript

What changes when audio, visuals, timing, application state, and even parts of media analysis become things developers can compile, inspect, test, diff, and version?

Most software teams are comfortable treating configuration as code.

Infrastructure became code.

Build pipelines became code.

UI increasingly became declarative.

But media still often enters a project as a collection of opaque artifacts:

success-final.wav
success-final-v2.wav
success-final-v2-fixed.wav
notification-new.wav
animation-final.webm
animation-final-really-final.webm
Enter fullscreen mode Exit fullscreen mode

That workflow is perfectly reasonable when the media itself is the final product.

It becomes less comfortable when media is actually part of software behavior.

A monitoring system needs a warning cue.

A test suite needs a repeatable audio fixture.

An application needs synchronized audio and visual status feedback.

A developer tool needs generated notification sounds.

A data-driven experience needs its media to respond to state.

A transcription experiment needs to turn suitable audio back into something editable.

In those situations, the interesting question is not:

“How do I store this media file?”

It becomes:

“How much of this media can I describe as source?”

That is the direction behind SoundScript: an open-source .NET project for programmable audio and synchronized media.

Rather than treating every sound, score, visual cue, or media state as an opaque binary, SoundScript lets developers work with readable descriptions, compiled representations, deterministic renderers, queryable timelines, and application-facing APIs.

This article explores that idea from the perspective of a developer building something—not from the perspective of release notes.


Start with something deliberately small

Imagine an application needs a short success cue.

You could ship a WAV file.

Or you could describe the musical intent:

tempo 120

track success {
    instrument piano
    mf

    C4 e
    E4 e
    G4 q
}
Enter fullscreen mode Exit fullscreen mode

Then compile it from .NET:

using SoundScript;

var compilation = SoundScriptEngine.Compile("""
    tempo 120

    track success {
        instrument piano
        mf

        C4 e
        E4 e
        G4 q
    }
    """);

File.WriteAllBytes("success.wav", compilation.RenderWave());
File.WriteAllBytes("success.mid", compilation.RenderMidi());
Enter fullscreen mode Exit fullscreen mode

The result is still media.

The difference is that now the reason the media sounds the way it does is visible in source control.

If somebody changes the last note from G4 to C5, that change can appear in a diff.

If the tempo changes from 120 to 108, you can review the exact intention.

If a regression test expects a known deterministic rendering, the source that generated it lives beside the test.

Why this helps

Media stops being only an artifact and becomes partly an explainable software input. That makes review, reuse, debugging, testing, and maintenance much closer to ordinary engineering workflows.

This does not mean every piece of audio should become code.

It means code becomes another useful representation when the media participates in application logic.


From notes to a small musical language

A three-note cue is useful, but real authoring needs structure.

SoundScript's music language includes pitches, rests, ties, standard and numeric durations, dotted notes, tuplets, grace notes, chords, dynamics, articulations, time signatures, tempo changes, tracks, reusable blocks, patterns, phrases, loops, imports, layers, and orchestration helpers.

A more structured score might look like this:

perform expressive

tempo 96
time 4/4

pattern arp {
    up
}

block resolution {
    C4 q
    E4 q
    G4 q
    C5 q
}

track lead {
    instrument flute
    mf

    phrase {
        articulation legato
        crescendo

        C4 q
        D4 q
        E4 h
    }

    play resolution
}
Enter fullscreen mode Exit fullscreen mode

This matters because musical intent is not just a list of MIDI numbers.

There is a difference between:

C4 D4 E4
Enter fullscreen mode Exit fullscreen mode

and:

phrase {
    articulation legato
    crescendo
    C4 q
    D4 q
    E4 h
}
Enter fullscreen mode Exit fullscreen mode

The second version says something about how the phrase should behave.

SoundScript keeps a distinction between the written structure and the renderer interpreting it.

That becomes important when the same source can travel toward different outputs.

                   SoundScript source
                          │
                          ▼
                 Parser / interpretation
                          │
                ┌─────────┴─────────┐
                ▼                   ▼
              MIDI                Wave
          event-oriented      rendered audio
Enter fullscreen mode Exit fullscreen mode

The two outputs share musical intent but are not identical representations.

MIDI carries events.

Wave carries samples.

Some expressive information naturally behaves differently across those worlds.

Why this helps

A developer can keep intent above output format. The score does not have to be designed exclusively around a single file type from the beginning.

The full language reference is available in the SoundScript documentation.


Determinism is more useful than it sounds

“Deterministic media” can sound like an academic concern until you put media inside tests, automation, CI, build systems, or generated application assets.

Suppose a test creates a notification sound.

If every run changes slightly because humanization uses uncontrolled randomness, byte-level regression testing quickly becomes unpleasant.

SoundScript instead uses controlled deterministic behavior for supported rendering paths.

Humanization can still exist:

track piano {
    humanize timing=0.02 velocity=0.08 seed=42

    C4 q
    E4 q
    G4 q
}
Enter fullscreen mode Exit fullscreen mode

The important difference is that variation can be repeatable.

When a seed is not explicitly provided in supported Wave features, SoundScript can derive deterministic variation from content rather than reaching for clock-based randomness.

That makes “humanized” and “reproducible” less contradictory than they first appear.

Why this helps

Deterministic media is especially valuable for test fixtures, CI comparisons, generated assets, reproducible demos, regression investigation, and build pipelines. When output changes, you can ask whether the source or renderer actually changed instead of wondering whether randomness changed it.

There is an important boundary here.

Repeatability depends on the relevant fixed inputs: source, assets, renderer version, configuration, and supported deterministic execution path.

Compressed-media decoding and codec-based video encoding have additional environmental variables.

The goal is not to claim that every multimedia byte will magically be identical everywhere.

The goal is to make determinism a deliberate engineering property where it is defensible.


Direct audio: not everything needs to become MIDI first

A programmable music language does not necessarily have to route every operation through MIDI.

SoundScript also has a direct Wave backend.

That path can render a script directly to WAV and adds capabilities that make more sense after audio samples actually exist.

For example:

effect delay time=0.25 feedback=0.35 mix=0.25
effect filter type=lowpass cutoff=2400

track cue {
    humanize timing=0.015 velocity=0.05 seed=12

    C4 q
    E4 q
    G4 h
}
Enter fullscreen mode Exit fullscreen mode

The Wave path supports areas such as:

Wave rendering
├── direct script → WAV
├── mono / stereo output
├── expressive performance
├── deterministic humanization
├── delay
├── low-pass filtering
├── high-pass filtering
├── samples / recorded WAV stems
├── synthetic speak/prosody cues
├── external vocal stems
└── unpitched percussion events
Enter fullscreen mode Exit fullscreen mode

Why keep this separate from MIDI?

Because MIDI has no final mixed audio buffer on which to apply a post-mix delay or filter.

That separation avoids pretending every media concept belongs in every backend.

Why this helps

Backend-specific features can remain explicit instead of silently degrading. When an operation fundamentally requires audio samples, it belongs on an audio path rather than being awkwardly disguised as a MIDI feature.

See the Wave grammar documentation.


What if timbre were styled?

One of the more unusual pieces of the project is SoundCSS.

The rough analogy is intentional:

HTML describes structure.

CSS describes presentation.

SoundScript describes musical/media structure.

SoundCSS describes aspects of sound character.

A stylesheet can define phoneme-oriented properties such as harmonics, formants, transient behavior, brightness, noise components, resonance, and smoothing.

Conceptually:

aa {
    formant1: 700Hz;
    formant2: 1100Hz;

    harmonic1: 0.9;
    harmonic2: 0.6;
    harmonic3: 0.3;

    smoothness: 0.9;
}
Enter fullscreen mode Exit fullscreen mode

The project also supports word-level rules for things such as delivery style, pitch shift, speed, energy, timbre, persona, emotion, breath, vibrato, and other bounded transforms.

For example:

"initialize" {
    persona: robot;
    pitch: -2;
    vibrato: none;
    timbre: flat;
}
Enter fullscreen mode Exit fullscreen mode

This is not an attempt to replace modern neural speech synthesis.

It is a declarative way to experiment with reproducible transformations in the offline rendering pipeline.

Why this helps

Style-like media configuration separates what should be said or played from how a renderer should color it. That makes it easier to reuse content while experimenting with presentation.

The detailed model is documented in SoundCSS.


Text can become musical structure

There are several very different things people might mean by “text to sound.”

One is speech synthesis.

Another is turning language into musical structure.

SoundScript includes deterministic text-to-melody paths.

For example, the compose workflow uses syllables and phonemes as input to musical gesture generation.

A separate prosody path uses word stress and sentence contour to influence pitch.

Conceptually:

Text
 │
 ├── syllables
 │
 ├── phonemes
 │
 └── prosodic structure
       │
       ▼
 musical gestures
       │
       ▼
 SoundScript score / MIDI
Enter fullscreen mode Exit fullscreen mode

That means text can become editable musical source rather than terminating as an opaque generated asset.

This distinction matters.

Text-to-melody is not text-to-speech.

A phoneme composer generates notes.

A vocal pipeline can associate written lyrics with written pitches.

A Wave speak path can generate synthetic phoneme/prosody audio cues.

A WordBank-based path can work with recorded or generated word stems.

Those are separate tools because they solve separate problems.

Why this helps

Keeping text-to-melody, lyric alignment, synthetic cues, and vocal rendering separate avoids one oversized “AI voice” abstraction. Developers can choose the representation they actually need and inspect the intermediate result.


Lyrics can live next to pitches

Another part of the project treats vocals as symbolic data.

Consider:

voice lead {
    vocal choir
    mf

    sing "Twinkle twinkle little star"
         C4 q C4 q G4 q G4 q A4 q A4 q G4 h
}
Enter fullscreen mode Exit fullscreen mode

The engine aligns syllables with notes and emits lyric metadata in MIDI.

That makes the file useful to tools that understand karaoke-style lyric events.

The lyrics themselves are not audio embedded inside MIDI.

They are structured metadata associated with musical events.

The Playground can provide a browser speech preview, while deterministic MIDI remains separate from browser-dependent speech synthesis.

Why this helps

Symbolic lyric timing is reusable. The same musical/lyrical structure can travel into a DAW, karaoke workflow, another synthesizer, a visual lyric display, or a later rendering stage without forcing one speech engine into the core model.

See the vocal documentation.


Audio is only one rail

At some point the project crossed an important boundary.

Instead of asking only:

“What sound should happen?”

the language can also ask:

“What should exist visually at this moment?”

A simplified media program can combine a musical cue and a visual element:

tempo 120

track cue {
    C4 q
    E4 q
    G4 h
}

sync audio

visual "indicator" for 4s {
    shape circle
    fill "#16a34a"

    animate x 200 -> 1080 over 4s
}
Enter fullscreen mode Exit fullscreen mode

The important idea is not merely that a video can eventually be rendered.

It is that the visual timeline itself is queryable.

From .NET:

using SoundScript;
using SoundScript.Media;

var compilation = SoundScriptEngine.Compile(source);

var media = compilation.CompileMedia();

var scene = media.SceneAt(TimeSpan.FromSeconds(2));

string json = TemporalVisualJson.Serialize(scene);
string svg = TemporalSvgRenderer.Render(scene);
byte[] wav = media.RenderAudio();
Enter fullscreen mode Exit fullscreen mode

That makes the core model:

                    compiled media
                          │
             ┌────────────┴────────────┐
             │                         │
             ▼                         ▼
           Audio                    Timeline
          WAV/PCM                  SceneAt(t)
             │                         │
             └──────── shared t ───────┘
                          │
                          ▼
                    scene model
                ┌─────────┼─────────┐
                ▼         ▼         ▼
              JSON       SVG      custom UI
Enter fullscreen mode Exit fullscreen mode

This is different from treating video as the primary abstraction.

The application can ask:

What exists at t = 2.0 seconds?
Enter fullscreen mode Exit fullscreen mode

without first generating a video and decoding frame 60.

Why this helps

SceneAt(t) turns media timing into application data. That is useful for synchronized dashboards, simulations, accessibility layers, tests, alternate renderers, custom UIs, and systems where the application—not the media file—owns playback.

See the programmable media runtime.


The timeline is not tied to one renderer

Once a scene can be queried as structured state, the host has options.

The same conceptual media state can feed:

SceneAt(t)
├── JSON
├── SVG
├── HTML
├── Canvas
├── application-native UI
└── sampled frames → WebM
Enter fullscreen mode Exit fullscreen mode

That architectural separation is subtle but useful.

A browser does not need to implement a second interpretation of the SoundScript timing model.

A desktop application does not need to parse the source again.

A video exporter does not need to invent its own semantics.

They all consume evaluated state.

Why this helps

One timing model can support multiple presentation technologies. That reduces the chance of the browser, CLI, server, and video renderer slowly developing different interpretations of the same source.


Then comes the interesting part: application state

Static media is useful.

But applications are not static.

A monitoring system might have:

Normal
Warning
Critical
Enter fullscreen mode Exit fullscreen mode

A game entity might have changing intensity.

A simulation may expose continuously changing numeric state.

Rebuilding source strings every time a value changes would technically work, but it is not a particularly clean host API.

SoundScript therefore supports a constrained runtime-parameter model.

Consider:

param intensity = 0.25
param xpos = 200

perform expressive
tempo 120

track cue {
    gain intensity

    C4 q
    E4 q
    G4 h
}

visual "indicator" for 4s {
    shape circle

    set x xpos
    set y 360
    set width 120
    set height 120
    set opacity intensity
}
Enter fullscreen mode Exit fullscreen mode

The source describes a fixed structure.

The application changes approved numeric values.

using SoundScript;

var runtime = SoundScriptEngine.CompileRuntime(source);

runtime.SetMany(new Dictionary<string, decimal>
{
    ["intensity"] = 0.90m,
    ["xpos"] = 900m
});

var snapshot = runtime.Bind();

File.WriteAllBytes("critical.wav", snapshot.RenderAudio());

var scene = snapshot.SceneAt(
    TimeSpan.FromSeconds(2));
Enter fullscreen mode Exit fullscreen mode

The source was not regenerated.

The structure was not reparsed for the update.

The host supplied state.

The runtime validated it.

The snapshot froze a coherent view.

That last part is important.

If the host updates the runtime again, a previously bound snapshot does not silently mutate.

Compiled structure
       │
       ├── Runtime instance A
       │      ├── state
       │      └── Bind() → Snapshot A
       │
       └── Runtime instance B
              ├── state
              └── Bind() → Snapshot B
Enter fullscreen mode Exit fullscreen mode

Why this helps

Snapshots provide a clean boundary between mutable application state and stable media output. A render can finish against the state it was given instead of observing half of one update and half of another.

The runtime is deliberately constrained.

Parameters do not rewrite arbitrary notes, imports, tempo, timing, names, or structural decisions.

Supported use is intentionally more like data binding than self-modifying source code.

That limitation is valuable.

Why this helps

Restricting runtime parameters keeps the compile-time structure understandable. The host gets adaptability without turning every media operation into dynamic code generation.

The complete model is described in the runtime parameters guide.


A monitoring example makes the model concrete

Imagine a build or infrastructure dashboard.

Instead of playing the same alert for everything, application state can become both visual and audible.

Healthy
  ↓
low intensity
green indicator
quiet cue

Warning
  ↓
medium intensity
larger/brighter indicator
stronger cue

Critical
  ↓
high intensity
prominent indicator
louder cue
Enter fullscreen mode Exit fullscreen mode

That same state could drive:

                 Application state
                        │
                        ▼
               validated parameters
                        │
                        ▼
                  bound snapshot
                  ┌─────┴─────┐
                  ▼           ▼
                audio       visual
                  │           │
                  └─────┬─────┘
                        ▼
                 shared experience
Enter fullscreen mode Exit fullscreen mode

The interesting point is not the warning sound itself.

It is that one piece of typed state can influence multiple media channels coherently.

Why this helps

In status-oriented applications, audio and visuals can reinforce each other from the same state instead of being maintained as unrelated assets and UI logic.


Media can become test data

This is one of the less glamorous use cases, but perhaps one of the most practical.

Tests often need inputs.

A media-processing system might require:

a 440 Hz cue
a known sequence of notes
a fixed-duration WAV
a deterministic alert pattern
a synchronized visual timeline
a repeatable transcription input
Enter fullscreen mode Exit fullscreen mode

Traditionally you check binary fixture files into the repository.

Sometimes that is exactly the correct solution.

But generated fixtures have a useful property:

the fixture can explain itself.

A source such as:

tempo 120

track fixture {
    C4 q
    E4 q
    G4 h
}
Enter fullscreen mode Exit fullscreen mode

is much easier to inspect than a binary WAV when the test expectation is “three known notes at known durations.”

A build step can render the media when necessary.

Why this helps

Programmatic fixtures reduce the gap between what a test says it needs and what the binary fixture actually contains. That can make failures easier to investigate and fixtures easier to evolve intentionally.

SoundScript includes examples oriented toward deterministic fixtures and DevOps-style sonification in its documentation and sample projects.


CI can listen too

Once audio is programmable, non-media systems can use it.

One small example is sonification: mapping application or pipeline states to sound.

A CI system could represent:

Build started     → short neutral cue
Tests passed      → ascending confirmation
Tests failed      → contrasting alert
Deployment ready  → completion phrase
Enter fullscreen mode Exit fullscreen mode

This is not intended to replace logs.

It creates another channel.

The interesting part is that the mapping can live in source and be tested like any other behavior.

Why this helps

Sonification can provide ambient feedback when developers are already overloaded visually. It is especially useful as a complement to logs and dashboards rather than a replacement for them.


The reverse direction: audio back into editable source

So far, the flow has mostly been:

source → media
Enter fullscreen mode Exit fullscreen mode

But sometimes developers begin with media.

SoundScript also contains a transcription subsystem.

For suitable monophonic input, the goal is:

audio
  │
  ▼
analysis
  │
  ▼
musical observations
  │
  ▼
editable SoundScript
  │
  ▼
render / compare / modify
Enter fullscreen mode Exit fullscreen mode

A .NET example looks like this:

using SoundScript.Transcription;

var audio = PcmWaveInput.Decode(
    File.ReadAllBytes("melody.wav"));

var result =
    await new TranscriptionEngine().TranscribeAsync(
        audio,
        new TranscriptionOptions(
            Tempo: 120,
            Instrument: 73),
        TranscriptionMode.Monophonic);

Console.WriteLine(result.Suitability.Status);

var source =
    new SoundScriptOutput().Source(result.Score);

File.WriteAllText("melody.ss", source);
Enter fullscreen mode Exit fullscreen mode

The emphasis on suitability matters.

Audio transcription is not deterministic truth recovery.

A waveform does not contain a hidden .ss file waiting to be extracted perfectly.

Pitch estimation, segmentation, tempo inference, mixtures, room effects, noise, instrumentation, and performance variation make the problem ambiguous.

SoundScript therefore separates the supported monophonic baseline from experimental analysis modes.

Transcription
├── Monophonic
│   └── supported baseline for suitable single-line material
│
├── Extract Melody
│   └── experimental dominant-line estimation
│
├── Polyphonic / Piano
│   └── experimental simultaneous-note analysis
│
├── Mixed Roles
│   └── experimental melody / harmony / bass estimates
│
└── Percussion / Rhythm
    └── experimental transient / unpitched events
Enter fullscreen mode Exit fullscreen mode

Compressed desktop media can be decoded through FFmpeg-supported paths, while browser decoding depends on available browser codecs.

Analysis is bounded to short clips rather than pretending to be an unlimited music-understanding service.

Why this helps

Treating transcription as evidence plus editable output is much safer than pretending an estimated score is ground truth. Developers can inspect diagnostics, correct the generated source, and keep the human review step where the signal is ambiguous.

See the transcription documentation.


The Playground makes the model easier to understand

A language is much easier to understand when you can modify something and immediately experience the result.

The SoundScript Playground provides browser workflows around several parts of the system:

Playground
├── Music & Wave
├── text-to-melody
├── prosody
├── Wave rendering
├── audio/visual timelines
├── runtime parameter examples
├── transcription
├── source editing
├── diagnostics
├── timeline scrubbing
├── play / pause / resume / restart
└── export workflows
Enter fullscreen mode Exit fullscreen mode

The editor also contains developer-oriented assistance such as validation, formatting, completion, hover information, outlines, find/replace, transposition helpers, duration editing, and visual authoring conveniences.

The point is not to become a full DAW inside a browser.

The Playground is closer to an interactive executable documentation surface for the language and runtime.

Why this helps

Developers can explore the model before committing to an integration. A concept like SceneAt(t) or a runtime parameter is easier to understand when you can scrub time or modify a value and immediately see the state change.

Try it at soundscript.net/playground.


The .NET API keeps the core inside your process

For .NET applications, the package exposes a programmatic facade.

The current public library is available as:

dotnet add package SoundScript --version 16.0.1
Enter fullscreen mode Exit fullscreen mode

The basic workflow is intentionally straightforward:

var compilation =
    SoundScriptEngine.Compile(source);

byte[] midi = compilation.RenderMidi();
byte[] wav = compilation.RenderWave();
Enter fullscreen mode Exit fullscreen mode

For synchronized media:

var media =
    compilation.CompileMedia();

byte[] audio =
    media.RenderAudio();

var scene =
    media.SceneAt(TimeSpan.FromSeconds(1.5));
Enter fullscreen mode Exit fullscreen mode

For adaptive state:

var runtime =
    SoundScriptEngine.CompileRuntime(source);

runtime.Set("intensity", 0.8m);

var snapshot =
    runtime.Bind();
Enter fullscreen mode Exit fullscreen mode

Compilation does not need to launch a CLI subprocess.

That distinction matters when SoundScript becomes part of a server, worker, desktop application, testing utility, or internal tool.

Why this helps

In-process APIs avoid shell parsing and process orchestration for normal application use. Media generation can participate directly in typed application code, exception handling, testing, and dependency management.

See the .NET API guide and NuGet documentation.


The CLI solves a different problem

A library is useful when your program owns the workflow.

A CLI is useful when your shell, script, or CI pipeline owns it.

The project exposes commands for areas such as:

Music
├── run
├── compose
├── prosody
├── render
└── wave

Media
├── visual
└── video

Analysis
└── transcribe

Automation
├── validate
└── inspect

Vocals
├── vocal generate
├── vocal batch
├── wordbank ensure
└── wordbank normalize
Enter fullscreen mode Exit fullscreen mode

Validation and inspection are particularly useful in automation because they can return structured diagnostics and stable exit behavior.

A workflow can ask:

Does this source compile?
Which outputs are supported?
What is its duration?
Which visuals exist?
Are dependencies available?
Can FFmpeg produce the requested video format?
Enter fullscreen mode Exit fullscreen mode

before committing to a full render.

Why this helps

Separating validation from rendering lets CI fail early. You do not need to generate every media artifact just to discover that the source or environment is invalid.

The current public library is published to NuGet; the CLI is documented as a source-build workflow for the public version rather than a separately published NuGet CLI tool.

See the CLI reference.


The capability map

By this point, the project is easier to understand as a set of connected developer capabilities than as a list of releases.

SoundScript
│
├── Musical authoring
│   ├── notes / rests / ties
│   ├── durations / dotted notes / tuplets / grace notes
│   ├── chords / voicings
│   ├── tempo / meter
│   ├── tracks / melodies / sequences
│   ├── loops / blocks / imports
│   ├── patterns
│   ├── phrases
│   ├── dynamics / velocity
│   ├── articulations
│   ├── deterministic humanization
│   ├── GM instruments
│   ├── layers
│   └── orchestration
│
├── Audio
│   ├── MIDI
│   ├── direct Wave rendering
│   ├── stereo rendering
│   ├── expressive performance
│   ├── samples / stems
│   ├── delay / filtering
│   ├── synthetic speak cues
│   ├── unpitched percussion
│   └── SoundCSS timbre
│
├── Text and vocals
│   ├── text → melody
│   ├── word-level prosody
│   ├── editable .ss output
│   ├── lyric-to-note alignment
│   ├── karaoke MIDI metadata
│   ├── vocal stem workflows
│   ├── WordBank
│   ├── pronunciation transforms
│   └── continuous vocal rendering
│
├── Programmable visuals
│   ├── timed visuals
│   ├── overlaps / waits
│   ├── audio synchronization
│   ├── shapes / text
│   ├── fills / strokes / styles
│   ├── animations
│   ├── exact SceneAt(t)
│   ├── typed scene model
│   ├── JSON
│   ├── SVG
│   └── WebM export
│
├── Adaptive runtime
│   ├── compile once
│   ├── typed decimal parameters
│   ├── validation
│   ├── atomic SetMany
│   ├── reset
│   ├── independent instances
│   ├── immutable bound snapshots
│   ├── audio binding
│   └── visual property binding
│
├── Transcription
│   ├── monophonic baseline
│   ├── melody extraction [experimental]
│   ├── polyphonic/piano [experimental]
│   ├── mixed roles [experimental]
│   ├── percussion [experimental]
│   ├── diagnostics
│   ├── suitability
│   ├── editable source
│   └── reconstruction comparison
│
├── Developer tooling
│   ├── .NET 10 library
│   ├── CLI
│   ├── validation
│   ├── inspection
│   ├── JSON diagnostics
│   ├── Playground
│   ├── examples
│   ├── deterministic fixtures
│   └── CI-oriented workflows
│
└── Labs
    ├── VideoLab
    └── Signal Lab
Enter fullscreen mode Exit fullscreen mode

That map is useful because the individual features make more sense once the connecting idea is visible:

Source → validated structure → deterministic interpretation → inspectable state → chosen output.


Labs: experiments should look like experiments

Not every idea belongs in the stable package immediately.

SoundScript keeps experimental work in Labs, outside the main supported product surface.

That separation is important because experimental code can answer architecture questions without silently becoming a compatibility promise.

VideoLab

VideoLab explores programmable composition of actual video/audio assets using an independent experimental model.

Its work includes areas such as:

JSON composition
       │
       ▼
immutable composition
       │
       ▼
runtime bindings
       │
       ▼
snapshot
       │
       ▼
SceneAt(frame)
       │
       ▼
FFmpeg
       │
       ├── MP4
       └── WebM
Enter fullscreen mode Exit fullscreen mode

The experiment covers clips, transforms, crops, opacity, audio gain, expressions, easing, reusable effects, conditions, parameterized variants, data-driven sequences, batches, and review-annotation scenarios.

It is deliberately bounded.

It is not presented as a full nonlinear video editor.

Why this helps

A separate lab allows the project to explore new composition semantics without contaminating the compatibility surface of the production language.

Explore the public VideoLab area at soundscript.net/labs/videolab.

Signal Lab

Signal Lab explores another direction: signals rather than music.

Its small experimental language can generate deterministic signals such as sine, square, chirp, and sweep waveforms and analyze them with operations including FFT, peak detection, RMS, and frequency estimation.

Outputs include structured JSON plus WAV, PCM16, and Float32 data.

Its scope is deliberately narrow and simulated.

It does not claim hardware capture, calibration, machine control, or measurement certification.

Why this helps

Keeping signal experiments isolated makes it possible to test whether SoundScript-like ideas generalize beyond music without prematurely expanding the production product.

See SoundScript Labs.


Where this architecture becomes especially useful

The value of programmable media varies enormously by project.

For a professionally mastered album, source-based synthesis may not be the interesting part.

For application media, however, several situations stand out.

Test automation

A test can generate the media it expects instead of depending entirely on unexplained fixture binaries.

Benefit: fixtures become easier to review, regenerate, and version.

Monitoring and operational software

One state can influence both audible and visual feedback.

Benefit: status communication can stay synchronized across multiple channels.

Developer tools

Build status, diagnostics, workflows, or simulations can emit meaningful cues.

Benefit: audio can supplement crowded visual interfaces.

Games and interactive applications

A compiled structure can create independent stateful instances and bound snapshots.

Benefit: one media definition can support many entities or sessions without source regeneration.

Procedural content

Inputs can produce repeatable outputs rather than requiring a hand-authored asset for every variant.

Benefit: content scales with data.

Education and experimentation

Music, signal, timing, and rendering concepts are visible in source and inspectable at intermediate stages.

Benefit: the path from intent to artifact becomes easier to study.

Accessibility experiments

Queryable media state creates opportunities for alternate representations rather than assuming one fixed rendered output.

Benefit: applications can interpret the same timing/state model differently for different users or interfaces.

CI and reproducibility

Validation, inspection, deterministic rendering, and structured diagnostics fit naturally into automated workflows.

Benefit: media becomes more compatible with ordinary software delivery practices.


What SoundScript deliberately does not try to hide

A programmable-media system becomes less useful if every boundary is marketed away.

Several distinctions matter.

SoundScript is not a DAW replacement.

It is not a neural music generator.

Its synthetic vocal paths are not a replacement for production speech synthesis.

Experimental transcription is not perfect music understanding.

VideoLab is not a production nonlinear editor.

Signal Lab is not measurement hardware.

Runtime parameters are not unrestricted self-modifying code.

WebM export still depends on FFmpeg.

Browser speech and browser media decoding depend on the browser environment.

These boundaries are features of the engineering story, not weaknesses to conceal.

Why this helps

Explicit limitations tell developers where the abstraction ends. That makes integration decisions more reliable than an API that appears universal until edge cases reveal otherwise.


The deeper idea: version the reason, not just the result

Imagine reviewing these two commits.

Commit A

Replace notification.wav
Enter fullscreen mode Exit fullscreen mode

Commit B

- tempo 120
+ tempo 108

track warning {
-   C4 q E4 q G4 h
+   C4 q Eb4 q G4 h
}
Enter fullscreen mode Exit fullscreen mode

The first commit tells you that the media changed.

The second begins to tell you why.

That difference is the part of programmable media I find most interesting.

Developers already expect software behavior to be explainable through code.

Media embedded in applications is also behavior.

Sometimes the right representation is still a handcrafted WAV, PNG, MP4, or professionally produced asset.

But sometimes the better representation is:

intent
+
source
+
parameters
+
assets
+
renderer
Enter fullscreen mode Exit fullscreen mode

with the final media treated as an output.

Why this helps

Versioning the explanation makes media changes easier to review, reproduce, compare, regenerate, and connect to the application behavior that required them.


A practical way to explore SoundScript

If you want to understand the project, I would not begin by trying every feature.

Start with one tiny outcome.

Create a three-note cue.

Change its instrument.

Add a phrase.

Render it to WAV.

Add one visual.

Query the scene at two seconds.

Replace one constant with a runtime parameter.

Bind a new value from C#.

Only then explore text composition, vocals, transcription, or Labs.

The conceptual progression is:

sound
  ↓
structured sound
  ↓
rendered audio
  ↓
timed visuals
  ↓
queryable media
  ↓
application-bound media
Enter fullscreen mode Exit fullscreen mode

At each step the question is the same:

Can the media remain understandable as software?


Final thought

Software engineering has spent decades converting opaque operational processes into inspectable systems.

We did it with builds.

We did it with infrastructure.

We did it with deployments.

We did it with configuration.

Media will never become entirely code—and it should not.

Human performance, filmmaking, sound design, illustration, and production assets have value precisely because they are richer than a declarative model.

But application media occupies an interesting middle ground.

When a sound represents application state...

when a visual follows an application clock...

when a test requires known audio...

when a generated artifact must be reproducible...

when multiple variants come from the same structure...

when the system needs to inspect what exists at an exact time...

then treating at least part of that media as code becomes surprisingly powerful.

The biggest benefit is not that code can generate a WAV or a WebM.

The benefit is that the system can retain an inspectable explanation of how that media came to exist.

That is the idea I am exploring with SoundScript.

If that problem space interests you, start with the Playground, read the source, or try the .NET package.

Playground: soundscript.net/playground

Documentation: soundscript.net documentation

GitHub: github.com/dharangutti/sound-script

NuGet: SoundScript

Labs: soundscript.net/labs

Top comments (0)