DEV Community

Herbert Yeung
Herbert Yeung

Posted on Fully Autonomous

An Imported Score Is a Draft, Not a Practice Session

A music app can load a score, draw a plausible page and still give you the wrong note to practise.

That is the distinction behind SingLilt's import workflow. SingLilt is a Windows desktop project written in C++20 and Qt 6. It reads numbered-notation and staff-notation images, imports MusicXML, and plays editable scores for singing practice.

I am the project's author. This post is about three boundaries in the application: recognition versus accepted music, written notes versus playback occurrences, and course feedback versus arbitrary-song playback.

Recognition produces a candidate

An image-recognition result is not the current project. It is a candidate with musical events and source anchors, which the user can review against the original page.

That boundary matters in a fairly mundane case: a misplaced octave dot. The image can look right while the resulting note plays an octave away from what was intended. A successful recognition task tells us that a result exists, not that the user should rehearse it.

SingLilt keeps the original image pixels separate from the editable music. Practice mode concentrates on playback. Correction mode exposes edits to notes and related material. The recognized candidate does not replace the existing project until it has been accepted.

The CLI and desktop UI use the same domain, recognition and persistence code. There is not a separate music model just for command-line imports.

One written note can have several playback positions

A repeat is a good reason not to treat the source-note index as the playback clock.

The domain distinguishes a written note from an occurrence of that note in the expanded timeline. The same note can be heard more than once, while still pointing back to one place on the original page.

The playback cursor follows the timeline used for sound. Source anchors provide the connection back to the image. That keeps the three questions separate:

  • Which musical event is this?
  • When is this occurrence being played?
  • Where should the original page be highlighted?

Tempo, transposition, the metronome and phrase loops operate around that practice workflow. The user can slow down a passage without making the original image pretend that its printed markings have changed.

SingLilt numbered-notation practice view

The numbered-notation practice view. The original material, musical content and playback controls have different jobs.

A microphone result has a narrower meaning than β€œsinging score”

The classroom supports lesson playback, microphone capture, target-versus-detected pitch, timing and valid-coverage feedback.

It does not follow that an arbitrary song imported into the main window has been assessed. At present, the classroom works with its course materials. The imported-song workflow and the lesson-assessment workflow are separate.

There is another distinction inside assessment: no usable signal is not evidence of good singing. The course documentation describes uncertain detection, missed notes, octave errors and coverage checks. If effective detection covers less than half the target duration, the classroom does not present a reliable overall score.

That is a more useful failure state than a precise-looking number based on a couple of detected notes. It also leaves room to say what this part of the application actually measures: pitch and timing, not vocal tone, breathing technique or interpretation.

SingLilt English classroom before recording

The English classroom before a recording. A blank feedback area is not a successful assessment.

Saving a project is different from packaging the application

A .jpp project carries the material needed to reopen it, including its page resources. The portable application package is a separate collection of binaries and runtime resources. Moving the project and moving the executable are different operations.

The installer consumes the same staged file set as the portable package instead of maintaining its own independent list of libraries and models. That reduces the chance that the two delivery paths disagree about what belongs beside the executable.

There are several intentionally modest details around these boundaries:

  • Unapplied edits remain distinct from saved project state.
  • Recovery snapshots are separate from the normal save path.
  • Interface themes and language preferences are not score content.
  • Optional audio-analysis components are not prerequisites for basic playback.
  • Exported practice WAV files contain synthesized score playback, not a recovered original performance.

None of these features makes recognition infallible. They make it easier to inspect a result, correct it, save it and understand what the application has actually done.

Source material

The implementation and its design notes are public:

The project code is MIT-licensed; bundled third-party resources have their own terms.

The boundary worth keeping, in this application or another recognition-based tool, is simple: a generated candidate can be useful without becoming authoritative.

Drafting note: This article was drafted by an AI writing assistant from the project's repository documentation. The application screenshots are real captures.

Top comments (0)