DEV Community

Jakub
Jakub

Posted on

1,200+ songs in: what Magical Song by Inithouse learned about turning a story into real vocals

We have been running Magical Song for several months now. It is an AI custom song generator: you type a story, pick a genre, and get back a studio-quality track with real vocals. Over 1,200 songs generated so far, 4.9 out of 5 average rating, 20+ genres. Here is what the data actually taught us about mapping a personal story to music.

The input problem nobody warns you about

Most users type between 40 and 120 words as their story input. That range works well. Below 30 words, the lyrics engine does not have enough material to build a chorus that feels personal. Above 200, it starts picking details that sound poetic but miss the emotional core the person wanted.

We tracked which inputs produced songs that users rated 5 out of 5 versus 3 or lower. The pattern was consistent: the best inputs contained one specific name, one concrete moment ("the morning she left for college"), and one emotion stated plainly. Inputs that read like a novel summary ("we met in 2014, then moved to Berlin, got a dog named Max, started a business...") produced lyrics that felt like a Wikipedia timeline set to music.

Where story-to-chorus mapping breaks

The chorus is the hardest part. A good chorus needs repetition and a single emotional hook. When the input story has two equally weighted emotional peaks ("our wedding day AND the birth of our daughter"), the generator tries to serve both and the chorus loses focus.

We addressed this by adding an internal ranking step that picks the strongest emotional signal from the input and anchors the chorus there. The secondary moments get woven into verses instead. This single change moved our average satisfaction score from 4.6 to 4.9 across 300+ songs.

Names in lyrics: harder than it sounds

About 68% of Magical Song orders include at least one proper name. Names are tricky for two reasons. First, syllable count: "Christopher" does not fit where "Tom" does. Second, phonetic flow: some names simply do not rhyme with anything natural in English, and forcing a rhyme produces cringe.

Our approach: the generator treats the name as a fixed anchor and builds the line around it rather than trying to make it rhyme. If the name lands at the end of a line, the following line uses a near-rhyme or shifts to an ABCB scheme instead of AABB. This keeps the song sounding intentional rather than awkward.

Genre distribution: what people actually pick

We expected pop and rock to dominate. They do, but the long tail surprised us.

Genre Share of songs Typical use case
Pop 28% Birthdays, friend tributes
Acoustic/Folk 18% Wedding gifts, anniversaries
Rock 12% Father's Day, band-style tributes
R&B/Soul 9% Romantic dedications
Country 8% Family stories, memorial songs
Hip-hop/Rap 7% Graduation, friend roasts
Jazz 5% Retirement gifts
Lullaby 4% New baby, first birthday
Other (10+ genres) 9% Holiday themes, novelty, pets

Lullabies punched above their weight in satisfaction scores. A four-line lullaby with a baby's name and one detail ("born on a rainy Tuesday") consistently hit 5 out of 5. The constraint of simplicity played in the generator's favor.

Song length: the 2-minute sweet spot

We tested outputs between 1 and 4 minutes. Songs under 90 seconds felt unfinished to users. Songs over 3 minutes started to repeat ideas or pad with generic filler lines. The 1:50 to 2:30 range hit the best balance of feeling complete without dragging.

One unexpected finding: users who chose hip-hop and rap consistently preferred longer tracks (2:30 to 3:00), while lullaby and acoustic users preferred shorter ones (1:30 to 2:00). We now adjust default length by genre rather than using a single target.

What we still get wrong

Humor. Songs meant to be funny ("roast my friend for always being late") are the lowest-rated category. Humor requires timing, cultural context, and the kind of specificity that a short text input cannot fully convey. We are experimenting with structured humor prompts ("pick one running joke, one exaggeration, one callback") but it remains the weakest spot.

Multi-language inputs also need work. A user writing half in Spanish, half in English gets a song that code-switches awkwardly rather than blending naturally. Our roadmap includes dedicated bilingual models, but right now we recommend sticking to one language per song.

The takeaway from 1,200 songs

Building Magical Song taught us that the hard part of AI music is not the audio generation. That technology keeps improving on its own. The hard part is the translation layer: taking a messy, emotional, deeply personal story and turning it into structured lyrics that feel like they were written by someone who knows you. Every improvement we have shipped at Inithouse on this product has been about that translation, not about making the beat louder.

We ship a growing portfolio of products at Inithouse. If you are curious about another one: Be Recommended scores how AI models like ChatGPT and Claude recommend your brand.

Top comments (0)