From Passion to Precision Building Music Splitter Pro and Vocal Splitter Pro with BS RoFormer and SATB
- Masatoshi Hirakata
- 7 days ago
- 8 min read
Some songs feel impossible to take apart because they were never meant to be taken apart. A Beach Boys harmony stack can arrive like sunlight through stained glass. A Beatles rhythm section can hide tiny guitar, piano, tambourine, and bass details in the same bright corner of the mix. That is exactly what made building Music-Splitter-Pro and Vocal-Splitter-Pro so addictive.
These projects grew out of a simple music fan’s question: what would it feel like to step inside a favourite record?
Not just listen to the song, but sit with the bass line. Hear the drum pocket without the guitars. Study how the keyboard supports the vocal. Notice how a lead singer moves just slightly ahead of the backing voices. For anyone who loves arranging, mixing, learning by ear, or just understanding why certain records work, source separation feels close to magic.
The magic, of course, is mostly maths, models, training data, signal processing, and a lot of listening. The work has been rewarding, but it has also made one thing clear: separating music is easy to explain and hard to do well.

Why I wanted to build these tools
My love of music has always been a mix of feeling and curiosity. A great song is emotional first. It hits before it explains itself. But after that first hit, I want to know how it works.
Why does the bass line feel so melodic without crowding the vocal? Why do the drums sound simple until you isolate the ghost notes? Why does a chorus suddenly widen, even when the chords barely change? Why do some backing vocals blend into one golden cloud while others keep their individual shape?
That curiosity led to two connected tools:
Music-Splitter-Pro
Aimed at separating a full track into core instrumental and vocal stems, such as bass, drums, guitars or keyboard, and vocals.
Vocal-Splitter-Pro
Focused on the more delicate task of separating vocal content, especially lead vocals from supporting chorus or harmony parts.
Music-Splitter-Pro is about opening the arrangement. Vocal-Splitter-Pro is about opening the voice stack.
They share the same spirit, but they are very different problems. Instrumental separation asks the model to identify sound sources with distinct textures. Vocal separation asks it to make decisions between sources that often share the same singer, similar tone, overlapping pitch ranges, and the same reverb.
That second challenge can be brutal.
The problem with classic records
Modern recordings often give separation models more clues. Contemporary productions may have tight low end, centred vocals, cleaner frequency slots, and clearer stereo placement. Older records are different.
Songs by bands like the Beach Boys and the Beatles are wonderful test cases because they are rich, dense, and human. They also break many assumptions that simple separation systems quietly rely on.
The Beach Boys, for example, are famous for layered vocal arrangements. Their harmonies often behave less like separate voices and more like one blended instrument. The parts overlap in pitch and tone. The lead line may sit inside the harmony rather than above it. The backing vocals can carry emotional weight equal to the main vocal.
The Beatles bring a different kind of challenge. Their recordings often contain:
Bass with strong melodic movement
Drums that share space with hand percussion, piano, or rhythm guitar
Guitars and keyboards that overlap in midrange frequencies
Vocals treated with double tracking, room sound, tape colour, and effects
Mix decisions shaped by the recording tools of the time
Those qualities are part of the charm. They are also why separation can produce strange artefacts. A guitar transient might leak into the vocal stem. A snare might leave a shadow in the keyboard stem. A harmony vocal might vanish because the model decides it belongs to the lead.
When the source material is beautiful and messy, the tool has to be both powerful and careful.

Building Music-Splitter-Pro with BS-RoFormer
For Music-Splitter-Pro, I chose BS-RoFormer as the core technology.
BS-RoFormer is well suited to music source separation because it treats the track as more than one flat signal. In plain terms, it can work across frequency bands and learn relationships over time. That matters because musical sources do not live in neat boxes.
A bass guitar is not only “low frequencies”. It has pick attack, fret noise, harmonics, and sometimes distortion higher up the spectrum. Drums are not only transients. A kick can have body, click, room tone, and low-end bloom. A vocal can occupy the centre, but its breath, sibilance, delay, and reverb can spread elsewhere.
A good splitter needs to recognise patterns, not just cut frequencies.
Music-Splitter-Pro currently focuses on four broad outputs:
Stem | What the tool is trying to capture | Why it is difficult |
Bass | Bass guitar or low-end melodic instrument | Harmonics overlap with guitar, piano, and vocal warmth |
Drums | Kick, snare, cymbals, percussion, room sound | Cymbals and vocal sibilance can blur together |
Guitars or keyboard | Midrange harmonic instruments | Piano, guitar, organ, and strings can share similar space |
Vocals | Lead and backing vocal content | Harmonies, reverb, and double tracking can confuse the model |
The strongest results so far have come from tracks where the arrangement leaves consistent space between parts. Bass and drums often separate well enough to study the groove. On many songs, the vocal stem becomes clear enough for transcription, remix sketches, or arrangement analysis.
The hard cases are the ones I love most.
A busy 1960s mix can make the “guitars or keyboard” stem feel like a musical attic. Everything interesting in the midrange wants to live there. Rhythm guitar, piano, organ, tambourine bleed, room reflections, and bits of vocal can all compete. Sometimes that stem is musically useful, even when it is not perfectly clean. Other times it needs more refinement.
The goal is not to pretend the model can undo history. It cannot return a mono bounce to the original studio multitrack. The goal is to create stems that are clear, useful, and honest enough to help people listen more closely.
What success sounds like
A successful split is not only a quiet noise floor or a pretty spectrogram. It has to feel musical.
When Music-Splitter-Pro works well, the bass stem still has intention. The drummer’s feel remains intact. The keyboard or guitar part keeps its rhythm and colour. The vocal is not hollowed out or full of watery artefacts.
I use a few listening checks throughout development:
Can the bass line be followed from start to finish?
Do the drums still feel like a performance rather than a set of clicks?
Does the vocal keep natural phrasing?
Are artefacts distracting, or only noticeable under close inspection?
Does each stem help reveal something about the arrangement?
That last point matters. A technically “clean” output can still be dull if it removes the musical detail that made the part worth hearing. Separation is not only extraction. It is preservation.
There have been genuine wins. Hearing a buried bass line step forward from a dense mix is thrilling. Isolating drums from a classic recording can reveal how much the groove depends on tiny timing choices. Pulling a vocal forward can expose phrasing and breath control that disappear in the full track.
Those moments keep the project moving.

Building Vocal-Splitter-Pro with SATB
If Music-Splitter-Pro is about separating the band, Vocal-Splitter-Pro is about separating the choir inside the record.
For this tool, I have been working with SATB as the guiding technology and structure. In musical terms, SATB refers to soprano, alto, tenor, and bass voice parts. In the context of Vocal-Splitter-Pro, the idea is to organise and separate vocal material by role and range so the tool can better understand stacked harmonies.
This is especially useful for songs where the vocal arrangement is not just “lead plus backing”. Many classic vocal parts behave like miniature choral writing. A high harmony might carry the emotional lift. A lower voice might stabilise the chord. A middle harmony might rub against the melody and create tension.
The hard part is that pop vocals do not follow clean choral rules. A lead singer can move through the same range as a harmony part. A backing vocal can briefly become the hook. A chorus can include multiple voices singing the same rhythm as the lead. Double-tracked vocals can look like two sources but act like one performance.
That makes lead-vocal separation one of the most difficult parts of the whole project.
The lead vocal and chorus problem
Separating lead vocals from chorus parts sounds simple until you listen closely.
A lead vocal often has clues:
It may sit in the centre of the stereo image.
It may be slightly louder.
It may carry the main lyric.
It may have a different compression or reverb treatment.
It may stay present through verses and choruses.
But those clues fall apart often.
In many Beach Boys-style harmony passages, the lead voice blends with the others by design. The production wants the voices to become one instrument. In Beatles-style arrangements, double tracking and group vocals can blur the line between lead, harmony, and response. Sometimes the lead is not the loudest thing. Sometimes a backing part shares the same words, timing, and pitch region.
The model then faces a musical judgement. Should that voice stay with the lead because it reinforces the melody? Should it move to the chorus stem because it is part of the group texture? If two singers are locked together, can the tool separate them without damaging both?
This is where artefacts become more personal. A little drum leakage in a bass stem may be acceptable. A damaged vocal is harder to ignore. Human hearing is very sensitive to voice quality. We notice smearing, lisping, phase-like tones, and missing consonants quickly.
The current progress is encouraging, but not finished. Vocal-Splitter-Pro can separate some lead lines cleanly, especially when the lead has a strong centre position and distinct tone. It also handles some harmony stacks well when the parts occupy different ranges.
The toughest sections are chorus peaks, stacked unisons, and close harmonies. When singers share too much pitch, timing, and tone, the model needs more context than a single moment in the waveform can provide. It has to understand the arrangement over time.
Lessons from development so far
The main lesson is that music separation is an engineering problem and a listening problem at the same time.
Metrics can point in the right direction, but the ear makes the final call. A model might score well and still produce a vocal stem that feels artificial. Another output might contain a little bleed but preserve the emotional shape of the performance. For music lovers, the second result can be more useful.
I have also learned to treat each era of recording differently. A clean modern production and a dense classic mix need different expectations. With older material, there may be tape saturation, bouncing, mono elements, phase quirks, and effects printed into the same track. Those are not bugs. They are part of the recording.
For developers working on similar tools, a few practical ideas have helped:
Build listening tests into every stage
Do not rely only on visual outputs or scores.
Use difficult songs early
Easy tracks can hide weak separation behaviour.
Inspect stems in context
Solo listening matters, but a stem also needs to work when recombined.
Name the failure clearly
“Bad separation” is too vague. Is it bleed, missing transients, vocal smearing, phase sound, or wrong source assignment?
Respect the music
The goal is not to strip a song bare. The goal is to learn from it.

Where the projects are heading next
Music-Splitter-Pro and Vocal-Splitter-Pro are still growing. The next stage is less about adding flashy features and more about improving trust.
For Music-Splitter-Pro, that means cleaner handling of midrange instruments, better reduction of cross-stem bleed, and more stable results on older recordings. I want the guitar or keyboard stem to become more musically readable, especially when piano, rhythm guitar, and organ all share the same space.
For Vocal-Splitter-Pro, the focus is clearer lead and chorus separation. The aim is not only to isolate a lead vocal, but to preserve the beauty of the surrounding voices. If the chorus stem loses its blend, the tool has missed the point. If the lead stem sounds clean but lifeless, that is not good enough either.
The long-term dream is a pair of tools that serve both fans and makers. A songwriter could study harmony movement. A bassist could learn a hidden line. A producer could examine how a classic arrangement creates width. A developer could test separation ideas against real musical problems.
More than anything, these projects have deepened my respect for the records that inspired them. The more I try to separate the parts, the more I hear how carefully they belong together.
That is the strange beauty of building tools like this. You start by trying to pull music apart. If the work goes well, you end up loving the whole song even more.