an oto guide · by ear nº5

Stems, by ear

A finished track is one wave. Inside it: a drum kit, a bassline, chords, a voice. Pulling them back apart is called separation. It was impossible for decades and is now a button on a DJ deck. This page lets you hear what that button does, why it is genuinely hard, and what it unlocks.

headphones on · nothing to install · every sound is made live in your browser

Same four bars. The second button is what this whole page is about: a voice, alone, that the listener was never given alone.

A confession that is also the lesson. Every sound here is synthesized live, so this page secretly holds the true separate parts. Real separation software starts with one finished wave and nothing else. Wherever a demo below cheats, the cheat is labeled, and the gap between the cheat and reality is exactly what chapter 04 is about.

01 · the multitrack

A mix starts as separate parts.

In the studio, a song is never one recording. It is dozens: a microphone on each drum, a line from the bass, layers of synths, take after take of the voice. Bounce related layers into grouped tracks, such as all the drums or all the vocals. Each group is called a stem. The stems are what the producer has and the listener doesn't.

Then comes release day, and everything is summed. As the signals guide put it: mixing is addition, sample by sample, two signals becoming one. Do that addition across every stem and you get the single stereo wave that goes to streaming, to vinyl, to your crate. The parts are gone. Only their sum survives.

Below is the moment before the sum. Four stems, drums, bass, chords, and voice, play one loop of house. This mixer is the studio's view. Keep its four colors in mind; they thread through every demo on this page.

rig 01 · the stems, before the sum

The screen shows the waveform of whatever is currently summed. Solo a stem to hear one part naked; tap the colored names to mute and unmute.

TRY: Solo the voice, then the bass, then un-solo and listen to all four at once. Notice something: once everything plays, your ear can still follow each part, but the waveform on screen is just one squiggle. Your ear separates. The file does not.
In Oto: everything a DJ tool receives is the squiggle, never the mixer. Every trick in the rest of this guide is about clawing back some version of this screen from the sum alone.

02 · why it's hard

You can't un-add.

Here is the whole difficulty in one line of arithmetic: 5 + 7 = 12, but 12 has forgotten. It is also 3 + 9, and 6 + 6, and a thousand other recipes. Addition destroys the addends. A mixed sample is one number that used to be four, and the file keeps no receipt.

Think of pouring four tubes of paint into one bucket and stirring. Nothing was lost: every molecule of pigment is still in there. Yet handing someone the bucket and asking for the original tubes back is a different kind of problem entirely.

"Fine," says every engineer who ever met this problem, "the parts live in different frequency ranges: I'll just split the spectrum." This breaks down quickly, and the rig below shows why. Each stem's live spectrum is drawn in its own color, on the same axes. Watch how much territory they share.

rig 02 · four spectra, one set of axes: bass left, treble right

Each color is one stem's spectrum, measured live. Where colors stack, multiple instruments occupy the same frequencies at the same instant. The mix stores their sum as a single value.

TRY: Start with only the voice on. Now add the chords. The violet and teal shapes land almost on top of each other: a voice and a chord live in the same midrange. Add drums and watch the hi-hats camp on the voice's consonants. There is no clean fence to build.

This is why separation is not a filtering problem. At 400 Hz, at second 23, the file holds one number that is voice plus chord plus the ring of a snare. Any tool that works purely by frequency must give you all three or none. To do better, a system has to recognize what a voice is. That became practical in the 2010s. Before modern separation, DJs and karaoke machines relied on two practical tricks.

03 · the old tricks

Fifty years of clever cheating.

Hack one: EQ carving. A voice mostly lives between roughly 200 Hz and 4 kHz, so slap a band-pass on the mix and call it a vocal isolator. You will hear the problem instantly below: a filter selects frequencies, not sources. You get the voice, plus everyone else's midrange, minus the voice's own air and body.

Hack two is smarter: the center-channel trick. In most mixes the lead vocal is placed dead center, meaning it is identical in the left and right channels, while guitars, keys, and cymbals are panned off to the sides. So subtract right from left. Anything identical in both channels, including the entire center, cancels to silence. Radio engineers called it Out Of Phase Stereo; the karaoke industry of the 1970s was essentially built on it.

It really works, and it has two famous flaws. The kick and the bass are usually mixed center too, so they vanish along with the voice, leaving a thin, phasey ghost of a track. And the trick only runs in one direction: subtraction can remove the center, but there is no equivalent move that keeps only the center. The mid signal (left plus right) still contains everything.

rig 03 · one stereo mix, four ways to interrogate it

this rig's mix: drums panned left · chords panned right · bass and voice dead center

Switch modes while it plays. The spectrum shows whatever survives.

TRY: Play the full mix, then hit L − R. The voice is gone through pure subtraction. Listen to what left with it: the bass disappeared too, and the result sounds hollow. Now try L + R hoping to keep only the voice: no luck, everything is still there. Finally, try the EQ carve. The voice is muffled and the rest of the mix still leaks through.
In Oto: the karaoke trick is still useful as a test. If L − R makes a vocal vanish, that vocal is center-mixed and mono, which is a useful fact about a track. But nothing in this chapter can hand you an acapella. For that, the machine has to stop doing arithmetic and start listening.

04 · machine listening

A spectrogram makes sources visible.

Run the Fourier transform over and over as a track plays and you get a spectrogram: a photograph of sound, time running sideways, frequency running up, brightness showing energy. And on that photograph, instruments look like something. A voice appears as parallel harmonic bands that bend together with its pitch and vibrato. A drum hit is a vertical strike of lightning across every frequency at once. A bassline is a thick glowing floor.

Reframed as a picture, separation becomes a problem computer vision already knew: cutting a figure out of a crowded photo. The machine estimates a mask that assigns part of each spectrogram pixel to the voice. It applies that mask and turns the selected energy back into sound.

The reason this finally worked in the late 2010s is the answer key. Researchers assembled songs with their true multitracks, the studio view from chapter 01. A neural network could then guess a mask, compare it with the real stem, and improve over millions of examples. That recipe gave the world Spleeter in 2019, then Open-Unmix, then Demucs, which learned to work partly on the raw waveform itself. Descendants of these models are what now run inside DJ software, in real time, on a laptop.

The lab below stages that whole history on one loop. The cheat, labeled: the "today" setting plays our true synthesized stem, because a real modern model gets remarkably close to truth and we happen to own the truth. The "2019" setting deliberately damages the stem the way early masks did. The EQ setting is chapter 03, for scale.

rig 04 · the separation lab: output drawn as a scrolling spectrogram

what to pull out
how to pull it

Watch the spectrogram while you switch: ribbons are pitched notes, vertical strikes are drums. The 2019 voice flickers and warbles; today's barely does.

TRY: Target the voice and step through the three eras. The 1975 method still contains much of the original mix. The 2019 method produces a recognizable voice with a watery shimmer and a faint hi-hat ghost. Today is close enough to the truth to play on a big system. Then target the instrumental on the 2019 setting and listen closely: the voice isn't gone, it's buried, wobbling quietly inside the track that supposedly lost it.

05 · the artifacts

What failure sounds like.

Separation errors have a vocabulary, and once you learn it you will hear it in every free stem-splitter on the internet. All of it comes from one cause: the mask has to make a call on every pixel, and where sounds truly share pixels, every call loses something.

Bleed occurs when the mask includes part of another source, such as a hi-hat or string pad inside an isolated vocal. Wateriness is the underwater, warbling quality of early acapellas. It comes from the mask flickering between yes and no, which carves unstable holes in the sound. Researchers politely call the residue musical noise. Smearing affects brief broadband sounds such as cymbals, claps, and vocal consonants. The mask softens their edges, so attacks lose definition.

Notice the pattern: the failures cluster where instruments genuinely overlap, which chapter 02 showed is most of the picture. A shaker and the top of a voice are almost the same pixels. Perfect separation of truly shared energy is not an engineering milestone waiting to be reached; some of it is information the sum simply no longer holds.

TRY: Go back up to rig 04, set 2019, target the voice, and collect the full set by ear: the warble (wateriness), the hi-hat ghost (bleed), and the softened attack at the start of each note (smearing). Then switch to today and notice how much quieter each artifact becomes. Listen closely; they are reduced, not gone.
In Oto: artifact-spotting is a machine's job too. Separation quality varies by track. Dense mixes and busy cymbals usually separate less cleanly, so a stem should carry a confidence score the way a key tag carries one in the keys guide. A watery acapella is fine for cueing in headphones and a risk on a festival rig. A tool that knows the difference should say so before you air it.

06 · the payoff

Every track becomes a multitrack.

Now the fun part. If any finished track can be split live, then every record you own quietly becomes the mixer from chapter 01, and a whole set of moves that used to require luck, label connections, or a studio become buttons:

The instant acapella. DJs once hunted rare acapella releases and B-sides. Now the vocal of anything is one gesture away. And so is the instrumental, which is how a booth kills a vocal to talk over a break.

The clean blend. The oldest trainwreck in DJing is two kicks fighting during a transition. Mute the incoming track's drums until the handover moment and the fight never starts. Same trick for two vocalists talking over each other.

The mashup, live. One track's voice over another track's groove, assembled on the fly. Tempo and key still rule: the wheel does not care how you obtained your acapella. But within those rules, your library offers many more combinations. This is the feature that moved stems from research papers into Serato, Traktor, djay, and rekordbox, and it is the crowd-pleaser below.

rig 05 · the stem deck: two tracks, one clock, eight switches

track a · house
track b · breaks
one-tap moves

Both tracks share a tempo and a key. The keys guide explains why 8A matters here. Flip stems mid-loop; the clock never stops.

TRY: Play track A for a few bars, then tap A's voice over B. A familiar voice now rides a different groove. You just built that live remix with five stems and one clock. Then hit drop to acapella and count how long the naked voice holds the room before you need the drums back. Every stem DJ learns that number fast.

One honest footnote for the toybox: separation makes acapellas available, not licensed. A model can un-mix a song; it cannot un-mix the rights. The law here is still catching up to the button, and working DJs should know both facts.

07 · how oto uses this

Transitions become handoffs.

Everything upstream of this chapter changes what a transition is. Without stems, a blend is a crossfade: two full mixes trading loudness, with EQ as a blunt referee. With stems, a blend is a handoff: track B's drums take over while track A's voice stays; the bass changes hands on a downbeat; the old chords dissolve last. That is arrangement, performed live.

Words map onto stems naturally. "Keep her voice, lose the drums." "Bring the new bassline under this vocal." In a stem-aware system, those sentences are not vibes to interpret: they are literally mixer moves, the four faders of rig 01 addressed by name.

Stems open doorways. The keys guide showed that drum-only sections are free doorways between distant keys, because drums have no pitch to clash. Separation means a doorway no longer has to exist in the arrangement: mute the melodic stems and you have built one anywhere in the track. A planner that can make doorways, not just find them, drafts far bolder journeys.

And confidence stays in the loop. Because separation quality varies, every extracted stem in Oto carries the artifact question from chapter 05: is this acapella booth-quality or headphone-quality? Drafts use clean stems, route around watery ones, and tell you which is which. A weak mask changes the plan instead of putting a damaged stem into the room.

In short, one wave goes in and four faders come out, each with a confidence score. That gives you a night you can direct in sentences.

08 · keep going

Where the trail continues.

Oto Labs · keep listening

Continue by ear.

Oto · private beta

Direct the handoff.

Keep the voice, hand over the drums, or open a doorway anywhere. Oto turns those plain instructions into stem moves.

Join the waitlist