an oto guide · by ear nº5
A finished track is one wave. Inside it: a drum kit, a bassline, chords, a voice. Pulling them back apart is called separation. It was impossible for decades and is now a button on a DJ deck. This page lets you hear what that button does, why it is genuinely hard, and what it unlocks.
headphones on · nothing to install · every sound is made live in your browser
Same four bars. The second button is what this whole page is about: a voice, alone, that the listener was never given alone.
01 · the multitrack
In the studio, a song is never one recording. It is dozens: a microphone on each drum, a line from the bass, layers of synths, take after take of the voice. Bounce related layers into grouped tracks, such as all the drums or all the vocals. Each group is called a stem. The stems are what the producer has and the listener doesn't.
Then comes release day, and everything is summed. As the signals guide put it: mixing is addition, sample by sample, two signals becoming one. Do that addition across every stem and you get the single stereo wave that goes to streaming, to vinyl, to your crate. The parts are gone. Only their sum survives.
Below is the moment before the sum. Four stems, drums, bass, chords, and voice, play one loop of house. This mixer is the studio's view. Keep its four colors in mind; they thread through every demo on this page.
rig 01 · the stems, before the sum
The screen shows the waveform of whatever is currently summed. Solo a stem to hear one part naked; tap the colored names to mute and unmute.
02 · why it's hard
Here is the whole difficulty in one line of arithmetic: 5 + 7 = 12, but 12 has forgotten. It is also 3 + 9, and 6 + 6, and a thousand other recipes. Addition destroys the addends. A mixed sample is one number that used to be four, and the file keeps no receipt.
Think of pouring four tubes of paint into one bucket and stirring. Nothing was lost: every molecule of pigment is still in there. Yet handing someone the bucket and asking for the original tubes back is a different kind of problem entirely.
"Fine," says every engineer who ever met this problem, "the parts live in different frequency ranges: I'll just split the spectrum." This breaks down quickly, and the rig below shows why. Each stem's live spectrum is drawn in its own color, on the same axes. Watch how much territory they share.
rig 02 · four spectra, one set of axes: bass left, treble right
Each color is one stem's spectrum, measured live. Where colors stack, multiple instruments occupy the same frequencies at the same instant. The mix stores their sum as a single value.
This is why separation is not a filtering problem. At 400 Hz, at second 23, the file holds one number that is voice plus chord plus the ring of a snare. Any tool that works purely by frequency must give you all three or none. To do better, a system has to recognize what a voice is. That became practical in the 2010s. Before modern separation, DJs and karaoke machines relied on two practical tricks.
03 · the old tricks
Hack one: EQ carving. A voice mostly lives between roughly 200 Hz and 4 kHz, so slap a band-pass on the mix and call it a vocal isolator. You will hear the problem instantly below: a filter selects frequencies, not sources. You get the voice, plus everyone else's midrange, minus the voice's own air and body.
Hack two is smarter: the center-channel trick. In most mixes the lead vocal is placed dead center, meaning it is identical in the left and right channels, while guitars, keys, and cymbals are panned off to the sides. So subtract right from left. Anything identical in both channels, including the entire center, cancels to silence. Radio engineers called it Out Of Phase Stereo; the karaoke industry of the 1970s was essentially built on it.
It really works, and it has two famous flaws. The kick and the bass are usually mixed center too, so they vanish along with the voice, leaving a thin, phasey ghost of a track. And the trick only runs in one direction: subtraction can remove the center, but there is no equivalent move that keeps only the center. The mid signal (left plus right) still contains everything.
rig 03 · one stereo mix, four ways to interrogate it
Switch modes while it plays. The spectrum shows whatever survives.
04 · machine listening
Run the Fourier transform over and over as a track plays and you get a spectrogram: a photograph of sound, time running sideways, frequency running up, brightness showing energy. And on that photograph, instruments look like something. A voice appears as parallel harmonic bands that bend together with its pitch and vibrato. A drum hit is a vertical strike of lightning across every frequency at once. A bassline is a thick glowing floor.
Reframed as a picture, separation becomes a problem computer vision already knew: cutting a figure out of a crowded photo. The machine estimates a mask that assigns part of each spectrogram pixel to the voice. It applies that mask and turns the selected energy back into sound.
The reason this finally worked in the late 2010s is the answer key. Researchers assembled songs with their true multitracks, the studio view from chapter 01. A neural network could then guess a mask, compare it with the real stem, and improve over millions of examples. That recipe gave the world Spleeter in 2019, then Open-Unmix, then Demucs, which learned to work partly on the raw waveform itself. Descendants of these models are what now run inside DJ software, in real time, on a laptop.
The lab below stages that whole history on one loop. The cheat, labeled: the "today" setting plays our true synthesized stem, because a real modern model gets remarkably close to truth and we happen to own the truth. The "2019" setting deliberately damages the stem the way early masks did. The EQ setting is chapter 03, for scale.
rig 04 · the separation lab: output drawn as a scrolling spectrogram
Watch the spectrogram while you switch: ribbons are pitched notes, vertical strikes are drums. The 2019 voice flickers and warbles; today's barely does.
05 · the artifacts
Separation errors have a vocabulary, and once you learn it you will hear it in every free stem-splitter on the internet. All of it comes from one cause: the mask has to make a call on every pixel, and where sounds truly share pixels, every call loses something.
Bleed occurs when the mask includes part of another source, such as a hi-hat or string pad inside an isolated vocal. Wateriness is the underwater, warbling quality of early acapellas. It comes from the mask flickering between yes and no, which carves unstable holes in the sound. Researchers politely call the residue musical noise. Smearing affects brief broadband sounds such as cymbals, claps, and vocal consonants. The mask softens their edges, so attacks lose definition.
Notice the pattern: the failures cluster where instruments genuinely overlap, which chapter 02 showed is most of the picture. A shaker and the top of a voice are almost the same pixels. Perfect separation of truly shared energy is not an engineering milestone waiting to be reached; some of it is information the sum simply no longer holds.
06 · the payoff
Now the fun part. If any finished track can be split live, then every record you own quietly becomes the mixer from chapter 01, and a whole set of moves that used to require luck, label connections, or a studio become buttons:
The instant acapella. DJs once hunted rare acapella releases and B-sides. Now the vocal of anything is one gesture away. And so is the instrumental, which is how a booth kills a vocal to talk over a break.
The clean blend. The oldest trainwreck in DJing is two kicks fighting during a transition. Mute the incoming track's drums until the handover moment and the fight never starts. Same trick for two vocalists talking over each other.
The mashup, live. One track's voice over another track's groove, assembled on the fly. Tempo and key still rule: the wheel does not care how you obtained your acapella. But within those rules, your library offers many more combinations. This is the feature that moved stems from research papers into Serato, Traktor, djay, and rekordbox, and it is the crowd-pleaser below.
rig 05 · the stem deck: two tracks, one clock, eight switches
Both tracks share a tempo and a key. The keys guide explains why 8A matters here. Flip stems mid-loop; the clock never stops.
One honest footnote for the toybox: separation makes acapellas available, not licensed. A model can un-mix a song; it cannot un-mix the rights. The law here is still catching up to the button, and working DJs should know both facts.
07 · how oto uses this
Everything upstream of this chapter changes what a transition is. Without stems, a blend is a crossfade: two full mixes trading loudness, with EQ as a blunt referee. With stems, a blend is a handoff: track B's drums take over while track A's voice stays; the bass changes hands on a downbeat; the old chords dissolve last. That is arrangement, performed live.
Words map onto stems naturally. "Keep her voice, lose the drums." "Bring the new bassline under this vocal." In a stem-aware system, those sentences are not vibes to interpret: they are literally mixer moves, the four faders of rig 01 addressed by name.
Stems open doorways. The keys guide showed that drum-only sections are free doorways between distant keys, because drums have no pitch to clash. Separation means a doorway no longer has to exist in the arrangement: mute the melodic stems and you have built one anywhere in the track. A planner that can make doorways, not just find them, drafts far bolder journeys.
And confidence stays in the loop. Because separation quality varies, every extracted stem in Oto carries the artifact question from chapter 05: is this acapella booth-quality or headphone-quality? Drafts use clean stems, route around watery ones, and tell you which is which. A weak mask changes the plan instead of putting a damaged stem into the room.
In short, one wave goes in and four faders come out, each with a confidence score. That gives you a night you can direct in sentences.
08 · keep going
Oto Labs · keep listening
Oto · private beta
Keep the voice, hand over the drums, or open a doorway anywhere. Oto turns those plain instructions into stem moves.
Join the waitlist