AI stem separation works by converting audio into a visual frequency map called a spectrogram, running that data through a trained neural network that has learned to recognize the spectral fingerprints of different instruments, and then reconstructing separate audio signals from the network's predictions. The whole process takes seconds to minutes depending on whether it runs on a cloud GPU or your local machine.

You've got a stereo mix and you want the vocal out. Or the drums. Or you want to pull the bass line from a track you didn't produce and can't find the stems for. Ten years ago that meant phase cancellation tricks that left you with something barely usable. Now you run it through a model like HTDemucs or Mel-Roformer and get a vocal track that sits cleanly enough in a new mix.

That outcome feels like magic. It isn't. There's real engineering underneath it, and understanding that engineering tells you exactly why results are brilliant on some tracks and frustrating on others. It also tells you how to fix the artifacts that survive the separation.

This guide breaks down what's actually happening inside these tools, which architectures matter, where the models fall apart, and what you do after the separation is done.

How Does an AI Model Actually "See" Audio?

Raw audio is a waveform: a single line that describes air pressure over time. Every instrument, every voice, and every room reflection are all summed together in that one line. A neural network can't isolate sources directly from a waveform with much accuracy.

So the first step is conversion to a spectrogram.

A spectrogram is a 2D image. Time runs left to right. Frequency runs bottom to top. Brightness or colour represents loudness at each frequency at each moment. A kick drum looks different from a snare. A sung vowel looks different from a picked guitar string. Each instrument leaves a visual pattern the network can learn to recognize.

This is where the phrase "spectral fingerprint" becomes literal, not metaphorical. A soprano's formants show up as bright horizontal bands in the 2-4 kHz range. A bass guitar's fundamental sits below 200 Hz with harmonics stacking upward in a predictable pattern. The model learns these shapes across thousands of training tracks.

Feature Extraction: What the Model Learns to Notice

Training means feeding the model annotated multi-track sessions where the ground truth is known: here is the mix, here is the isolated drum stem, here is the isolated vocal. The model adjusts its internal weights until its predictions match the known stems as closely as possible.

What it learns is not rules a human wrote. It learns correlations: spectral patterns, amplitude envelopes, harmonic relationships, timing cues. A vocal breathes in before a phrase. A kick drum has a sharp transient followed by a pitched tail. A hi-hat has energy above 8 kHz with no harmonic structure.

After training on tens of thousands of tracks, the model can generalize. Feed it a mix it's never heard and it predicts which frequency content belongs to which source.

Which Neural Network Architectures Power Modern Stem Separation?

Three families dominate in 2025.

U-Net

U-Net is a convolutional architecture originally developed for medical image segmentation. It encodes a spectrogram down to a compressed representation, then decodes back up to full resolution, with skip connections between matching encoder and decoder layers. Those skip connections preserve fine detail that would otherwise wash out in the bottleneck.

Spleeter, the open-source model from Deezer, used a U-Net variant and was the first tool that showed producers this technology was worth taking seriously. Fast. Decent results. Its SDR numbers on the MUSDB18 benchmark sit around 6-7 dB for vocals. That was impressive in 2019. It's the floor now.

HTDemucs

HTDemucs is a hybrid transformer architecture from Meta Research. It operates simultaneously in the time domain and the frequency domain, combining a convolutional waveform encoder with a spectrogram encoder, then fusing both through a cross-domain transformer. This dual-domain approach means it captures both fine transient detail (better in the time domain) and harmonic structure (better in the frequency domain).

HTDemucs achieves 9.00 dB SDR on the standard MUSDB18 benchmark. That's a real number on a real test set. For context, the benchmark measures how much of the target signal survives versus how much bleed and artifact gets added.

Mel-Roformer and Spectral Masking Transformers

The newest generation uses attention-based transformer architectures operating on mel-scale spectrograms. Mel-Roformer achieves approximately 10.9 dB SDR for vocal separation on MUSDB18. That's a meaningful jump over HTDemucs, and you hear it on breathy, whispery, or falsetto vocals that used to confuse earlier models.

Attention mechanisms let the model look at the whole track context when predicting any one moment. A vocal that's buried in a dense chorus gets separated more cleanly because the model can reference how that voice sounded in the sparse verse.

Music AI models built on this family deliver around 15% higher SDR compared to older architectures. That 15% translates to audibly fewer artifacts in professional use.

What Happens After the Neural Network Makes Its Predictions?

The network outputs a mask. That mask describes, for every time-frequency bin in the spectrogram, how much of the energy belongs to the target stem. Multiply the original spectrogram by the mask, and you get the target stem's spectrogram. Do this for each stem simultaneously, and the masks should sum to approximately 1.0 across all bins.

The final step is synthesis: converting that masked spectrogram back into audio. This is called the inverse short-time Fourier transform (ISTFT). It's where some artifacts are introduced, because the masking process can't perfectly reconstruct phase information.

That's why separated stems sometimes have a faint, watery smear on sustained notes or a slight metallic edge on the vocal. The frequency content is right. The phase reconstruction is approximate.

How Many Stems Can You Get?

Standard 4-stem models output: vocals, drums, bass, and other (everything else). 6-stem models add guitar and piano as separate outputs. Some specialized models trained only on vocals achieve higher SDR than multi-stem generalist models because they're not splitting attention across four or six sources at once.

We ran a dense pop mix through a 6-stem model last month. Vocals came out clean. Drums were excellent. The "piano" stem captured the Rhodes but leaked significant vocal reverb tail into it because the reverb's decay occupied the same frequency range as the piano's sustain. That's a real limitation, not a bug. It's a physics problem the model partially solves and partially can't.

Cloud GPU vs. Local Processing: Which Should You Use?

This is a tradeoff between speed, privacy, and cost. Not one answer fits everyone.

Cloud Processing

Tools like LALAL.AI and Moises run models on remote servers. LALAL.AI charges around £10/month or £84/year, with a free trial tier. Moises runs $5.99/month or $39.99/year. StemSplit uses a pay-per-minute model at $0.10/minute, so a 3-minute song costs $0.30.

Cloud processing is fast. A 3-minute track through a browser-based service typically takes 3-5 minutes on standard queue times, less on off-peak hours. You don't need a powerful computer.

The downside: your audio leaves your machine. For unreleased masters, that's a real consideration. Read the terms of service before uploading anything that isn't already public.

Local Processing

Running HTDemucs or Mel-Roformer locally through tools like Demucs (command line), or through DAW-integrated options like the stem separation built into Logic Pro 10.7+, keeps your audio private. A modern CPU will process a 4-minute track in 2-5 minutes for HTDemucs. An Apple Silicon Mac or a machine with a dedicated GPU cuts that to under 60 seconds.

Logic Pro's built-in separation uses its own model. It's not HTDemucs. We find the drum stem output is slightly cleaner than its vocal output when compared against standalone HTDemucs results. Good enough for production use without leaving your session.

For open-source local separation with the latest models, Demucs on GitHub is the honest answer. It's free. It runs locally. Setup takes 20 minutes if you're comfortable with the terminal.

Worth Bookmarking

  • Demucs (Facebook Research), Free, open-source, runs locally. HTDemucs and newer models. Best for privacy-conscious users willing to use a command line.
  • LALAL.AI, The most polished browser-based option. Clean UI, fast turnaround, multiple stem options.
  • Moises, Best for musicians who want separation plus playback tools, tempo/key detection, and a mobile app in one subscription.

Where Do These Models Still Struggle?

This is the section most explainers skip. We won't.

Synthetic and Electronic Instruments

Models trained on acoustic instrument recordings struggle with synthesizers. A Moog bass and a bass guitar occupy similar frequency ranges, but a synth pad has no acoustic fingerprint the model was trained to recognize as a specific source. It may route synth pads into the "other" stem, split them between two stems, or let bleed through into the vocal.

We fed a synth-heavy electronic track to three different tools. Vocal separation was excellent on all three. The synth arpeggiation leaked into the drum stem on two of them. Annoying, but fixable in post.

Dense Mixes

Separation quality drops when three or more sources occupy the same frequency band at the same time. A wall-of-sound chorus with heavy reverb, stacked guitars, and a vocal gives the model overlapping spectral fingerprints with no clean edges to separate. SDR numbers measured on sparse or well-separated mixes don't reflect performance on modern dense pop productions.

Practical expectation: if your mix has a lot of mid-frequency clutter in the 800Hz-2kHz range, plan for some artifact cleanup.

Post-Processing: What to Do With the Artifacts That Survive

A clean separation is the start, not the end. Here's the corrective workflow we use.

On the vocal stem: high-pass at 80-100 Hz to remove low-frequency bleed from bass and kick that the model didn't fully separate. Cut any resonant buildup around 300-500 Hz with a narrow Q (1.2-1.6) where bleed tends to accumulate. A gentle de-esser at 5-8 kHz catches any metallic artifact from the ISTFT reconstruction.

On the drum stem: transient shaping tightens any attack smear. We typically add 1.5-2 ms attack on the transient shaper to let the natural crack through, then suppress the sustain by 3-4 dB if the stem has obvious wash from cymbals bleeding into the drum body.

On any stem with low-level artifact hum: a gentle noise gate with a -40 dB threshold removes the spectral residue that sits below the music but above silence.

Summary

AI stem separation converts audio to spectrograms, runs them through trained neural networks that recognize spectral fingerprints, applies frequency masks, and reconstructs separate audio streams. Modern architectures like HTDemucs (9.00 dB SDR) and Mel-Roformer (~10.9 dB SDR) produce results usable in professional workflows. Cloud tools are fast and accessible. Local tools like Demucs are free and private. Quality degrades on dense mixes and synthetic sources. Post-processing with high-pass filters, transient shapers, and targeted EQ cuts handles most of what the model leaves behind. Understand the tool's limits and you'll stop being surprised when it hits them.

Frequently Asked Questions

What is SDR and why does it matter for stem separation?

SDR stands for Signal-to-Distortion Ratio. It measures how much of the target stem (the signal) makes it through versus how much bleed, artifact, and reconstruction error (the distortion) gets added. Higher is better. The MUSDB18 benchmark is the standard test set: Mel-Roformer scores around 10.9 dB on vocals, HTDemucs scores 9.00 dB overall. A 1-2 dB improvement in SDR is audible in professional use, especially on breathy or quiet passages.

Can I use separated stems in a commercial release?

The technical answer and the legal answer are different. Technically, yes. Legally, it depends entirely on copyright ownership of the original track. Separating stems from a track you don't own doesn't transfer any rights to those stems. If you're separating your own productions for remix or re-release, there's no issue. If you're separating someone else's music, speak to a music lawyer before using the output commercially.

Why does the bass stem sometimes contain kick drum bleed?

Bass guitar and kick drum share significant overlap in the 40-120 Hz range. When both sources are loud in that band simultaneously, the model's frequency mask can't assign those bins exclusively to one source. Some kick energy leaks into the bass stem and vice versa. The fix is a high-pass filter on the bass stem at around 35-40 Hz and a low-cut on the kick stem around 50-60 Hz, combined with gentle multiband compression to control the overlap region without destroying the tone of either stem.