🧠

Mel-Band RoFormer Karaoke Model — How 913MB Isolates Lead Vocals Only

RoPE + mel-scale frequency decomposition + hierarchical time-freq Transformer — with code and config

Model File Info

Item Value
Filename mel_band_roformer_karaoke_aufr33_viperx_sdr_10.1956.ckpt
Size 913 MB
Format PyTorch pickle checkpoint
Config config_mel_band_roformer_karaoke.yaml (1.72 kB)
Hosted HuggingFace

RoFormer — Rotary Position Embedding

RoFormer uses RoPE instead of standard position encodings. RoPE rotates Query and Key vectors by position, naturally decaying attention between distant tokens.

from rotary_embedding_torch import RotaryEmbedding
time_rotary_embed = RotaryEmbedding(dim=64)
# In Attention.forward():
q = self.rotary_embed.rotate_queries_or_keys(q)
k = self.rotary_embed.rotate_queries_or_keys(k)

Separate RoPE for time and frequency axes.

Mel-Band vs Band-Split

BS-RoFormer Mel-Band RoFormer
Band split Heuristic, non-overlapping, 62 bands Mel-scale, 50% overlap, 60 bands
Perceptual weighting No Yes (matches human hearing)
Parameters (L=6) 72.2M 84.2M

Mel scale reflects how humans perceive frequency — narrower bands at low frequencies, wider at high. 50% overlap reduces artifacts at band boundaries.

Full Inference Pipeline

  1. STFT: stereo audio → complex spectrogram [batch, 1025, time, 2]
  2. Mel-Band Projection: 1025 freq bins → 60 mel bands → linear projection → [batch, time, 60, 384]
  3. Hierarchical Transformer (6 layers): alternating time-axis and freq-axis Transformers with RoPE
  4. Mask Estimation: per-band MLP → complex mask → multiply with original STFT
  5. ISTFT: reconstruct time-domain audio

Why Karaoke-Specific

Training config:

instruments: [karaoke, other]
target_instrument: karaoke

karaoke = instrumental + backing vocals. other = lead vocals only. Model outputs the karaoke stem; residual gives clean lead vocals.

SDR 10.1956 in Context

Model Vocal SDR
Spleeter (2020) ~5.91 dB
HTDemucs (2022) ~9.20 dB
This model 10.20 dB (lead only — harder task)
Mel-RoFormer L=6 (paper) 11.21 dB (all vocals)

Testing

pip install audio-separator[gpu]
audio-separator song.mp3 --model_filename mel_band_roformer_karaoke_aufr33_viperx_sdr_10.1956.ckpt

Key source files:

  • mel_band_roformer.py (471 lines) — full architecture

  • mdxc_separator.py — inference orchestration

  • separator.py — entry point

Key Concepts

1

STFT — convert stereo audio to complex spectrogram (n_fft=2048, hop=441)

2

Mel-band projection — decompose 1025 freq bins into 60 overlapping mel bands + linear embedding

3

Hierarchical Transformer ×6 — alternating time-axis → freq-axis Transformer (RoPE)

4

Mask estimation — per-band MLP generates complex mask → multiply with original STFT

5

ISTFT — reconstruct time-domain waveform from masked spectrogram

6

Residual — original - karaoke (model output) = lead vocals

Use Cases

Karaoke accompaniment — remove lead vocals only, keep backing vocals and chorus Vocal extraction — isolate clean lead vocals (used as WhisperX transcription input) ML architecture study — understanding Transformer-based audio processing pipelines