This is the third iteration of the setup. I'm using Meta's free and open source DEMUCS for stem extraction, and librosa for analyzing the audio data. Do I know what any of that means? Hell no. But Gemini/ChatGPT/Claude knows and that's good enough for my little Python SOP!