Excited to share our @ISMIR2021 paper: Sequence-to-Sequence Piano Transcription with Transformers
Generic encoder-decoder Transformer + Spectrogram inputs + MIDI-like event outputs = SotA results!
arxiv.org/abs/2107.09142
With @iansimon, @rigeljs, @ethanmanilow, and @jesseengel
Jul 21, 2021 · 12:07 AM UTC
5
33
132
We use the "small" encoder-decoder architecture from the T5 paper (@colinraffel et al.) and feed it spectrograms as continuous inputs.
Outputs are MIDI-like events, not piano rolls: Note On/Off, Velocity, Time Shift, EOS
Decoding is just autoregressive greedy argmax.
1
6
Generic architecture with event-based output means that changing the task (e.g., to Onset only) is as simple as changing the target labels. No need for new output heads, loss design, or decoding procedures.
We show performance equivalent to architectures custom designed for AMT.
1
6
To handle Transformer memory limits, we split input spectrograms into ~4-second segments. Each is transcribed independently.
Surprisingly, the model rarely makes mistakes at predicting note-off events for notes where the note-on event happened in a different segment.
1
7
We're excited about the possibilities for creating new Music Information Retrieval models by focusing on dataset creation and labeling rather than custom model design.
1
8
The success of this seq2seq setup is also on the path toward using large model techniques from NLP for MIR. Imagine the GPT-3 of MIR!
Rodrigo Castellon and @chrisdonahuey also have some great recent work in this direction with CALM: arxiv.org/abs/2107.05677
11






