Excited to share our @ISMIR2021 paper: Sequence-to-Sequence Piano Transcription with Transformers Generic encoder-decoder Transformer + Spectrogram inputs + MIDI-like event outputs = SotA results! arxiv.org/abs/2107.09142 With @iansimon, @rigeljs, @ethanmanilow, and @jesseengel

Jul 21, 2021 · 12:07 AM UTC

5
33
132
We use the "small" encoder-decoder architecture from the T5 paper (@colinraffel et al.) and feed it spectrograms as continuous inputs. Outputs are MIDI-like events, not piano rolls: Note On/Off, Velocity, Time Shift, EOS Decoding is just autoregressive greedy argmax.
1
6
Generic architecture with event-based output means that changing the task (e.g., to Onset only) is as simple as changing the target labels. No need for new output heads, loss design, or decoding procedures. We show performance equivalent to architectures custom designed for AMT.
1
6
To handle Transformer memory limits, we split input spectrograms into ~4-second segments. Each is transcribed independently. Surprisingly, the model rarely makes mistakes at predicting note-off events for notes where the note-on event happened in a different segment.
1
7
We're excited about the possibilities for creating new Music Information Retrieval models by focusing on dataset creation and labeling rather than custom model design.
1
8
Sort replies: Relevant Recent Liked
Awesome work! Using absolute time offsets from the beginning of the segment is clever :)
1
4
Thanks! @iansimon gets the credit for that insight!
3
Is there a demo page for this? Thanks in advance.
1
Sorry, not yet, but we do plan to release code soon!
1
2
This looks very interesting! Congrats! It would be great to have this kind of model on TensorFlow Hub! 😉
1