Test
1
UNSUPERVISED BARWISE MUSIC COMPRESSION FOR PATTERN UNCOVERING AND STRUCTURAL SEGMENTATION
Axel Marmoret, PhD Student at IRISA, Rennes, France
Centre for Digital Music, Queen Mary University, London
axel.marmoret@irisa.fr
Hi!
3
Rennes
PhD co-supervised:
Frédéric Bimbot
Jérémy Cohen
Verse
Chorus
Verse
Solo
Chorus
PhD subject, in short:
Unsupervised Barwise Music Compression for Pattern Uncovering and Structural Segmentation
4
Music
Barwise processing
UnsupervisedCompression
Structural Segmentation
Pattern uncovering
Feature representation
Audio signal
STFT, Mel Spectrogram, Chromagram, …
madmom toolbox
[Böck+ 2016]
[Böck+ 2016] Böck, S et al. Madmom: A new python audio and music signal processing library. In 2016 Proc. of the 24th ACM international conference on Multimedia.
Guideline of this talk
5
UnsupervisedCompression
Structural Segmentation
Pattern uncovering
STRUCTURAL SEGMENTATION
6
Music structure
7
Verse
Chorus
Verse
Solo
Chorus
A
B
A
C
B’
a
Song organisation
High-scale: sectional level
Low-scale:
bar level
b
c
c
a
b
d
f
c
c’
e
4
5
4
5
1
2
3
3
1
2
3
3
4
5
4
9
8
8
6
7
Boundary retrieval
Why structural segmentation?
8
For structure analysis
As a mid-level feature
As a quantitative evaluation
To quantitatively evaluate the relevance of a model, applied to music analysis
Structure as a compression scheme
9
Redundancy 🡺 compression schemes
[Nieto+ 2020] Nieto, O et al., Audio-based music structure analysis: Current trends, open challenges, and applications. 2020 Transactions of the International Society for Music Information Retrieval.
🡺 constraints can account for regularity
Barwise point of view
10
Original
spectrogram
Barwise
processing
Barwise
spectrograms
11
Barwise autosimilarity
Convolutive "block-matching" (CBM) algorithm�[Marmoret+ 2022]
gitlab.inria.fr/amarmore/autosimilarity_segmentation
12
[Marmoret+ 2022] Marmoret, A., Cohen, J. E., & Bimbot, F. Barwise Compression Schemes for Audio-Based Music Structure Analysis. 2022 arXiv preprint arXiv:2202.04981.
Autosimilarity matrix
Bar indexes
Bar indexes
CBM Algorithm, in practice
13
[Goto+ 2002] Goto, M., Hashiguchi, H., Nishimura, T., & Oka, R. RWC Music Database: Popular, Classical and Jazz Music Databases. In 2002 ISMIR.
[Foote 2000] Foote, J. Automatic Audio Segmentation using a Measure of Audio Novelty. In 2000 Proc. IEEE International Conference on Multimedia and Expo.
[Nieto+ 2013] Nieto, O., & Jehan, T. Convex non-negative matrix factorization for automatic music structure identification. In 2013 IEEE ICASSP.
[McFee+ 2014] McFee, B., & Ellis, D. Analyzing Song Structure with Spectral Clustering. In 2014 ISMIR.
[Grill+ 2015] Grill, T., & Schlüter, J. Music Boundary Detection Using Neural Networks on Combined Features and Two-Level Annotations. In 2015 ISMIR.
NONNEGATIVE �TUCKER DECOMPOSITION
14
TFB tensor
15
Original
spectrogram
Barwise spectrograms
Nonnegative Matrix Factorization example: Music Transcription [Smaragdis+ 2003]
16
Image taken from [Wu+ 2022] Wu, H., Marmoret, A., & Cohen, J. E. Semi-Supervised Convolutive NMF for Automatic Music Transcription. 2022, arXiv preprint arXiv:2202.04989.
[Smaragdis+ 2003] P. Smaragdis and J. Brown, “Non-negative matrix factorization for polyphonic music transcription,” in 2003 Proc. IEEE Workshop Applicat. Signal Process. Audio Acoust. (WASPAA)
Nonnegative Matrix Factorization example
17
Mixing templates
18
19
♩
♪
♩
♩
♩
♩
♩
♩
♩
♩
♩
=
Hi-hat
Snare
Kick
Mixing patterns
20
Bars, expressed with patterns
21
NTD: Nonnegative Tucker Decomposition�[Kim+ 2007, Smith+ 2018, Marmoret+ 2020, Marmoret+ 2021]
22
X
G
[Kim+ 2007] Kim, Y. D., & Choi, S.Nonnegative tucker decomposition. In 2007 IEEE conference on computer vision and pattern recognition.
[Smith+ 2018] Smith, J. B., & Goto, M.. Nonnegative tensor factorization for source separation of loops in audio. In 2018 IEEE ICASSP.
[Marmoret+ 2020] Marmoret, A., Cohen, J., Bertin, N., & Bimbot, F. Uncovering Audio Patterns in Music with Nonnegative Tucker Decomposition for Structural Segmentation. In 2020 ISMIR.
[Marmoret+ 2021] Marmoret, A., et al.. Nonnegative Tucker Decomposition with Beta-divergence for Music Structure Analysis of audio signals 2021,. arXiv preprint arXiv:2110.14434.
NTD in practice: chromagram
23
G
G
X
Song: "Come Together" by The Beatles
24
25
Autosimilarities: feature vs NTD
26
Segmentation results
27
28
Griffin-Lim
spectrogram
Audio signal
[Marmoret+ 2021] Marmoret, A., Voorwinden, F., Leplat, V., Cohen, J. E., & Bimbot, F. Nonnegative Tucker Decomposition with Beta-divergence for Music Structure Analysis of audio signals, 2021 arXiv preprint arXiv:2110.14434.
[Gillis 2020] Gillis, N., Nonnegative Matrix Factorization, Society for Industrial and Applied Mathematics, Philadelphia, PA, 2020.
[Févotte+ 2009] Févotte, C., Bertin, N., & Durrieu J.L., Nonnegative matrix factorization with the Itakura-Saito divergence: With application to music analysis, Neural computation 2009
STFT
Audio signal
Audio signal
STFT
Audio signal
spectrogram
NTD
Griffin-Lim
(One musical pattern)
BARWISE COMPRESSION�SCHEMES
29
Barwise compression
Barwise spectrograms
30
Q
NTD
PCA
Autoencoders
… and potentially more!
TFB tensor
31
Original
spectrogram
Barwise spectrograms
Barwise TF (Time-Frequency)
unfolding
30
Original
spectrogram
Barwise spectrograms
Barwise TF
PCA
33
PCA projection
Single-Song AutoEncoders (SSAE) [Marmoret+ 2022]
34
Encoder
Decoder
[Marmoret+ 2022] Marmoret, A., Cohen, J. E., & Bimbot, F. Barwise Compression Schemes for Audio-Based Music Structure Analysis. 2022 arXiv preprint arXiv:2202.04981.
Segmentation results [Marmoret+ 2022]
35
AUTOENCODING NTD
Work in progress!
36
NTD matricisation
37
Barwise TF matrix
Latent
representation
"Matricisation"
("unfolding" of the tensor)
2 matrix products
Kronecker product
2 fully-connected layers
NTD-AutoEncoder:�NTD as a 2-layer decoder
38
Encoder
(?)
Convolutions,
Fully-connected, …
Can be initialized with NTD results, at random… or in between
Why mixing both?
39
Performance-Interpretability
Expressiveness
Technical tools
Preliminary results
40
Decoder initalisation: | Random | NTD |
| | |
Summing up
41
Unsupervised Compression
Structural Segmentation
Pattern uncovering
… ?
Music
Barwise processing
Feature representation
Audio signal
STFT, Mel Spectrogram, Chromagram, …
madmom toolbox
[Böck+ 2016]
Compressed representations hold structural information
Interpretable outputs
Open problems
42
References
43
References
44
References
45
QUESTIONS?
46
Barwise point of view
47
[Goto+ 2002] Goto, M., Hashiguchi, H., Nishimura, T., & Oka, R. RWC Music Database: Popular, Classical and Jazz Music Databases. In 2002 ISMIR.
[Böck+ 2016] Böck, S et al. Madmom: A new python audio and music signal processing library. In 2016 Proc. of the 24th ACM international conference on Multimedia.
[Foote 2000] Foote, J. Automatic Audio Segmentation using a Measure of Audio Novelty. In 2000 Proc. IEEE International Conference on Multimedia and Expo.
[Nieto+ 2013] Nieto, O., & Jehan, T. Convex non-negative matrix factorization for automatic music structure identification. In 2013 IEEE ICASSP.
[McFee+ 2014] McFee, B., & Ellis, D. Analyzing Song Structure with Spectral Clustering. In 2014 ISMIR.
Segmentation results: NTD β-divergence�[Marmoret+ 2021]�
48
[Marmoret+ 2021] Marmoret, A., Voorwinden, F., Leplat, V., Cohen, J. E., & Bimbot, F. Nonnegative Tucker Decomposition with Beta-divergence for Music Structure Analysis of audio signals. 2021 arXiv preprint arXiv:2110.14434.
SSAE Architecture
49
…
3
3
3x3 Convolutions
3x3 Convolutions
Transposed Convolution 3x3
Decoder
Encoder
3
3
ReLU + Max-pooling 2x2
ReLU + Fully-connected
Transposed Convolution 3x3
Kronecker product