1 of 50

Test

  • First
  • Second
  • Third

1

2 of 50

UNSUPERVISED BARWISE MUSIC COMPRESSION FOR PATTERN UNCOVERING AND STRUCTURAL SEGMENTATION

Axel Marmoret, PhD Student at IRISA, Rennes, France

Centre for Digital Music, Queen Mary University, London

axel.marmoret@irisa.fr

3 of 50

Hi!

3

Rennes

PhD co-supervised:

Frédéric Bimbot

Jérémy Cohen

Verse

Chorus

Verse

Solo

Chorus

PhD subject, in short:

4 of 50

Unsupervised Barwise Music Compression for Pattern Uncovering and Structural Segmentation

4

Music

Barwise processing

UnsupervisedCompression

Structural Segmentation

Pattern uncovering

Feature representation

Audio signal

STFT, Mel Spectrogram, Chromagram, …

madmom toolbox

[Böck+ 2016]

[Böck+ 2016] Böck, S et al. Madmom: A new python audio and music signal processing library. In 2016 Proc. of the 24th ACM international conference on Multimedia.

5 of 50

Guideline of this talk

  1. Structural segmentation

  • Nonnegative Tucker Decomposition (NTD)

  • Barwise compression schemes

  • AutoEncoding NTD

5

UnsupervisedCompression

Structural Segmentation

Pattern uncovering

6 of 50

STRUCTURAL SEGMENTATION

6

7 of 50

Music structure

  • Segmenting a song in structural elements
    • Simplified representation of the organisation of a song

7

Verse

Chorus

Verse

Solo

Chorus

A

B

A

C

B’

a

Song organisation

High-scale: sectional level

Low-scale:

bar level

b

c

c

a

b

d

f

c

c’

e

4

5

4

5

1

2

3

3

1

2

3

3

4

5

4

9

8

8

6

7

Boundary retrieval

8 of 50

Why structural segmentation?

8

For structure analysis

As a mid-level feature

As a quantitative evaluation

  • Automatic summary
  • Music analysis/Musicology

  • Genre/music recognition
  • Recommendation
  • Cover detection

To quantitatively evaluate the relevance of a model, applied to music analysis

9 of 50

Structure as a compression scheme

  • Structure is generally retrieved via [Nieto+ 2020]:
    • Homogeneity
    • Novelty
    • Repetition
    • Regularity

9

Redundancy 🡺 compression schemes

[Nieto+ 2020] Nieto, O et al., Audio-based music structure analysis: Current trends, open challenges, and applications. 2020 Transactions of the International Society for Music Information Retrieval.

🡺 constraints can account for regularity

10 of 50

Barwise point of view

  • Hypothesis 1: patterns tend to develop at the bar scale

  • Hypothesis 2: boundaries fall on downbeats

  • Each bar contains exactly s frames ("subdivision" parameter)

10

Original

spectrogram

 

Barwise

processing

Barwise

spectrograms

 

 

 

 

11 of 50

  •  

11

Barwise autosimilarity

12 of 50

Convolutive "block-matching" (CBM) algorithm�[Marmoret+ 2022]

  • Idea: framing square blocks of high similarity in the diagonal

  • Each segment has a cost: similarity average

  • Optimal segmentation: maximizes the average similarity of each block (segments) by dynamic programming

  • Open-source code:

gitlab.inria.fr/amarmore/autosimilarity_segmentation

12

[Marmoret+ 2022] Marmoret, A., Cohen, J. E., & Bimbot, F. Barwise Compression Schemes for Audio-Based Music Structure Analysis. 2022 arXiv preprint arXiv:2202.04981.

Autosimilarity matrix

Bar indexes

Bar indexes

13 of 50

CBM Algorithm, in practice

  • Segmentation results:
    • Boundary: valid if close to annotation, with a tolerance, leading to Precision, Recall, Fmeasure
    • Results for state-of-the-art on RWC Pop [Goto+ 2002]

  • Objective: improve these results with compression-based similarity notions

13

[Goto+ 2002] Goto, M., Hashiguchi, H., Nishimura, T., & Oka, R. RWC Music Database: Popular, Classical and Jazz Music Databases. In 2002 ISMIR.

[Foote 2000] Foote, J. Automatic Audio Segmentation using a Measure of Audio Novelty. In 2000 Proc. IEEE International Conference on Multimedia and Expo.

[Nieto+ 2013] Nieto, O., & Jehan, T. Convex non-negative matrix factorization for automatic music structure identification. In 2013 IEEE ICASSP.

[McFee+ 2014] McFee, B., & Ellis, D. Analyzing Song Structure with Spectral Clustering. In 2014 ISMIR.

[Grill+ 2015] Grill, T., & Schlüter, J. Music Boundary Detection Using Neural Networks on Combined Features and Two-Level Annotations. In 2015 ISMIR.

14 of 50

NONNEGATIVE �TUCKER DECOMPOSITION

14

15 of 50

TFB tensor

  • Tensors: generalization of matrices to higher dimensions

  • Time-Frequency-Bar tensor: 3 dimensions
    • 3-rd order tensor representing barwise spectrograms

15

Original

spectrogram

Barwise spectrograms

 

 

 

 

 

16 of 50

Nonnegative Matrix Factorization example: Music Transcription [Smaragdis+ 2003]

16

Image taken from [Wu+ 2022] Wu, H., Marmoret, A., & Cohen, J. E. Semi-Supervised Convolutive NMF for Automatic Music Transcription. 2022, arXiv preprint arXiv:2202.04989.

[Smaragdis+ 2003] P. Smaragdis and J. Brown, “Non-negative matrix factorization for polyphonic music transcription,” in 2003 Proc. IEEE Workshop Applicat. Signal Process. Audio Acoust. (WASPAA)

17 of 50

Nonnegative Matrix Factorization example

17

 

 

 

 

 

18 of 50

Mixing templates

18

 

 

 

 

 

 

19 of 50

 

19

 

 

=

 

Hi-hat

Snare

Kick

20 of 50

Mixing patterns

20

 

 

 

 

 

 

 

21 of 50

Bars, expressed with patterns

21

 

 

 

 

 

 

 

22 of 50

NTD: Nonnegative Tucker Decomposition�[Kim+ 2007, Smith+ 2018, Marmoret+ 2020, Marmoret+ 2021]

22

X

 

 

 

G

 

 

 

 

 

 

 

 

 

 

 

 

 

 

[Kim+ 2007] Kim, Y. D., & Choi, S.Nonnegative tucker decomposition. In 2007 IEEE conference on computer vision and pattern recognition.

[Smith+ 2018] Smith, J. B., & Goto, M.. Nonnegative tensor factorization for source separation of loops in audio. In 2018 IEEE ICASSP.

[Marmoret+ 2020] Marmoret, A., Cohen, J., Bertin, N., & Bimbot, F. Uncovering Audio Patterns in Music with Nonnegative Tucker Decomposition for Structural Segmentation. In 2020 ISMIR.

[Marmoret+ 2021] Marmoret, A., et al.. Nonnegative Tucker Decomposition with Beta-divergence for Music Structure Analysis of audio signals 2021,. arXiv preprint arXiv:2110.14434.

 

23 of 50

NTD in practice: chromagram

23

G

 

 

G

 

 

X

Song: "Come Together" by The Beatles

24 of 50

 

  • Most bars appear as a sparse combination of musical patterns

24

25 of 50

 

25

 

26 of 50

Autosimilarities: feature vs NTD

26

27 of 50

Segmentation results

  • NTD segmentation results on chromagrams outperforms unsupervised state-of-the-art

27

28 of 50

 

28

Griffin-Lim

spectrogram

Audio signal

[Marmoret+ 2021] Marmoret, A., Voorwinden, F., Leplat, V., Cohen, J. E., & Bimbot, F. Nonnegative Tucker Decomposition with Beta-divergence for Music Structure Analysis of audio signals, 2021 arXiv preprint arXiv:2110.14434.

[Gillis 2020] Gillis, N., Nonnegative Matrix Factorization, Society for Industrial and Applied Mathematics, Philadelphia, PA, 2020.

[Févotte+ 2009] Févotte, C., Bertin, N., & Durrieu J.L., Nonnegative matrix factorization with the Itakura-Saito divergence: With application to music analysis, Neural computation 2009

STFT

Audio signal

 

Audio signal

STFT

Audio signal

spectrogram

NTD

Griffin-Lim

 

(One musical pattern)

29 of 50

BARWISE COMPRESSION�SCHEMES

29

30 of 50

Barwise compression

Barwise spectrograms

30

Q

NTD

PCA

Autoencoders

… and potentially more!

31 of 50

TFB tensor

31

Original

spectrogram

Barwise spectrograms

 

 

 

 

 

32 of 50

Barwise TF (Time-Frequency)

  • In tensor algebra:

unfolding

30

Original

spectrogram

Barwise spectrograms

 

 

 

 

 

Barwise TF

 

 

33 of 50

PCA

  •  

33

 

 

 

 

 

 

PCA projection

 

34 of 50

Single-Song AutoEncoders (SSAE) [Marmoret+ 2022]

  •  

34

 

 

Encoder

 

Decoder

 

 

 

 

[Marmoret+ 2022] Marmoret, A., Cohen, J. E., & Bimbot, F. Barwise Compression Schemes for Audio-Based Music Structure Analysis. 2022 arXiv preprint arXiv:2202.04981.

35 of 50

Segmentation results [Marmoret+ 2022]

  • SSAE outperforms supervised state-of-the-art for F3
  • For other techniques: results between blind and supervised methods

35

36 of 50

AUTOENCODING NTD

Work in progress!

36

37 of 50

NTD matricisation

  •  

37

 

Barwise TF matrix

Latent

representation

"Matricisation"

("unfolding" of the tensor)

2 matrix products

Kronecker product

 

 

 

 

 

 

 

 

 

 

 

 

2 fully-connected layers

38 of 50

NTD-AutoEncoder:�NTD as a 2-layer decoder

38

 

 

Encoder

(?)

 

 

 

Convolutions,

Fully-connected, …

Can be initialized with NTD results, at random… or in between

 

 

 

 

 

39 of 50

Why mixing both?

39

Performance-Interpretability

  • Combining:
    • segmentation results of AutoEncoders
    • interpretability of NTD

  • Objective: interpretable patterns in the decoder

Expressiveness

  • Greater subspace of solutions with AE than with NTD
    • activations functions
    • encoder design
  • Ease of implementing constraints (ex: sigmoid in the latent space)

  • Objective: improve the NTD outputs, and follow state-of-the-art research

Technical tools

  • Using toolboxes (PyTorch, TensorFlow) with high-performance computing

  • Objective: benefiting from recent developments in Neural networks computation

40 of 50

Preliminary results

  •  

40

Decoder initalisation:

Random

NTD

41 of 50

Summing up

41

Unsupervised Compression

Structural Segmentation

Pattern uncovering

… ?

Music

Barwise processing

Feature representation

Audio signal

STFT, Mel Spectrogram, Chromagram, …

madmom toolbox

[Böck+ 2016]

Compressed representations hold structural information

Interpretable outputs

  • NTD
  • PCA
  • AutoEncoders
  • NTD-AE

42 of 50

Open problems

  • Dimensioning the compression schemes
    • No obvious link between dimension and segmentation results
    • Particularly important for NTD: three hyperparameters

  • Using supervision and prior knowledge to improve compression
    • Transfer learning for AE, dictionary learning for NTD
    • Constraints on factors/layers adapted to representation

  • Results should be consolidated
    • Almost exclusively results on RWC Pop

42

43 of 50

References

  • [Böck+ 2016] Böck, S et al. Madmom: A new python audio and music signal processing library. In 2016 Proc. of the 24th ACM international conference on Multimedia.
  • [Févotte+ 2009] Févotte, C., Bertin, N., & Durrieu J.L., Nonnegative matrix factorization with the Itakura-Saito divergence: With application to music analysis, Neural computation 2009.
  • [Foote 2000] Foote, J. Automatic Audio Segmentation using a Measure of Audio Novelty. In 2000 Proc. IEEE International Conference on Multimedia and Expo.
  • [Gillis 2020] Gillis, N., Nonnegative Matrix Factorization, Society for Industrial and Applied Mathematics, Philadelphia, PA, 2020.
  • [Goto+ 2002] Goto, M., Hashiguchi, H., Nishimura, T., & Oka, R. RWC Music Database: Popular, Classical and Jazz Music Databases. In 2002 ISMIR.

43

44 of 50

References

  • [Grill+ 2015] Grill, T., & Schlüter, J. Music Boundary Detection Using Neural Networks on Combined Features and Two-Level Annotations. In 2015 ISMIR.
  • [Kim+ 2007] Kim, Y. D., & Choi, S.Nonnegative tucker decomposition. In 2007 IEEE conference on computer vision and pattern recognition.
  • [Marmoret+ 2020] Marmoret, A., Cohen, J., Bertin, N., & Bimbot, F. Uncovering Audio Patterns in Music with Nonnegative Tucker Decomposition for Structural Segmentation. In 2020 ISMIR.
  • [Marmoret+ 2021] Marmoret, A., Voorwinden, F., Leplat, V., Cohen, J. E., & Bimbot, F. Nonnegative Tucker Decomposition with Beta-divergence for Music Structure Analysis of audio signals. 2021 arXiv preprint arXiv:2110.14434.
  • [Marmoret+ 2022] Marmoret, A., Cohen, J. E., & Bimbot, F. Barwise Compression Schemes for Audio-Based Music Structure Analysis. 2022 arXiv preprint arXiv:2202.04981.

44

45 of 50

References

  • [McFee+ 2014] McFee, B., & Ellis, D. Analyzing Song Structure with Spectral Clustering. In 2014 ISMIR.
  • [Nieto+ 2013] Nieto, O., & Jehan, T. Convex non-negative matrix factorization for automatic music structure identification. In 2013 IEEE ICASSP.
  • [Nieto+ 2020] Nieto, O et al., (2020). Audio-based music structure analysis: Current trends, open challenges, and applications. Transactions of the International Society for Music Information Retrieval, 3(1).
  • [Smaragdis + 2003] P. Smaragdis and J. Brown, “Non-negative matrix factorization for polyphonic music transcription,” in 2003 Proc. IEEE Workshop Applicat. Signal Process. Audio Acoust. (WASPAA)
  • [Smith+ 2018] Smith, J. B., & Goto, M.. Nonnegative tensor factorization for source separation of loops in audio. In 2018 IEEE ICASSP.
  • [Wu+ 2022] Wu, H., Marmoret, A., & Cohen, J. E. Semi-Supervised Convolutive NMF for Automatic Music Transcription. 2022, arXiv preprint arXiv:2202.04989.

45

46 of 50

QUESTIONS?

46

47 of 50

Barwise point of view

  • Hypothesis 1: patterns tend to develop at the barscale

  • Hypothesis 2: boundaries fall on downbeats
    • Boundary: valid if close to annotation, with a tolerance, leading to Precision, Recall, F1
    • Results for some state-of-the-art on RWC Pop [Goto+ 2002] with madmom [Böck+ 2016]

47

[Goto+ 2002] Goto, M., Hashiguchi, H., Nishimura, T., & Oka, R. RWC Music Database: Popular, Classical and Jazz Music Databases. In 2002 ISMIR.

[Böck+ 2016] Böck, S et al. Madmom: A new python audio and music signal processing library. In 2016 Proc. of the 24th ACM international conference on Multimedia.

[Foote 2000] Foote, J. Automatic Audio Segmentation using a Measure of Audio Novelty. In 2000 Proc. IEEE International Conference on Multimedia and Expo.

[Nieto+ 2013] Nieto, O., & Jehan, T. Convex non-negative matrix factorization for automatic music structure identification. In 2013 IEEE ICASSP.

[McFee+ 2014] McFee, B., & Ellis, D. Analyzing Song Structure with Spectral Clustering. In 2014 ISMIR.

48 of 50

Segmentation results: NTD β-divergence�[Marmoret+ 2021]

  • Better segmentation scores than with chromagrams

48

[Marmoret+ 2021] Marmoret, A., Voorwinden, F., Leplat, V., Cohen, J. E., & Bimbot, F. Nonnegative Tucker Decomposition with Beta-divergence for Music Structure Analysis of audio signals. 2021 arXiv preprint arXiv:2110.14434.

49 of 50

SSAE Architecture

49

3

3

3x3 Convolutions

3x3 Convolutions

 

Transposed Convolution 3x3

Decoder

Encoder

3

3

ReLU + Max-pooling 2x2

ReLU + Fully-connected

Transposed Convolution 3x3

50 of 50

Kronecker product