1 of 94

Local

Music

Recommendation

Douglas Turnbull

Ithaca College

Ithaca, New York

USA

2 of 94

3 of 94

4 of 94

5 of 94

6 of 94

Part 1

Recommender Systems

7 of 94

8 of 94

9 of 94

Personalized Radio

10 of 94

Celestial Jukebox

11 of 94

Recommender System (RecSys)

A computational system that

recommends

specific items to a

specific user.

12 of 94

Recommender System (RecSys)

A computational system that

recommends

specific [artists, songs, events] to a

specific [listener, fan].

13 of 94

Content-based

(CB)

Analyze Substance of Items

Collaborative

Filtering

(CF)

Analyze User Interactions with Items

Hybrid

Combine CB and CF Approaches

14 of 94

Content-Based

(CB)

Analyze Substance of Items

  1. Human Annotation
  2. Social Tagging
  3. Text-Mining Reviews
  4. Computer Vision/Audition

15 of 94

Collaborative

Filtering

(CF)

Analyze User Interactions with Items

  • Explicit Feedback
  • Implicit Feedback

16 of 94

Cold Start” Problems

New Item has no user ratings

Solution: Content-based RecSys

New User has no interaction with the RS

Solutions: Social Login, UX Onboarding

17 of 94

CF Matrix Representation

ri,u = how much user u likes (or dislikes) item i

item i

user u

|U|x|I|

ru,i

18 of 94

Neighborhood Algorithms

  • Collect Joe’s favorite movies
  • Find users who also like the same movies
  • Recommend movies that they like to Joe

Two Ways to Compute

  • User-User Similarity
  • Item-Item Similarity

19 of 94

User-User Neighborhood

  1. Represent Joe as a column vector
  2. Calculated the similarity between Joe’s vector and all users
    • e.g., cosine distance
  3. Find k most similar “neighbors”
  4. Average rows from neighbors
  5. Rank items

X

|U|x|I|

|I|x1

|I|x1

=

|U|x1

|U|x|I|

X

20 of 94

Item-Item Neighborhood

  • Compute Item-Item Similarity Matrix
  • Recommend most similar items to items in the user model

XT

|I| x |U|

X

|U| x |I|

G

|I| x |I|

=

G

uT

|I|x1

User Model

21 of 94

Problems with using Raw Data Matrix X

Size: millions of items, millions of users

Sparse: inefficient storage + computation

Noisy: lots of meaningless junk

Synonymy: “hip hop” vs. “rap”

Polysemy: “romantic”

Matrix Factorization “fixes” these problems

22 of 94

Part 2 Matrix Factorization

23 of 94

24 of 94

25 of 94

Embedding Items and Users into Latent “Semantic” Space

26 of 94

Matrix Factorization

minU,Iratings (ru,i - xuT yi)2

k - selected dimension of embedded space

item i

user u

|U| x |I|

X

ru,i

xuT

|U| x k

U

k x |I|

yi

I

.

.

27 of 94

Matrix Factorization

minU,Iratings (ru,i - xuT yi)2

+ λ (||U||2 + ||I||2)

λ - regularization parameter

item i

user u

|U| x |I|

X

ru,i

xuT

|U| x k

U

k x |I|

yi

I

.

.

28 of 94

Matrix Factorization

minU,Iratings (ru,i - xuT yi)2

+ λ (||U||2 + ||I||2)

Alternating Least Squares (ALS)

  • Fix I → solve for U
  • Fix U → solve for I

29 of 94

Implicit Data

minU,Iu,i cu,i(ru,i - xuT yi)2

+ λ (||U||2 + ||I||2)

ri,u = 1 when user u has interacted with item i

0 otherwise

ci,u = confidence in rating ri,u

Ci,u= 50 when ri,u = 1

Ci,u= 1 when ri,u = 0

30 of 94

Explicit vs. Implicit Data Loss Function

Implicit Feedback

minU,Iu,i cu,i(ru,i - xuT yi)

+ λ (||U||2 + ||I||2)

Explicit Ratings

minU,Iratings (ru,i - xuT yi)

+ λ (||U||2 + ||I||2)

31 of 94

Part 3Long-Tail

Music Recommendation

32 of 94

33 of 94

34 of 94

35 of 94

36 of 94

Most Popular

Users - Playlists

Items - Tracks

Least Popular

Short

Head

Mid

Body

Long

Tail

minU,Iu,i cu,i(ru,i - xuT yi)2

37 of 94

Popularity Bias

Small minority of Short-Head have big impact on loss function.

minU,Iu,i cu,i(ru,i - xuT yi)

Item-Balanced Loss Function:

Scale cu,i so each track has same weight.

cu,i ← cu,i / ||c*,i||

38 of 94

Real Plot from Gabe

Million Playlist Dataset

1M Playlists

2.2M Tracks

Retain tracks with on at least 10 playlists

→ 370K tracks

One-Third Partitions:

Short-Head → 2K tracks

Mid-Body → 18K tracks

Long-Tail → 350K tracks

39 of 94

Training Playlists

Test Playlist

Mask

Project

Score

Rank

Segment

Rank

40 of 94

Item-Balanced Loss Function

AUC

Overall

Standard ALS

0.9817

Item-Balanced ALS

0.9846

Improvement

0.0029

Error Reduction

16%

41 of 94

Item-Balanced Loss Function

AUC

Overall

Short

Head

Mid

Body

Standard ALS

0.9817

0.9448

0.9621

Item-Balanced ALS

0.9846

0.9493

0.9643

Improvement

0.0029

0.0045

0.0021

Error Reduction

16%

8%

6%

42 of 94

Item-Balanced Loss Function

AUC

Overall

Short

Head

Mid

Body

Long

Tail

Standard ALS

0.9817

0.9448

0.9621

0.9471

Item-Balanced ALS

0.9846

0.9493

0.9643

0.9627

Improvement

0.0029

0.0045

0.0021

0.0157

Error Reduction

16%

8%

6%

30%

43 of 94

Long-tail Recommendation Takeaways

Item-Balanced Loss Function results in

  • big improvement for long-tail tracks
  • overall improvement

Neighborhood-based approaches can outperform Matrix Factorization for long-tail data

  • see recent publications

44 of 94

Part 3

Local

Music Recommendation

45 of 94

46 of 94

47 of 94

48 of 94

49 of 94

50 of 94

Kevin Kinsella - Grassroots 2013 – Mark Anbinder

51 of 94

52 of 94

53 of 94

54 of 94

55 of 94

56 of 94

57 of 94

58 of 94

59 of 94

60 of 94

61 of 94

Why MegsRadio Failed?

  • too many music streaming giants
    • YouTube, Spotify, Apple Music, Google Play

  • personalized Internet radio is dead
    • Long live celestial jukeboxes

  • our inability to maintain professional-grade service
    • Needed to be more responsive to users
    • Students graduate too quickly

  • playlist algorithm is not tuned correctly
    • too many WTF music recommendations

  • obtaining music is hard

62 of 94

63 of 94

64 of 94

65 of 94

66 of 94

67 of 94

68 of 94

69 of 94

70 of 94

71 of 94

72 of 94

73 of 94

74 of 94

Event �Information

1, Artist Information

2. Artist Similarities

3. User Preferences

Identify Local Artists

Event Recommendation

Playlist Generation

75 of 94

76 of 94

77 of 94

78 of 94

Northeastern Music Informatics Special Interest Group

Ithaca College

Februrary 2020

79 of 94

JimiLab @ Ithaca College:

Just Inspired Music Innovation

Douglas Turnbull

Ithaca College

dougturnbull.org

80 of 94

(Truncated) Singular Value Decomposition

k is a hyperparameter

Xk = best k-rank approximation of X with respect to |X - Xk|2 (Frobenius norm)

https://en.wikipedia.org/wiki/Latent_semantic_analysis

u

users

i

items

|I| x |U|

X

i’

|I| x k

Uk

σ1

σk

Σk

k x k

k x |U|

u’

VTk

≈ Xk =

81 of 94

Kevin Kinsella - Grassroots 2013 – Mark Anbinder

82 of 94

83 of 94

84 of 94

85 of 94

86 of 94

Popularity Bias - Algorithmic

  1. Popular Items have more Ratings
  2. More Ratings More Signal → More Confidence
  3. Structural Bias in Model

Add Matrix Visualization HereasdfAAAA

87 of 94

Event �Information

1M User Playlists

Identify Local Artists

Evaluate

Local Music Recommendation

88 of 94

Local Music Data

  • 8 Cities in the USA
  • ~36 Artists Per City
  • 99.999% Sparse

89 of 94

Local Music Experiment

6 RecSys Algorithms

  • Baselines: Random, Popularity
  • Neighborhood: ItemItem, UserUser
  • Matrix Factorization: ALS, BPS

3 Evaluation Metrics

  • NCDG, Precision@R, Precision@1

90 of 94

91 of 94

(Surprising) Results #1

Item-Item Neighborhood works better than Matrix Factorization

  • Track & Artist Recommendation
  • Perhaps due to extreme sparsity

92 of 94

Per-Track Loss Function

Since popular tracks have more ratings, the model designed to overfit them, while underfitting obscure tracks.

Solution: Alter the loss function so that each track has some influence.

93 of 94

Popularity Bias - Commercial

94 of 94