1 of 52

Unsupervised Learning

& Reinforcement Learning

CSIR Modelling and Digital Science

Nyalleng Moorosi

Borrowed from: Dr. Vukosi Marivate - Senior Data Scientist

2 of 52

Machine Learning

Labels? What Labels?

2

3 of 52

Unsupervised Learning

“the automatic discovery of regularities in data through the use of computer algorithms and with the use of these regularities to take actions such as classifying the data into different categories” [Bishop]

3

4 of 52

Supervised -> Unsupervised

We have input data x, we don't have have labels, what to do?

4

5 of 52

Why?

  • We can still find patterns with the data
    • Patterns could still be useful
    • Find regularity
    • Explain Data
    • Reduce Dimensions
    • Segment data
    • Feed into Supervised Learning
  • Finding anomalies/novelty

5

6 of 52

Motivations, Real World

Customer segmentation for better donation solicitation.

6

7 of 52

Motivations, Real World

2 weeks of Twitter data Data collected from Gauteng, Word2Vec + tSNE

7

8 of 52

Motivations, Real World

Closer look at some of the Word2Vec Groupings. How can we find these clusters?

8

9 of 52

When?

  • Most data is not labelled, so when dealing with some (most data) problems you will deal with this challenge

  • Exploratory Data analysis

Wikipedia

9

10 of 52

k-means Clustering

Given

We would like to partition data into k sets

Goal: Minimise Within-Cluster sum of squares

10

11 of 52

k-means Algorithm

Assign

Update

11

12 of 52

K-Means Algorithm

Wikipedia

12

13 of 52

K-Means Discussion

Pros:

  • Intuitive
  • Scalable

To note

  • Initialisation is important
  • Local minima
  • Choice of k
  • Geometry of clusters is set.
  • Normalisation before running.

Uses:

  • Anomaly Detection
  • Exploratory Analysis

Wikipedia

13

14 of 52

K-Means Choosing k

Within Cluster Sum of Squares

Silhouette Coefficient

14

15 of 52

K-means demo

http://bit.ly/phonepricedata

15

16 of 52

Expectation Maximisation (Gaussian Mixture Models)

Given

We would like to partition data into k sets

But, allow mixed memberships to each set/cluster, further each set will be made up of a multivariate normal distribution.

16

17 of 52

Expectation Maximisation

(Gaussian Mixture Models)

17

18 of 52

Expectation Maximisation

(Gaussian Mixture Models)

18

19 of 52

Expectation Maximisation (Gaussian Mixture Models)

EM Algorithm

E-Step:

  • Calculate the membership of each point to the components/cluster

M-Step:

  • Update model parameters given e-step

19

20 of 52

Expectation Maximisation (Gaussian Mixture Models)

E-Step:

M-Step:

Source: Wikipedia

20

21 of 52

Expectation Maximisation (Gaussian Mixture Models)

Results

Wikipedia

21

22 of 52

Expectation Maximisation (Gaussian Mixture Models): Discussion

Pros:

  • Geometry is flexible
  • Soft CLustering

To note:

  • Initialisation is important
  • Local minima
  • Choice of k
  • EM more computationally expensive.

Uses:

  • Anomaly Detection
  • Exploratory Analysis

sklearn

22

23 of 52

k-means soft clustering

Minimise

Where

Source: Wikipedia

23

24 of 52

More Clustering

Can we just use the similarity of points to each other to cluster?

Answer: YES

24

25 of 52

Hierarchal Clustering

Bottom Up Clustering

Need

  • Similarity Measure
  • Linkage Criteria

25

26 of 52

Hierarchal Clustering: �Similarity Measures

26

27 of 52

Hierarchal Clustering: �Linkage Criteria

How to merge?

Single Linkage

  • Smallest distance between items in two groups

Complete Linkage

  • Largest distance between items in two groups

27

28 of 52

Hierarchal Clustering: �Visualisation

Dendogram

Source: Vukosi Marivate and Nyalleng Moorosi SA job market analysis

28

29 of 52

Hierarchal Clustering: �Where to cut?

29

30 of 52

Hierarchal Clustering: �Discussion

Pros:

  • Intuitive
  • Scalable

To note

  • Distance measures impact clustering
  • Linkage criteria
  • Local minima
  • Where to stop

Uses:

  • Grouping

30

31 of 52

Practical: MNIST Digit Recognition

Hand-written digits, ignore labels.

31

32 of 52

Practical: The news problem

You are given a large number of documents, you would like to characterise what these documents are talking about, but labelling is expensive. Can we use unsupervised learning to find out the latent structure of these documents?

32

33 of 52

Advanced Topics

General

  • t-Distributed Stochastic Neighbor Embedding

Text

  • Latent Dirichlet Allocation
  • Word2Vec
  • Doc2Vec

Graphs

33

34 of 52

Advanced Topics

34

35 of 52

Advanced Topics

35

36 of 52

Reinforcement Learning

Striving to solve the AI problem, one reward at a time.

36

37 of 52

The Apartment Domain

Source: V Marivate: IMPROVED EMPIRICAL METHODS IN REINFORCEMENT-LEARNING EVALUATION

37

38 of 52

Reinforcement Learing

38

39 of 52

Reinforcement Learning

39

40 of 52

Reinforcement Learning

40

41 of 52

Value Iteration

Initialise all Values (V’s) to zero. Then iteratively do

Dynamic Programming

41

42 of 52

Online Paradigm

Q Learning

e-Greedy

42

43 of 52

Online Paradigm

Learn while in the world:

  • Access to real world/simulator
  • Actions have consequences
  • Learn and exploit as quickly as reasonable

Q Learning Algorithm:

Action choices:

  • e-Greedy
    • Best action , rest of time random

43

44 of 52

Offline Learning

Learn from collected data:

  • No access to real world/simulator
  • Actions have consequences when policy deployed
  • Infinite access to data
  • Bias from the collection policy

LSTDQ Learning Algorithm

44

45 of 52

Offline Learning

To Note:

  • Function Approximation
    • Not just tabular representation
    • Linear/Non-Linear
    • Deep Learning

  • Imperfect Information

  • Applications

45

46 of 52

RL + Unsupervised Learning???

Source: V Marivate: IMPROVED EMPIRICAL METHODS IN REINFORCEMENT-LEARNING EVALUATION

46

47 of 52

RL + Unsupervised Learning???

47

48 of 52

RL + Unsupervised Learning???

48

49 of 52

RL + Unsupervised Learning???

49

50 of 52

RL + Unsupervised Learning???

Source: V Marivate: IMPROVED EMPIRICAL METHODS IN REINFORCEMENT-LEARNING EVALUATION

50

51 of 52

Bibliography

  • Introduction to Machine Learning (Ethem Alpaydın)

  • Pattern Recognition (Christopher Bishop)

  • The Elements of Statistical Learning (Trevor Hastie et. al.)

  • Reinforcement Learning: An Introduction (Richard Sutton et. al.)

51

52 of 52

Always looking for:

Students, Collaborations, Shared Passions

so get in touch

@vukosi

http://www.vima.co.za

vmarivate@csir.co.za

52