Evaluation �and Credibility
How much should we believe in what was learned?
Lectures in DM
Academic year 2020/21
Jamolbek Mattiev
Outline
© Branko Kavšek, 2020/21
slide 2
12th lecture
Evaluation
Introduction
© Branko Kavšek, 2020/21
slide 3
12th lecture
Evaluation
Evaluation issues
© Branko Kavšek, 2020/21
slide 4
12th lecture
Evaluation
Classifier error rate
© Branko Kavšek, 2020/21
slide 5
12th lecture
Evaluation
Evaluation on “LARGE” data, 1
If many (thousands) of examples are available, including several hundred examples from each class, then how can we evaluate our classifier method?
© Branko Kavšek, 2020/21
slide 6
12th lecture
Evaluation
Evaluation on “LARGE” data, 2
© Branko Kavšek, 2020/21
slide 7
12th lecture
Evaluation
Classification Step 1: �Split data into train and test sets
© Branko Kavšek, 2020/21
slide 8
12th lecture
Results Known
+
+
-
-
+
THE PAST
Data
Training set
Testing set
Evaluation
Classification Step 2: �Build a model on a training set
© Branko Kavšek, 2020/21
slide 9
12th lecture
Training set
Results Known
+
+
-
-
+
THE PAST
Data
Model Builder
Testing set
Evaluation
Classification Step 3:�Evaluate on test set (Re-train?)
© Branko Kavšek, 2020/21
slide 10
12th lecture
Data
Predictions
Y
N
Results Known
Training set
Testing set
+
+
-
-
+
Model Builder
Evaluate
+
-
+
-
Evaluation
Unbalanced data
© Branko Kavšek, 2020/21
slide 11
12th lecture
Evaluation
Handling unbalanced data – how?
If we have two classes that are very unbalanced, then how can we evaluate our classifier method?
© Branko Kavšek, 2020/21
slide 12
12th lecture
Evaluation
Balancing unbalanced data, 1
© Branko Kavšek, 2020/21
slide 13
12th lecture
Evaluation
Balancing unbalanced data, 2
© Branko Kavšek, 2020/21
slide 14
12th lecture
Evaluation
A note on parameter tuning
© Branko Kavšek, 2020/21
slide 15
12th lecture
Evaluation
Making the most of the data
© Branko Kavšek, 2020/21
slide 16
12th lecture
Evaluation
Classification: Train, Validation, Test split
© Branko Kavšek, 2020/21
slide 17
12th lecture
Data
Predictions
Y
N
Results Known
Training set
Validation set
+
+
-
-
+
Model Builder
Evaluate
+
-
+
-
Final Model
Final Test Set
+
-
+
-
Final Evaluation
Model
Builder
Evaluation
*Predicting performance
© Branko Kavšek, 2020/21
slide 18
12th lecture
Evaluation
*Confidence intervals
© Branko Kavšek, 2020/21
slide 19
12th lecture
Evaluation
*Mean and variance (also Mod 7)
© Branko Kavšek, 2020/21
slide 20
12th lecture
Evaluation
*Confidence limits
© Branko Kavšek, 2020/21
slide 21
12th lecture
Pr[X ≥ z] | z |
0.1% | 3.09 |
0.5% | 2.58 |
1% | 2.33 |
5% | 1.65 |
10% | 1.28 |
20% | 0.84 |
40% | 0.25 |
–1 0 1 1.65
Evaluation
*Transforming f
© Branko Kavšek, 2020/21
slide 22
12th lecture
Evaluation
*Examples
(should be taken with a grain of salt)
© Branko Kavšek, 2020/21
slide 23
12th lecture
Evaluation
Evaluation on “small” data, 1
© Branko Kavšek, 2020/21
slide 24
12th lecture
Evaluation
Evaluation on “small” data, 2
© Branko Kavšek, 2020/21
slide 25
12th lecture
Evaluation
Repeated holdout method, 1
© Branko Kavšek, 2020/21
slide 26
12th lecture
Evaluation
Repeated holdout method, 2
© Branko Kavšek, 2020/21
slide 27
12th lecture
Evaluation
Cross-validation
© Branko Kavšek, 2020/21
slide 28
12th lecture
Evaluation
© Branko Kavšek, 2020/21
slide 29
12th lecture
Cross-validation example:
Evaluation
More on cross-validation
© Branko Kavšek, 2020/21
slide 30
12th lecture
Evaluation
Leave-One-Out cross-validation
© Branko Kavšek, 2020/21
slide 31
12th lecture
Evaluation
Leave-One-Out-CV and stratification
© Branko Kavšek, 2020/21
slide 32
12th lecture
Evaluation
*The bootstrap
© Branko Kavšek, 2020/21
slide 33
12th lecture
Evaluation
*The 0.632 bootstrap
© Branko Kavšek, 2020/21
slide 34
12th lecture
Evaluation
*Estimating error�with the bootstrap
© Branko Kavšek, 2020/21
slide 35
12th lecture
Evaluation
*More on the bootstrap
© Branko Kavšek, 2020/21
slide 36
12th lecture
Evaluation
Comparing data mining schemes
© Branko Kavšek, 2020/21
slide 37
12th lecture
Evaluation
Significance tests
© Branko Kavšek, 2020/21
slide 38
12th lecture
Evaluation
*Paired t-test
© Branko Kavšek, 2020/21
slide 39
12th lecture
William Gosset
Born: 1876 in Canterbury; Died: 1937 in Beaconsfield, England.�Obtained a post as a chemist in the Guinness brewery in Dublin in 1899. Invented the t-test to handle small samples for quality control in brewing. Wrote under the name "Student".
Evaluation
*Distribution of the means
© Branko Kavšek, 2020/21
slide 40
12th lecture
Evaluation
*Student’s distribution
© Branko Kavšek, 2020/21
slide 41
12th lecture
Pr[X ≥ z] | z |
0.1% | 4.30 |
0.5% | 3.25 |
1% | 2.82 |
5% | 1.83 |
10% | 1.38 |
20% | 0.88 |
Pr[X ≥ z] | z |
0.1% | 3.09 |
0.5% | 2.58 |
1% | 2.33 |
5% | 1.65 |
10% | 1.28 |
20% | 0.84 |
9 degrees of freedom normal distribution
Evaluation
*Distribution of the differences
© Branko Kavšek, 2020/21
slide 42
12th lecture
Evaluation
*Performing the test
© Branko Kavšek, 2020/21
slide 43
12th lecture
Evaluation
Unpaired observations
© Branko Kavšek, 2020/21
slide 44
12th lecture
Evaluation
*Interpreting the result
© Branko Kavšek, 2020/21
slide 45
12th lecture
Evaluation
T-statistic many uses
© Branko Kavšek, 2020/21
slide 46
12th lecture
Evaluation
*Predicting probabilities
© Branko Kavšek, 2020/21
slide 47
12th lecture
Evaluation
*Quadratic loss function
© Branko Kavšek, 2020/21
slide 48
12th lecture
Evaluation
*Informational loss function
© Branko Kavšek, 2020/21
slide 49
12th lecture
Evaluation
*Discussion
© Branko Kavšek, 2020/21
slide 50
12th lecture
Evaluation
Evaluation Summary:
© Branko Kavšek, 2020/21
slide 51
12th lecture
Evaluation