1 of 23

Live Lecture 7.1

Introduction to Classification

Summer 2020

DATA 8

Spring 2020

2 of 23

Announcements

  • HW 12 released today, due on 08/06 at 11:59pm PDT
  • Project 2 released today, due on 08/09 at 11:59pm PDT
  • Last discussion tomorrow
  • Updated lab and discussion section schedule for week 7 available here
  • Week 8 Lectures:
    • Monday 10 am: Guest lecture by Ziad Obermeyer
    • Friday 1 pm: Data Science Career/Alumni Panel

3 of 23

Agenda

  1. Introduction to Classification
  2. Classifiers
  3. K Nearest Neighbors
  4. Evaluating Classifier Performance

4 of 23

Predicting an Observation’s Outcome

  • Two Types of Prediction:
    • Regression = Continuous numerical outcome
    • Classification = Categorical outcome

  • Given incomplete information, one way of making predictions:
  • Find others who are similar to that observation
  • Record their outcomes
  • Use those outcomes as a basis for prediction

5 of 23

Machine Learning Algorithm

  • A mathematical model
  • “Trained” on sample data
  • Makes predictions or decisions about the outcome of an observation

6 of 23

Classification Example: Spam Emails

7 of 23

Examples (cont.)

  • Hacking detection: Amazon, Paypal, LinkedIn, etc.
  • Medical alert systems: Early detection of sepsis, breast cancer tumor classification
  • Marketing: Targeted advertising on Facebook and Amazon
  • Insurance and banking: Identifying risky, would-be clients
  • Self-driving cars: computer vision
  • Etc.

8 of 23

Agenda

  • Introduction to Classification
  • Classifiers
  • K Nearest Neighbors
  • Evaluating Classifier Performance

9 of 23

Training a Classifier

Classifier

Variable(s) of an observation

Predicted label of the observation

Population

Labels

Sample

Training

Set

Test

Set

Model the association between attributes & labels

Estimate the accuracy of the classifier

10 of 23

Training a Classifier

Population

Attributes

Labels

Training

Set

Test

Set

Classifier

Variables of an observation

Predicted label of the observation

11 of 23

Agenda

  • Introduction to Classification
  • Classifiers
  • K Nearest Neighbors
  • Evaluating Classifier Performance

12 of 23

Sketch of K Nearest Neighbors

Variable 1

Variable 2

Legend:

Class 1:

Class 2:

13 of 23

Sketch of K Nearest Neighbors

Variable 1

Variable 2

Legend:

Class 1:

Class 2:

14 of 23

Sketch of K Nearest Neighbors

Variable 1

Variable 2

Legend:

Class 1:

Class 2:

?

15 of 23

Distance: Measure of Similarity

(x₀, y₀)

(x₁, y₁)

y₀ - y₁

x₀ - x₁

16 of 23

Distance Between Two Points

  • Two attributes x and y:
  • Three attributes x, y, and z:
  • and so on ...

17 of 23

Standardize if Necessary

Chronic Kidney Disease data set

  • If the attributes are on very different numerical scales, distance can be affected
  • In such a situation, it is a good idea to convert all the variables to standard units

18 of 23

Finding the K Nearest Neighbors (KNN)

To find the k nearest neighbors of a new observation:

  1. Find the distance between the observation and each example in the training set
  2. Augment the training data table with a column containing all the distances
  3. Sort the augmented table in increasing order of the distances
  4. Take the top k rows of the sorted table
  5. Take a majority vote of the k nearest neighbors

19 of 23

Sketch of K Nearest Neighbors

Variable 1

Variable 2

Legend:

Class 1:

Class 2:

20 of 23

Agenda

  • Introduction to Classification
  • Classifiers
  • K Nearest Neighbors
  • Evaluating Classifier Performance

21 of 23

Accuracy of a Classifier

The accuracy of a classifier on a labeled data set is the proportion of examples that are labeled correctly

Need to compare classifier predictions to true labels

If the labeled data set is sampled at random from a population, then we can infer accuracy on that population

Sample

Labels

Training

Set

Test

Set

Population

22 of 23

Start with a Representative Sample

  • Both the training and test sets must accurately represent the population on which you use your classifier
  • Overfitting happens when a classifier does very well on the training set, but can’t do as well on the test set
  • Finding K: Select the K which maximizes test set accuracy

23 of 23

Questions?