1 of 14

Machine Learning:

Random Forest Classification

2 of 14

4 Key Pieces

  1. Data split

  • Train models

  • Models make Predictions

  • Majority Wins

2

3 of 14

1. Split the Data

Sampling with Replacement

4 of 14

Split the Data

  • The full dataset gets split into a number of smaller datasets

  • Each of these smaller datasets have small differences, making them more diverse

  • Sampling is done WITH replacement so some data points may be repeated/left out

4

5 of 14

2. Train Models

Train models on their respective datasets

6 of 14

Train Models

  • Each sample will have a decision tree trained on it

  • Trees are trained in parallel (independent of one another)

  • Vocab:
    • Branch (the lines indicating splits)
    • Leaf node is a node with no branches

6

7 of 14

3. Models Make Predictions

Each model makes its own prediction from its training

8 of 14

Models Make Predictions

  • Each model based on its own training on its own subset of the dataset makes a prediction

  • These predictions are a classification

  • Different trees may make different class label predictions

8

9 of 14

4. Final Prediction

Majority Wins

10 of 14

Majority Wins

  • The prediction of each model is factored in

  • The class label that is most predicted is the final output of the model

  • The model output is what the model predicts is the classification (not guaranteed to be the actual class label)

10

11 of 14

Practice Tree

11

12 of 14

12

13 of 14

13

Three different trees – take majority vote

14 of 14

Now…

Let’s Dance!!!

14