1 of 75

Predicting Health Outcomes

Correlation Matrices

2 of 75

What are the indicators of healthy hearts?

Respond to questions here: https://bit.ly/CA-Pre-Assessment

3 of 75

How is machine learning different from statistics?

4 of 75

How are machine learning and artificial intelligence similar to each other? How are they different?

5 of 75

What does Machine Learning do?

6 of 75

In what everyday scenarios can machine learning be used?

7 of 75

Streaming Platforms

8 of 75

Marketing and Advertising

9 of 75

Transportation

10 of 75

Banking and Finance

11 of 75

Historical Research

12 of 75

Music

13 of 75

Natural Language Processing

14 of 75

Biology Research

15 of 75

Medicine

16 of 75

“Cardiovascular disease - like heart attacks and stroke - may seem like an elderly disease. But studies have shown that cardiovascular events are happening more and more frequently in people as young as 20.”

17 of 75

P100 wellness study correlations

Framingham dataset correlations

18 of 75

Correlation or Causation?

Liam collected data on the sales of ice cream cones and air conditioners in his hometown. He found that when ice cream sales were low, air conditioner sales tended to be low and that when ice cream sales were high, air conditioner sales tended to be high.

19 of 75

Correlation or Causation?

Correlation: there is a relationship or pattern between the values of two variables.

Example: Liam found that when ice cream sales were low, air conditioner sales tended to be low and that when ice cream sales were high, air conditioner sales tended to be high.

Causation: one event causes another event to occur. Causation can only be determined from an appropriately designed experiment. In such experiments, similar groups receive different treatments, and the outcomes of each group are studied. We can only conclude that a treatment causes an effect if the groups have noticeably different outcomes

Example: Liam did another study and found that when ice cream sales were low, temperature was low and that when temperatures were high, ice cream sales were high, and air conditioner sales tended to be high.

20 of 75

3 types of correlation:

Positive correlation

As x increases, y tends to increase.

Negative correlation

As x increases, y tends to decrease.

No correlation

As x increases, y tends to stay the same or have no clear pattern

21 of 75

By the end of this activity, you’ll create something like this!

22 of 75

Feature vs. Variable

Definition:

23 of 75

The Framingham Heart Study Background

24 of 75

How can machine learning be used to accurately predict a patient’s risk of developing CVD?

25 of 75

Address “null variables” and “outliers”

Classify variables

Select features

Visualize your data

Calculate correlation

Correlation Roadmap

1

2

3

4

5

26 of 75

Classify variables

1

27 of 75

Classify variables

1

Your turn:

Classify the variables in your

Framingham dataset

Systems Medicine, 2019

28 of 75

Address “null variables” and “outliers”

Classify variables

Select features

Visualize your data

Calculate correlation

Correlation Roadmap

2

3

4

5

29 of 75

Address “null variables”

2

30 of 75

Address “null variables”

2

31 of 75

Address “null variables”

Your turn:

Find and address null values in your dataset.

2

Systems Medicine, 2019

32 of 75

Address “outliers”

2

33 of 75

Address “outliers”

2

34 of 75

Address “outliers”

2

35 of 75

Address “outliers”

2

Continuous variables only!

Input this formula into D:19

=OR(MAX(D2:D11)>D17, MIN(D2:D11)<D15)

36 of 75

Address “outliers”

Your turn:

Find and address outliers in your dataset.

2

Systems Medicine, 2019

37 of 75

Address “null variables” and “outliers”

Classify variables

Select features

Visualize your data

Calculate correlation

Correlation Roadmap

3

4

5

38 of 75

Calculate correlation

Positive correlation

As x increases, y tends to increase.

Negative correlation

As x increases, y tends to decrease.

No correlation

As x increases, y tends to stay the same or have no clear pattern

Pearson’s correlation coefficient closer to 1

Pearson’s correlation coefficient closer to -1

Pearson’s correlation coefficient closer to 0

3

39 of 75

Calculate correlation

3

40 of 75

Pearson’s Correlation formula

Calculate correlation

3

  • Pearson’s finds the coefficient “r” for each point - a value between 1 and -1. This determines the relationship of each point to the mean point of the data set. In practice the greater the codevience the great the correlation coefficient.

  • Answers the question: are the points in a close or a distant relationship? How scattered or distant (deviated) are these points �of data?

 

Sum of squared deviations (x)

Sum of squared deviations (y)

Codeviance

41 of 75

Calculate correlation

What is the purpose of the Pearson’s correlation formula? It tells us how closely our data aligns with the average x and y values, or how well x and y are correlated.

How does it mathematically achieve that purpose? We add all the differences from the average values of x and y in a standardized way.

3

 

42 of 75

43 of 75

44 of 75

Pearson Correlation

3

45 of 75

r = +1.0

r = +0.7

r = 0

r = -0.46

46 of 75

Calculate correlation

Your turn:

Create two scatter plots with your dataset. What do you notice?

3

Systems Medicine, 2019

47 of 75

Calculate correlation

3

48 of 75

Calculate correlation

Your turn:

Calculate correlation using Pearson’s Correlation Coefficient with your dataset.

3

Systems Medicine, 2019

49 of 75

Address “null variables” and “outliers”

Classify variables

Select features

Visualize your data

Calculate correlation

Correlation Roadmap

4

5

50 of 75

Visualize your data

4

51 of 75

Visualize your data

Your turn:

Create a heat map of your data correlation matrices.

4

Systems Medicine, 2019

52 of 75

Visualize your data

4

53 of 75

Bias, Errors, and/or Pitfalls

54 of 75

We’ve found a correlation that fits the above data… but will it hold up for a larger dataset like the one on the right?

55 of 75

Let’s apply what we’ve learned to an even larger dataset!

56 of 75

Your turn:

Create a correlation matrix using a larger dataset. �Then visualize the data.

Systems Medicine, 2019

57 of 75

Calculate correlation

3

Our outcome of CVD

58 of 75

So what?

59 of 75

Address “null variables” and “outliers”

Classify variables

Select features

Visualize your data

Calculate correlation

Correlation Roadmap

5

60 of 75

Select features

5

61 of 75

Select features

5

62 of 75

Select features

Orange sections: A features contribution in explaining the outcome

Red section: Mutual information provided by both features.

5

63 of 75

Select features

Weight

Age

BMI

Model

Weight

Age

5

64 of 75

Select features

5

65 of 75

Select features

5

66 of 75

Select features

5

67 of 75

Select features

5

68 of 75

Select features

Your turn:

Write down features correlated with each other (feature to feature) and features that are not correlated with CVD (feature to outcome).

5

Systems Medicine, 2019

69 of 75

Select features

5

  1. Using the list of Irrelavent and redundant features- Go to the “2c Feature Selection” tab.

  • Here you can delete the values from the features on your list, by deleting the values from both its column and row.

  • Replace Null cells with a “ -”

70 of 75

Select features

Your turn:

Remove variables correlated with one another. And create a Circos plot.

5

Systems Medicine, 2019

71 of 75

Circos Plot Visualization

72 of 75

What are the indicators of healthy hearts?

73 of 75

Address “null variables” and “outliers”

Classify variables

Select features

Visualize your data

Calculate correlation

Correlation Roadmap

74 of 75

Logistic Regression!

Next step…

75 of 75

What are the indicators of healthy hearts?

Respond to questions here: https://bit.ly/CAPostAssessment