1 of 204

Regression Analyses

Dr. Debasis Samanta

Associate Professor

Department of Computer Science & Engineering

2 of 204

This presentation includes…

  • Introduction

  • Correlation Analysis

  • Regression Analysis
    • Linear Regression
    • Non-Linear Regression
    • Auto-Regression

DSamanta@IIT Kharagpur

2

Regression Analysis

3 of 204

Introduction

DSamanta@IIT Kharagpur

3

Regression Analysis

4 of 204

Hypothesis Testing Strategies

DSamanta@IIT Kharagpur

4

Regression Analysis

Types of tests of hypotheses

Also called

distribution-free test of hypotheses

Non-parametric tests

Also called

standard test of hypotheses

Parametric tests

5 of 204

Parametric Tests : Assumptions

DSamanta@IIT Kharagpur

5

Regression Analysis

  • Observation come from a normal population

  • Sample size is small

  • Population parameters like mean, variance, etc. hold good.

  • Requires measurement dealing with interval scaled data.

Usually assume certain properties of the population from which we draw samples.

6 of 204

Hypothesis Testing: Non-Parametric Test

DSamanta@IIT Kharagpur

6

Regression Analysis

  • Does not come under any assumption.

  • Assumes nominal, ordinal as well as interval scale data.

Non-parametric tests

Note:

Non-parametric tests need entire population (or sample of very large size)

7 of 204

Data for Relationship Analysis

DSamanta@IIT Kharagpur

7

Regression Analysis

Example:

Univariate population: The population consisting of only one variable.

Example:

Here, statistical measures suffice to find a relationship.

Bivariate population: Here, the data happen to be with two variables.

.

8 of 204

Data for Relationship Analysis

DSamanta@IIT Kharagpur

8

Regression Analysis

Multivariate population: If the data happen to be one more than two variable.

.

Example:

? If we add another variable say viscosity in addition to Pressure, Volume or Temperature?

9 of 204

Measures of Relationship

DSamanta@IIT Kharagpur

9

Regression Analysis

Q1: Does there exist relation between two variables (in case of bivariate population) ?

    • If yes, of what degree?

Q2: Is there any relationship between one variable in one side and two or more variables on the other side (in case of multivariate population)?

    • If yes, of what degree and in which direction?

In case of bivariate and multivariate populations, usually, we have to answer two types of questions:

10 of 204

Measures of Relationship

DSamanta@IIT Kharagpur

10

Regression Analysis

Q1: Does there exist relation between two variables (in case of bivariate population) ?

Q2: Is there any relationship between one variable in one side and two or more variables on the other side (in case of multivariate population)?

In case of bivariate and multivariate populations, usually, we have to answer two types of questions:

Solution

?

11 of 204

Measures of Relationship

DSamanta@IIT Kharagpur

11

Regression Analysis

Q1: Does there exist relation between two variables (in case of bivariate population) ?

Q2: Is there any relationship between one variable in one side and two or more variables on the other side (in case of multivariate population)?

To find solutions to the above questions, two approaches are known.

Correlation Analysis

Regression Analysis

In case of bivariate and multivariate populations, usually, we have to answer two types of questions:

12 of 204

Correlation Analyses

DSamanta@IIT Kharagpur

12

Regression Analysis

13 of 204

Correlation Analysis

DSamanta@IIT Kharagpur

13

Regression Analysis

Example: Weight is correlated with height

In statistics, the word correlation is used to denote some form of association between two variables.

14 of 204

Correlation Analysis

DSamanta@IIT Kharagpur

14

Regression Analysis

Do you find any correlation between X and Y as shown in the table?

Note

Look at the table given below

  • In statistical learning, a correlation analysis makes sense only when a relationship makes sense.

15 of 204

Correlation Analysis

DSamanta@IIT Kharagpur

15

Regression Analysis

When the values of attribute A varies at random with B and vice-versa.

Correlation

If the value of the attribute A increases with the increase in the value of the attribute B and vice-versa.

If the value of the attribute A decreases with the increase in the value of the attribute B and vice-versa.

Zero correlation

Positive correlation

Negative correlation

16 of 204

Correlation Analysis

DSamanta@IIT Kharagpur

16

Regression Analysis

Positive correlation

Negative correlation

Zero correlation

17 of 204

Form of Correlation

DSamanta@IIT Kharagpur

17

Regression Analysis

A correlation is linear when two variables

change at constant rate.

Concerning the form of a correlation, it could be linear, non-linear, or monotonic.

Linear Correlation

In this case, the relationship between

the variables graph as a curved pattern

parabola, hyperbola … etc).

Non-linear Correlation

18 of 204

Form of Correlation

DSamanta@IIT Kharagpur

18

Regression Analysis

Concerning the form of a correlation , it could be linear, non-linear, or monotonic.

Monotonicity of a function

19 of 204

Form of Correlation

DSamanta@IIT Kharagpur

19

Regression Analysis

Monotonic correlation: In a monotonic relationship, the variables tend to move in the same relative direction or opposite direction, but not necessarily at a constant rate.

Concerning the form of a correlation , it could be linear, non-linear, or monotonic :

Monotonic and non-monotonic relations

20 of 204

Correlation Analysis

DSamanta@IIT Kharagpur

20

Regression Analysis

We need to measure the degree of correlation between two attributes.

Exam score

21 of 204

Correlation Coefficient

  •  

DSamanta@IIT Kharagpur

21

Regression Analysis

22 of 204

Correlation Coefficient

DSamanta@IIT Kharagpur

22

Regression Analysis

23 of 204

Correlation Coefficient

DSamanta@IIT Kharagpur

23

Regression Analysis

24 of 204

Correlation Coefficient

DSamanta@IIT Kharagpur

24

Regression Analysis

25 of 204

Measuring Correlation Coefficients

DSamanta@IIT Kharagpur

25

Regression Analysis

Karl Pearson’s coefficient

Find correlation coefficient between two numerical attributes

Charles Spearman’s coefficient

Find correlation coefficient between two ordinal attributes

Chi-square coefficient of correlation

Find correlation coefficient between two nominal attributes

Three methods to measure the correlation coefficients

26 of 204

DSamanta@IIT Kharagpur

26

Regression Analysis

Pearson’s Correlation Analysis

27 of 204

Karl Pearson’s Correlation Analysis

DSamanta@IIT Kharagpur

27

Regression Analysis

Monalisa Sarma

IIT KHARAGPUR

This is also called Pearson’s Product Moment Correlation

 

Definition : Karl Pearson’s correlation coefficient

28 of 204

Karl Pearson’s Coefficient of Correlation

DSamanta@IIT Kharagpur

28

Regression Analysis

A small study is conducted involving 17 infants to investigate the association between gestational age at birth, measured in weeks, and birth weight, measured in grams.

Example : Correlation of Gestational Age and Birth Weight

29 of 204

Karl Pearson’s coefficient of Correlation

DSamanta@IIT Kharagpur

29

Regression Analysis

A small study is conducted involving 17 infants to investigate the association between gestational age at birth, measured in weeks, and birth weight, measured in grams.

Example : Correlation of Gestational Age and Birth Weight

 

30 of 204

Karl Pearson’s coefficient of Correlation

DSamanta@IIT Kharagpur

30

Regression Analysis

 

For the given data

Conclusion: The sample’s correlation coefficient indicates a strong positive correlation between Gestational Age and Birth Weight.

31 of 204

Significance Test

DSamanta@IIT Kharagpur

31

Regression Analysis

 

Definition : Karl Pearson’s correlation coefficient

32 of 204

Karl Pearson’s Coefficient of Correlation

DSamanta@IIT Kharagpur

32

Regression Analysis

 

Significance Test

33 of 204

DSamanta@IIT Kharagpur

33

Regression Analysis

Rank Correlation Analysis

34 of 204

Charles Spearman’s Correlation Coefficient

  • This technique is applicable to determine the degree of correlation between two variables in case of ordinal data.

  • We can assign rank to the different values of a variable with ordinal data type.

DSamanta@IIT Kharagpur

34

Regression Analysis

This correlation measurement is also called Rank correlation

 

Example

Rank assigned

35 of 204

Charles Spearman’s Correlation Coefficient

  • The Spearman’s coefficient is often used as a statistical methods to aid either providing or disproving a hypothesis.

DSamanta@IIT Kharagpur

35

Regression Analysis

 

Definition 2: Charles Spearman’s correlation coefficient

36 of 204

Charles Spearman’s Coefficient of Correlation

Example 2: The hypothesis that the depth of a river does not progressively increase further from the bank.

A sample of size 10 is collected to test the hypothesis, using Spearman’s correlation coefficient.

DSamanta@IIT Kharagpur

36

Regression Analysis

37 of 204

Charles Spearman’s Coefficient of Correlation

Step 1: Assign rank to each data. It is customary to assign rank 1 to the largest data, and 2 to next largest and so on.

Note: If there are two or more samples with the same value, the mean rank should be used.

DSamanta@IIT Kharagpur

37

Regression Analysis

38 of 204

Charles Spearman’s Coefficient of Correlation

Step 2: The contingency table will look like

DSamanta@IIT Kharagpur

38

Regression Analysis

 

39 of 204

Charles Spearman’s Coefficient of Correlation

  •  

DSamanta@IIT Kharagpur

39

Regression Analysis

Spearaman’s rank correlation coefficient

40 of 204

Charles Spearman’s Coefficient of Correlation

  •  

DSamanta@IIT Kharagpur

40

Regression Analysis

41 of 204

DSamanta@IIT Kharagpur

41

Regression Analysis

χ2 Correlation Analysis

42 of 204

Chi-Squared Test of Correlation

  •  

DSamanta@IIT Kharagpur

42

Regression Analysis

43 of 204

 

Contingency Table

Given a data set, it is customary to draw a contingency table, whose structure is given below.

DSamanta@IIT Kharagpur

43

Regression Analysis

44 of 204

 

Entry into Contingency Table: Observed Frequency

In contingency table, an entry Oij denotes the event that attribute A takes on value ai and attribute B takes on value bj (i.e., A = ai, B = bj).

DSamanta@IIT Kharagpur

44

Regression Analysis

45 of 204

 

Entry into Contingency Table: Expected Frequency

In contingency table, an entry eij denotes the expected frequency, which can be calculated as

DSamanta@IIT Kharagpur

45

Regression Analysis

 

A

B

ai

bj

ai

bj

ai

bj

46 of 204

 

DSamanta@IIT Kharagpur

46

Regression Analysis

 

Definition 3: χ2-Value

47 of 204

 

  • The cell that contribute the most to the 𝛘2 value are those whose actual count is very different from the expected.

  • The 𝛘2 statistics tests the hypothesis that A and B are independent. The test is based on a significance level, with (n-1) ×(m-1) degrees of freedom., with a contingency table of size n×m

  • If the hypothesis can be rejected, then we say that A and B are statistically related or associated.

DSamanta@IIT Kharagpur

47

Regression Analysis

48 of 204

 

Example 3: Survey on Gender versus Hobby.

  • Suppose, a survey was conducted among a population of size 1500. In this survey, gender of each person and their hobby as either “book” or “computer” was noted. The survey result obtained in a table like the following.

  • We have to find if there is any association between Gender and Hobby of a people, that is, we are to test whether “gender” and “hobby” are correlated.

DSamanta@IIT Kharagpur

48

Regression Analysis

49 of 204

χ2 –Test

From the survey table, the observed frequency are counted and entered into the contingency table, which is shown below.

DSamanta@IIT Kharagpur

49

Regression Analysis

HOBBY

GENDER

Male

Female

Total

Book

Computer

Total

Example : Survey on Gender versus Hobby.

50 of 204

 

Example 7.3: Survey on Gender versus Hobby.

  • From the survey table, the observed frequency are counted and entered into the contingency table, which is shown below.

DSamanta@IIT Kharagpur

50

Regression Analysis

HOBBY

GENDER

Male

Female

Total

Book

Computer

Total

51 of 204

 

Example 3: Survey on Gender versus Hobby.

  • From the survey table, the expected frequency are counted and entered into the contingency table, which is shown below.

DSamanta@IIT Kharagpur

51

Regression Analysis

HOBBY

GENDER

Male

Female

Total

Book

Computer

Total

52 of 204

 

  •  

DSamanta@IIT Kharagpur

52

Regression Analysis

53 of 204

Significance Test for 𝛘2 -Test

DSamanta@IIT Kharagpur

53

Regression Analysis

 

Cramer’s V Test

54 of 204

More on Correlation Analyses

DSamanta@IIT Kharagpur

54

Regression Analysis

55 of 204

55

  • Binary variable to binary variable correlation
    • Tetrachoric correlation

  • Nominal/ categorical valued variable to binary variable correlation
    • Cramer’s V correlation

  • Continuous variable to binary variable correlation
    • Point-biserial correlation

56 of 204

  • Tetrachoric correlation is a measure of the association between two binary variables, that is, variables that can only take on two values like “yes” and “no” or “good” and “bad.”

  • Suppose, we have the following 2×2 table with two variables, x and y, that both take on two values:

56

CS 40003: Data Analytics

 

Tetrachoric correlation

 

57 of 204

  • Example:
  • Suppose, we want to know whether or not gender is associated with political party preference so we take a simple random sample of 47 voters and survey them on their political party preference.

57

 

Tetrachoric correlation: Example

58 of 204

  •  

58

 

Tetrachoric correlation: Example

 

59 of 204

59

CS 40003: Data Analytics

 

 

Cramer’s V correlation

60 of 204

  • Example:
  • Suppose, we want to know if there is any association between three different eye colors (blue, green and brown) and three regions (east, north and west). After surveying 50 random samples, the following data is obtained.

60

CS 40003: Data Analytics

Cramer’s V correlation: Example

Eye Color

Blue

Green

Brown

East

North

west

61 of 204

61

CS 40003: Data Analytics

Eye Color

Blue

Green

Brown

Row Total

East

19

North

13

west

18

Column Total

14

19

17

Grand total 50

Step 1:

Here all the frequencies are called observed frequency.

Add all values row wise and column wise.

Here

row totals are 19,13,18

column totals are 14, 19, 17

Grand Total is 50

Cramer’s V correlation: Example

62 of 204

62

CS 40003: Data Analytics

 

Eye Color

Blue

Green

Brown

Row Total

East

19

North

13

west

18

Column Total

14

17

Grand Total = 50

Cramer’s V correlation: Example

63 of 204

63

CS 40003: Data Analytics

 

Observed values

Eye Color

Blue

Green

Brown

East

North

west

Expected values

Eye Color

Blue

Green

Brown

East

5.32

North

west

 

Cramer’s V correlation: Example

64 of 204

64

CS 40003: Data Analytics

Step 4:

 

 

The correlation between three different eye colors (blue, green and brown) and three regions (east, north and west) is 0.25

It means eye color is weakly associated with the regions.

Cramer’s V correlation: Example

65 of 204

65

CS 40003: Data Analytics

Point-biserial correlation is a measure of the association between a continuous valued and a binary valued

variable.

 

Point-biserial correlation

66 of 204

66

CS 40003: Data Analytics

Example:

Suppose we want to know whether or not gender is associated with weekly expenditure of the students, where we take a simple random sample of 7 students and survey on them.

12

1

8

1

 

Point-biserial correlation: Example

67 of 204

  •  

67

CS 40003: Data Analytics

12

1

8

1

Here, the coefficient of correlation between gender and weekly expenditure of the students is 0.85.

It means gender is strongly associated with weekly expenditure of the students.

Point-biserial correlation: Example

68 of 204

Regression Analysis

DSamanta@IIT Kharagpur

68

Regression Analysis

69 of 204

Learning Strategies

DSamanta@IIT Kharagpur

69

Regression Analysis

There are two types of learning concepts:

Learning

Learning population parameters

Statistical Learning

Learning models

Machine Learning

70 of 204

Statistical Learning

DSamanta@IIT Kharagpur

70

Regression Analysis

Usually assumes certain properties of the population from which we draw samples:

    • Observation come from a normal population.
    • Sample size is small.
    • Population parameters like mean, variance, etc. are hold good.
    • Requires measurement equivalent to interval scaled data.

71 of 204

Machine Learning

DSamanta@IIT Kharagpur

71

Regression Analysis

  • Does not under any assumption
  • Works well with high volume high dimensional data

This learning strategy needs a very large sample data

Important Point

y = f(x) = ax2 + bx + c

72 of 204

DSamanta@IIT Kharagpur

72

Regression Analysis

Relationship Analysis

73 of 204

Relationship Analysis

  • Example: Wage Data

A large data regarding the wages for a group of employees from the eastern region of India is given.

In particular, we wish to understand the following relationships:

  • Employee’s age and wage: How wages vary with ages?

  • Calendar year and wage: How wages vary with time?

  • Employee’s age and education: Whether wages are anyway related with employees’ education levels?

DSamanta@IIT Kharagpur

73

Regression Analysis

74 of 204

Relationship Analysis

  • Example: Wage Data

  • Case I. Wage versus Age
    • From the data set, we have a graphical representations, which is as follows:

    • How wages vary with ages?

DSamanta@IIT Kharagpur

74

Regression Analysis

How wages vary with ages?

?

75 of 204

Relationship Analysis

  • Example: Wage Data
  • Employee’s age and wage: How wages vary with ages?

Interpretation: On the average, wage increases with age until about 60 years of age, at which point it begins to decline.

DSamanta@IIT Kharagpur

75

Regression Analysis

76 of 204

Relationship Analysis

  • Example: Wage Data

  • Case II. Wage versus Year
    • From the data set, we have a graphical representations, which is as follows:

DSamanta@IIT Kharagpur

76

Regression Analysis

How wages vary with time?

?

77 of 204

Relationship Analysis

  • Example: Wage Data
  • Wage and calendar year: How wages vary with years?

Interpretation: There is a slow but steady increase in the average wage between 2010 and 2016.

.

DSamanta@IIT Kharagpur

77

Regression Analysis

78 of 204

Relationship Analysis

  • Example: Wage Data

  • Case III. Wage versus Education
    • From the data set, we have a graphical representations, which is as follows:

DSamanta@IIT Kharagpur

78

Regression Analysis

?

Whether wages are related with education?

79 of 204

Relationship Analysis

  • Example: Wage Data
  • Wage and education level: Whether wages vary with employees’ education levels?

Interpretation: On the average, wage increases with the level of education.

DSamanta@IIT Kharagpur

79

Regression Analysis

80 of 204

Relationship Analysis

DSamanta@IIT Kharagpur

80

Regression Analysis

What more information can we get?

Given an employee’s wage can we predict his age?

Whether wage has any association with both year and education level?

... and what’s more?

81 of 204

DSamanta@IIT Kharagpur

81

Regression Analysis

Regression Analysis to Find Relationships

82 of 204

An Open Challenge!

  •  

DSamanta@IIT Kharagpur

82

Regression Analysis

83 of 204

Yahoo!

Just decide the values of a and b

(as if storing one point’s data only!)

Note: Here, tricks was to find a relationship among all the points.

DSamanta@IIT Kharagpur

83

Regression Analysis

84 of 204

Measures of Relationship

DSamanta@IIT Kharagpur

84

Regression Analysis

Univariate Population

Bivariate Population

Pressure

1

1.5

1.05

0.96

1.2

2.5

2.8

Multivariate Population

Pressure

1

1.5

1.05

0.96

1.2

2.5

2.8

85 of 204

Measures of Relationship

DSamanta@IIT Kharagpur

85

Regression Analysis

Relationship Analysis

Regression analysis

Linear Regression

Simple Linear Regression

Multiple Linear Regression

Non-linear Regression

Simple Non-linear Regression

Multiple Non-linear Regression

Auto-Regression Analysis

Logistic regression

Binary Logistic Regression

Multinomial Logistic Regression

86 of 204

Regression Analysis

DSamanta@IIT Kharagpur

86

Regression Analysis

The regression analysis is a statistical method to deal with the formulation of mathematical model depicting relationship amongst variables, which can be used for the purpose of prediction of the values of dependent variable, given the values of independent variable(s).

Definition

87 of 204

A Simple Example

DSamanta@IIT Kharagpur

87

Regression Analysis

How Exam Score is related to Hours of Study?

88 of 204

Regression Analyses

DSamanta@IIT Kharagpur

88

Regression Analysis

Regression Analysis

Linear Regression Models

Simple Linear Regression

Multiple Linear Regression

Non-linear Regression Models

89 of 204

DSamanta@IIT Kharagpur

89

Regression Analysis

Simple Linear Regression

90 of 204

Simple Linear Regression Model

DSamanta@IIT Kharagpur

90

Regression Analysis

 

In simple linear regression, we have only two variables:

 

Linear regression

 

91 of 204

Simple Linear Regression Model

DSamanta@IIT Kharagpur

91

Regression Analysis

In simple linear regression, we have only two variables:

 

Note

 

92 of 204

Regression Analysis

DSamanta@IIT Kharagpur

92

Regression Analysis

 

 

93 of 204

Regression Analysis

DSamanta@IIT Kharagpur

93

Regression Analysis

 

Note

94 of 204

True versus Fitted Regression Line

DSamanta@IIT Kharagpur

94

Regression Analysis

 

95 of 204

Least Square Method to estimate 𝛼 and 𝛽

DSamanta@IIT Kharagpur

95

Regression Analysis

 

Concept of Residuals

96 of 204

Least Square Method to estimate 𝛼 and 𝛽

DSamanta@IIT Kharagpur

96

Regression Analysis

 

Sum of Squares Error (SSE)

We need to minimize the value of SSE and hence to determine the parameters of a and b.

97 of 204

Least Square Method to estimate 𝛼 and 𝛽

DSamanta@IIT Kharagpur

97

Regression Analysis

 

Minimizing the Sum of Squares Error (SSE)

 

98 of 204

Least Square Method to estimate 𝛼 and 𝛽

DSamanta@IIT Kharagpur

98

Regression Analysis

Minimizing the Sum of Squares Error (SSE)

 

99 of 204

Least Square Method to estimate 𝛼 and 𝛽

DSamanta@IIT Kharagpur

99

Regression Analysis

Minimizing the Sum of Squares Error (SSE)

 

 

100 of 204

DSamanta@IIT Kharagpur

100

Regression Analysis

 

101 of 204

R2: Measure of Quality Fit

DSamanta@IIT Kharagpur

101

Regression Analysis

Coefficient of Determination

 

 

 

102 of 204

R2: Measure of Quality Fit

DSamanta@IIT Kharagpur

102

Regression Analysis

Coefficient of Determination

 

Note

103 of 204

DSamanta@IIT Kharagpur

103

Regression Analysis

Multiple Linear Regression

104 of 204

Regression Analyses

DSamanta@IIT Kharagpur

104

Regression Analysis

Regression Analysis

Linear Regression Models

Simple Linear Regression

Multiple Linear Regression

Non-linear Regression Models

105 of 204

Multiple Linear Regression

DSamanta@IIT Kharagpur

105

Regression Analysis

    • Multiple Regression Model: When more than one variable are independent variable, then the regression can be estimated as a multiple regression model

    • Multiple Linear Regression: When this model is linear in coefficients, it is called multiple linear regression model

Definition:

106 of 204

Multiple Linear Regression

DSamanta@IIT Kharagpur

106

Regression Analysis

 

Formulation:

107 of 204

Multiple Linear Regression

DSamanta@IIT Kharagpur

107

Regression Analysis

 

Estimating the coefficients: The data points

108 of 204

Multiple Linear Regression

DSamanta@IIT Kharagpur

108

Regression Analysis

Estimating the coefficients: The model formulation

 

109 of 204

Multiple Linear Regression

DSamanta@IIT Kharagpur

109

Regression Analysis

Estimating the coefficients: Minimization of the SSE

 

110 of 204

Multiple Linear Regression

DSamanta@IIT Kharagpur

110

Regression Analysis

Estimating the coefficients: Minimization of the SSE

 

 

111 of 204

DSamanta@IIT Kharagpur

111

Regression Analysis

Non-Linear Regression Model

112 of 204

Regression Analyses

DSamanta@IIT Kharagpur

112

Regression Analysis

Regression Analysis

Linear Regression Models

Simple Linear Regression

Multiple Linear Regression

Non-linear Regression Models

113 of 204

Non-linear Regression Model

DSamanta@IIT Kharagpur

113

Regression Analysis

Definition and Formulation:

 

114 of 204

Solving for Polynomial Regression Model

DSamanta@IIT Kharagpur

114

Regression Analysis

Model formulation:

 

 

115 of 204

Solving for Polynomial Regression Model

DSamanta@IIT Kharagpur

115

Regression Analysis

Transformation to Linear Regression:

 

116 of 204

Linear versus Non-Linear Regression

DSamanta@IIT Kharagpur

116

Regression Analysis

 

 

117 of 204

Linear versus Non-Linear Regression

DSamanta@IIT Kharagpur

117

Regression Analysis

 

 

X

Y

118 of 204

Multiple Non-Linear Regression

DSamanta@IIT Kharagpur

118

Regression Analysis

 

Issues with Multiple Non-Linear Regression

X

Y

Z

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

119 of 204

Today’s Topics of Learning…

  • Auto-Regression Analysis

  • Logistic Regression Analysis
    • Concept of Logistic Regression

    • Types of Logistic Regression

      • Binary Logistic Regression
        • One explanatory variable, two categories
        • Many explanatory variables, two outcomes

      • Multinomial Logistic Regression
        • Many explanatory variable, many categories

DSamanta@IIT Kharagpur

119

Regression Analysis

120 of 204

Auto-Regression Analysis

DSamanta@IIT Kharagpur

120

Regression Analysis

121 of 204

DSamanta@IIT Kharagpur

121

Regression Analysis

Introduction to

Time-Series Data

122 of 204

Time-series Data

DSamanta@IIT Kharagpur

122

Regression Analysis

Time-series data: The data collected on the same observational unit at multiple time periods

Example: Rate of price inflation

Time-series Data

123 of 204

Time-series Data

DSamanta@IIT Kharagpur

123

Regression Analysis

    • Aggregate consumption and GDP for a country (for example, 20 years of quarterly observations = 80 observations)

    • Yen/$, pound/$ and Euro/$ exchange rates (daily data for 1 year = 365 observations)

    • Cigarette consumption per capita in a state, by years

    • Rainfall data over a year or a period of years

    • Sales of tea from a tea shop in a season

Examples of time-series data:

124 of 204

Use of Time-series Data

DSamanta@IIT Kharagpur

124

Regression Analysis

  • To develop forecast model
    • What will be the rate of inflation in next year?

  • To estimate dynamic causal effects
    • If the rate of interest increases, what will be the effect on the rates of inflation and unemployment in 3 months? in 12 months?
    • What is the effect over time on electronics good consumption due to a hike in the excise duty?

  • Time dependent analysis
    • Rates of inflation and unemployment in the country can be observed over a time period.

Use of time-series data

125 of 204

Modeling with Time-series Data

DSamanta@IIT Kharagpur

125

Regression Analysis

  • Correlation over time
    • Serial correlation, also called autocorrelation
    • Calculating standard error

  • To estimate dynamic causal effects
    • Under which dynamic effects can be estimated?
    • How to estimate?

  • Forecasting model
    • Forecasting model build on regression model

126 of 204

Modeling with Time-series Data: Example

DSamanta@IIT Kharagpur

126

Regression Analysis

  • Can we predict the trend at a time say 2017?

Modeling with time-series data

127 of 204

DSamanta@IIT Kharagpur

127

Regression Analysis

Concept and Notations

128 of 204

Related Concepts and Notations

DSamanta@IIT Kharagpur

128

Regression Analysis

 

129 of 204

Related Concepts and Notations

DSamanta@IIT Kharagpur

129

Regression Analysis

 

130 of 204

Related Concepts and Notations

DSamanta@IIT Kharagpur

130

Regression Analysis

 

131 of 204

Autocorrelation coefficient

DSamanta@IIT Kharagpur

131

Regression Analysis

Autocorrelation

The correlation of a series with its own lagged values is called autocorrelation (also called serial correlation)

 

 

132 of 204

Covariance

DSamanta@IIT Kharagpur

132

Regression Analysis

 

 

Yt-j

. . .

Yt

x1

 

y1

x2

 

y2

. . .

 

. . .

xj

 

yj

.

 

.

.

 

.

.

 

.

xn

 

yn

 

133 of 204

Example: Autocorrelation

DSamanta@IIT Kharagpur

133

Regression Analysis

  • For the given data, say ρ1 = 0.84 between two given consecutive years
    • This implies that the Dollars per Pound is highly serially correlated

  • Similarly, we can determine ρ2 , ρ3 …. etc.

Example

134 of 204

DSamanta@IIT Kharagpur

134

Regression Analysis

Auto-Regression Model

135 of 204

Auto-Regression Model

DSamanta@IIT Kharagpur

135

Regression Analysis

An autoregressive model (also called AR model) is used to model a future behavior for a time-ordered data, using data from past behaviors.

      • Essentially, it is a linear regression analysis of a dependent variable using one or more variables(s) in a given time-series data.

Definition

 

136 of 204

Auto-Regression Model for Forecasting

DSamanta@IIT Kharagpur

136

Regression Analysis

 

Definition

137 of 204

 

DSamanta@IIT Kharagpur

137

Regression Analysis

 

 

 

138 of 204

Computing AR Coefficients

DSamanta@IIT Kharagpur

138

Regression Analysis

  • A number of techniques known for computing the AR coefficients
  • The most common method is called Least Squares Method (LSM)
  • The LSM is based upon the Yule-Walker equations

Computing AR(p) model

 

139 of 204

Computing AR Coefficients

DSamanta@IIT Kharagpur

139

Regression Analysis

 

Computing AR (p): Yule-Walker Equations

 

140 of 204

Logistic Regression

DSamanta@IIT Kharagpur

140

Regression Analysis

141 of 204

DSamanta@IIT Kharagpur

141

Regression Analysis

Introduction

  • Regression analysis and logistic regression
  • Concept of logistic regression
  • Types of logistic regression

142 of 204

Regression Analysis and Logistic Regression

    • Regression analysis
      • Based on the principle of Least Square Estimation (LSE)
        • The parameters are chosen to minimize the sum of squared errors (SSE)
      • Minimizes error in prediction
        • If the error distribution is normal with constant variance, the LSE estimates the parameters accurately; that is, model is the best possible and with the smallest standard errors
      • Applicable when dependent variable follows normal distribution

    • Logistic regression analysis
      • When a dependent variable does not follow normal distribution
      • Value of a dependent variable may be with 2, 3 or a fewer more outcomes

DSamanta@IIT Kharagpur

142

Regression Analysis

143 of 204

Regression Analysis and Logistic Regression

    • Logistic regression
      • Considers maximum likelihood estimation (MLE) to give better result.
      • In MLE, the likelihood is the probability of the observed data set given a set of proposed values for the parameters.
      • The principle of MLE is to estimate parameters by choosing parameter values that give the largest possible likelihood.

DSamanta@IIT Kharagpur

143

Regression Analysis

    • Regression analysis predicts a value of a dependent variable
    • Logistic regression predicts the probability of a given value of a dependent variable
    • Both estimates their respective model parameters

Note

144 of 204

Regression Analysis and Logistic Regression

DSamanta@IIT Kharagpur

144

Regression Analysis

A Regression model and Logistic Regression model

145 of 204

An Example

DSamanta@IIT Kharagpur

145

Regression Analysis

Hours (xi)

0.50

0.75

1.00

1.25

1.50

1.75

1.75

2.00

2.25

2.50

2.75

3.00

3.25

3.50

4.00

4.25

4.50

4.75

5.00

5.50

Pass (yi)

0

0

0

0

0

0

1

0

1

0

1

0

1

0

1

1

1

1

1

1

146 of 204

Concept of Logistic Regression

DSamanta@IIT Kharagpur

146

Regression Analysis

    • Developed and popularized primarily by Joseph Berkson in (1944), where he coined the term logit.

    • Logistic regression is a statistical method.
      • uses a logistic function to model a binary dependent variable.
      • although many more complex extensions exist.

    • Logistic regression (or logit regression) estimates the parameters of a logistic model.

Introduction

147 of 204

Concept of Logistic Regression

DSamanta@IIT Kharagpur

147

Regression Analysis

    • A binary logistic model has a dependent variable with two possible values, such as Pass or Fail, Happy or Sad etc.

    • It is represented by an indicator variable, where the two values are labeled ‘1’ and ‘0’

Introduction

1 0

148 of 204

Concept of Logistic Regression

DSamanta@IIT Kharagpur

148

Regression Analysis

 

What is logistic regression?

 

 

149 of 204

An Illustration

A sample is collected to examine the effect of toxic substance on tumor. A subject is examined for the toxic content in the body and then the presence (1) or absence (0) of tumors. The independent variable is the concentration of the toxic substance “Conc” . The number of subjects at each concentration (N) and the number having tumors “Tumor” is shown in the table.

Odds and ln(odds) are also included in the table.

DSamanta@IIT Kharagpur

149

Regression Analysis

Conc

N

Tumor

Odds

ln(odds)

0.0

50

2

0.0417

-3.18

2.1

54

5

0.1020

-3.28

5.4

46

5

0.1220

-2.10

8.0

51

10

0.2439

-1.41

15.0

50

40

4.0000

+1.39

19.5

52

42

4.2000

+1.44

 

150 of 204

An Illustration

  •  

DSamanta@IIT Kharagpur

150

Regression Analysis

Conc

N

Tumor

Odds

ln(odds)

0.0

50

2

0.0417

-3.18

2.1

54

5

0.1020

-3.28

5.4

46

5

0.1220

-2.10

8.0

51

10

0.2439

-1.41

15.0

50

40

4.0000

+1.39

19.5

52

42

4.2000

+1.44

151 of 204

An Illustration

  •  

DSamanta@IIT Kharagpur

151

Regression Analysis

 

152 of 204

Concept of Logistic Regression

DSamanta@IIT Kharagpur

152

Regression Analysis

    • The log-odds (the logarithm of the odds) is a linear combination of

      • one or more independent variables ("predictors").

      • the independent variables can each be a continuous variable (any real value).

    • The probability of the value can vary between 0 (such as certainly false) and 1 (such as certainly true).

Logit in Logistic Regression

153 of 204

Concept of Logistic Regression

DSamanta@IIT Kharagpur

153

Regression Analysis

    • Increasing one of the independent variables, multiplicatively scales the odds of the given outcome at a constant rate, with each independent variable having its own parameter;

      • for a binary dependent variable this generalizes the odds ratio.

    • Binary logistic regression: The dependent variable has two levels (categorical).

    • Multinomial logistic regression: Outputs with more than two levels.

    • Ordinal logistic regression: if the multiple categories are ordered, then it is called ordinal logistic regression.

Output of logistic function

154 of 204

Logistic Regression as Classifier

DSamanta@IIT Kharagpur

154

Regression Analysis

Logistic regression models the probabilities for classification problems with possible outcomes.

      • It’s an extension of the linear regression model for classification problems.
    • The logistic regression simply models probability of output in terms of input and does not perform statistical classification (it is not a classifier).

    • Although it can be used to make a classifier.
      • By choosing a cut-off value and classifying inputs with probability greater than the cutoff as one class, below the cut-off as the other

155 of 204

DSamanta@IIT Kharagpur

155

Regression Analysis

Logistic Regression Techniques

156 of 204

Types of Logistic Regression

DSamanta@IIT Kharagpur

156

Regression Analysis

Logistic regression

Binary logistic regression

One explanatory variable, two categories

Many explanatory variable, two categories

Multinomial logistic regression

Many explanatory variables, many categories

 

157 of 204

Types of Logistic Regression

DSamanta@IIT Kharagpur

157

Regression Analysis

    • Explanatory variable:

X: Hours Study

    • Outcome: Pass or Fail

Case 1: One explanatory variable, two categories

    • Explanatory variable

X1: Hours Study X2: 12th % Marks

    • Outcome: Pass or Fail

Case 2: Many explanatory variable, two categories

    • Explanatory variable:

X1: Hours Study X2: 12th % Marks X3: Age

    • Outcome: Bad, Good, Excellent

Case 3: Many explanatory variable, many categories

158 of 204

Binary Logistic Regression

DSamanta@IIT Kharagpur

158

Regression Analysis

  • Logistic Regression Analysis
    • Binary Logistic Regression
      • One explanatory variable, two categories

159 of 204

DSamanta@IIT Kharagpur

159

Regression Analysis

One explanatory variable,

two categories

160 of 204

One Explanatory Variable, Two Categories

DSamanta@IIT Kharagpur

160

Regression Analysis

A group of 20 students spends between 0 and 6 hours studying for an exam.

The table shows the number of hours each student spent studying, and whether they passed (1) or failed (0).

Hours (xi)

0.50

0.75

1.00

1.25

1.50

1.75

1.75

2.00

2.25

2.50

2.75

3.00

3.25

3.50

4.00

4.25

4.50

4.75

5.00

5.50

Pass (yi)

0

0

0

0

0

0

1

0

1

0

1

0

1

0

1

1

1

1

1

1

How does the number of hours spent studying affect the probability of the student passing the exam?

161 of 204

One Explanatory Variable, Two Categories

DSamanta@IIT Kharagpur

161

Regression Analysis

Hours (xi)

0.50

0.75

1.00

1.25

1.50

1.75

1.75

2.00

2.25

2.50

2.75

3.00

3.25

3.50

4.00

4.25

4.50

4.75

5.00

5.50

Pass (yi)

0

0

0

0

0

0

1

0

1

0

1

0

1

0

1

1

1

1

1

1

  • The reason for using logistic regression for this problem is that the values of the dependent variable, pass and fail, while represented by "1" and "0", are not cardinal numbers.

  • If the problem was changed so that pass/fail was replaced with the grade 0 to 100 (cardinal numbers), then simple regression analysis could be used.

How does the number of hours spent studying affect the probability of the student passing the exam?

Note

162 of 204

One Explanatory Variable, Two Categories

DSamanta@IIT Kharagpur

162

Regression Analysis

 

Hours (xi)

0.50

0.75

1.00

1.25

1.50

1.75

1.75

2.00

2.25

2.50

2.75

3.00

3.25

3.50

4.00

4.25

4.50

4.75

5.00

5.50

Pass (yi)

0

0

0

0

0

0

1

0

1

0

1

0

1

0

1

1

1

1

1

1

163 of 204

One Explanatory Variable, Two Categories

DSamanta@IIT Kharagpur

163

Regression Analysis

 

 

164 of 204

One Explanatory Variable, Two Categories

DSamanta@IIT Kharagpur

164

Regression Analysis

 

165 of 204

One Explanatory Variable, Two Categories

DSamanta@IIT Kharagpur

165

Regression Analysis

 

166 of 204

One Explanatory Variable, Two Categories

DSamanta@IIT Kharagpur

166

Regression Analysis

 

167 of 204

One Explanatory Variable, Two Categories

DSamanta@IIT Kharagpur

167

Regression Analysis

 

 

 

168 of 204

One Explanatory Variable, Two Categories

DSamanta@IIT Kharagpur

168

Regression Analysis

 

 

169 of 204

One Explanatory Variable, Two Categories

DSamanta@IIT Kharagpur

169

Regression Analysis

 

Coefficient

Std. Error

z-value

p-value (Wald)

Intercept (β0)

−4.0777

1.7610

−2.316

0.0206

Slope (β1)

1.5046

0.6287

2.393

0.0167

Hours (xi)

0.50

0.75

1.00

1.25

1.50

1.75

1.75

2.00

2.25

2.50

2.75

3.00

3.25

3.50

4.00

4.25

4.50

4.75

5.00

5.50

Pass (yi)

0

0

0

0

0

0

1

0

1

0

1

0

1

0

1

1

1

1

1

1

 

170 of 204

One Explanatory Variable, Two Categories

DSamanta@IIT Kharagpur

170

Regression Analysis

Coefficient

Std. Error

z-value

p-value (Wald)

Intercept (β0)

−4.0777

1.7610

−2.316

0.0206

Slope (β1)

1.5046

0.6287

2.393

0.0167

 

171 of 204

One Explanatory Variable, Two Categories

DSamanta@IIT Kharagpur

171

Regression Analysis

 

Hours of study (x)

Passing exam

Log-odds (t)

Odds (et)

Probability (p)

1

−2.57

0.076 ≈ 1:13.1

0.07

2

−1.07

0.34 ≈ 1:2.91

0.26

μ=2.71...

0

1

0.5

3

0.44

1.55

0.61

4

1.94

6.96

0.87

5

3.45

31.4

0.97

172 of 204

One Explanatory Variable, Two Categories

DSamanta@IIT Kharagpur

172

Regression Analysis

 

173 of 204

DSamanta@IIT Kharagpur

173

Regression Analysis

Many explanatory variable,

two categories

  • Logistic Regression Analysis
    • Binary Logistic Regression
      • Many explanatory variable, two categories

174 of 204

Many Explanatory Variable, Two Categories

DSamanta@IIT Kharagpur

174

Regression Analysis

 

  • An additional generalization has been introduced in which the base of the model b is not restricted to the Euler number e.

    • In most applications, the base b of the logarithm is usually taken to be e.

    • However, in some cases it can be easier to communicate results by working in base 2 or base 10.

175 of 204

Many Explanatory Variable, Two Categories

DSamanta@IIT Kharagpur

175

Regression Analysis

 

176 of 204

Many Explanatory Variable, Two Categories

DSamanta@IIT Kharagpur

176

Regression Analysis

 

177 of 204

Many Explanatory Variable, Two Categories

DSamanta@IIT Kharagpur

177

Regression Analysis

 

178 of 204

Many Explanatory Variable, Two Categories

DSamanta@IIT Kharagpur

178

Regression Analysis

 

179 of 204

Many Explanatory Variable, Two Categories

DSamanta@IIT Kharagpur

179

Regression Analysis

 

 

180 of 204

DSamanta@IIT Kharagpur

180

Regression Analysis

Multinomial Logistic Regression

181 of 204

Multinomial Logistic Regression

DSamanta@IIT Kharagpur

181

Regression Analysis

Binary

logistic regression

One explanatory variable, two categories

Many explanatory variable, two categories

Multinomial

logistic regression

Logistic

Regression

Many explanatory variable, many categories

182 of 204

Many Explanatory Variable, Many Categories

DSamanta@IIT Kharagpur

182

Regression Analysis

 

 

183 of 204

Many Explanatory Variable, Many Categories

DSamanta@IIT Kharagpur

183

Regression Analysis

 

 

184 of 204

Many Explanatory Variable, Many Categories

DSamanta@IIT Kharagpur

184

Regression Analysis

 

Note :

 

185 of 204

Many Explanatory Variable, Many Categories

DSamanta@IIT Kharagpur

185

Regression Analysis

 

186 of 204

Many Explanatory Variable, Many Categories

DSamanta@IIT Kharagpur

186

Regression Analysis

 

187 of 204

Many Explanatory Variable, Many Categories

DSamanta@IIT Kharagpur

187

Regression Analysis

 

188 of 204

DSamanta@IIT Kharagpur

188

Regression Analysis

Applications of Logistic Regression

189 of 204

Applications of Logistic Regression

DSamanta@IIT Kharagpur

189

Regression Analysis

  • In medical domains:
    • The Trauma and Injury Severity Score (TRISS), which is widely used to predict mortality in injured patients, was originally developed by Boyd et al. using logistic regression.

    • Many other medical scales used to assess severity of a patient have been developed using logistic regression.

    • Logistic regression may be used to predict the risk of developing a given disease (e.g. diabetes; coronary heart disease), based on observed characteristics of the patient (age, gender, body mass index, results of various blood tests, etc.)

190 of 204

Applications of Logistic Regression

DSamanta@IIT Kharagpur

190

Regression Analysis

  • In social sciences:
    • Another example might be to predict whether an Indian voter will vote National Congress or Communist Party or BJP or Any other party, based on age, income, gender, race, state of residence, votes in previous elections, etc.

  • In engineering:
    • The technique can also be used in engineering, especially for predicting the probability of failure of a given process, system or product.

  • In marketing:
    • It is also used in marketing applications, such as prediction of a customer's propensity to purchase a product or halt a subscription, etc.

191 of 204

DSamanta@IIT Kharagpur

191

Regression Analysis

Problems to Ponder

192 of 204

 

Statement: In a study of urban planning in a country, a survey was taken of 50 cities; 24 used Happiness Index (HI) and 26 did not. One part of the study was to investigate the relationship between the presence or absence of HI and the median family income of the city(x). The data are given in Table 1, with median income in order of $1000s.

DSamanta@IIT Kharagpur

192

Regression Analysis

HI

x

HI

x

HI

x

HI

x

0

9.2

0

10.5

1

9.6

1

12.5

0

9.2

0

10.5

1

10.1

1

12.6

0

9.3

0

10.9

1

10.3

1

12.6

0

9.4

0

11.0

1

10.9

1

12.6

0

9.5

0

11.2

1

10.9

1

12.9

0

9.5

0

11.2

1

11.1

1

12.9

0

9.5

0

11.5

1

11.1

1

12.9

0

9.6

0

11.7

1

11.1

1

12.9

0

9.7

0

11.8

1

11.5

1

13.1

0

9.7

0

12.1

1

11.8

1

13.2

0

9.8

0

12.3

1

11.9

1

13.5

0

9.8

0

12.5

1

12.1

0

9.9

0

12.9

1

12.2

Table 1: Data from Happiness Study

193 of 204

 

DSamanta@IIT Kharagpur

193

Regression Analysis

Statement: In a study of urban planning in Florida, a survey was taken of 50 cities; 24 used tax increment funding (TF) and 26 did not. One part of the study was to investigate the relationship between the presence or absence of TF and the median family income of the city(x). The data are given in the Table, with median income in order $1000s.

 

Income category

Income Category

Number of 0

Number of 1

Odds

ln(Odds))

9.5

13

1

-2.56395

10.5

3

4

0.287432

11.5

6

6

0

12.5

4

10

0.916291

13.5

0

3

1.386294

194 of 204

 

DSamanta@IIT Kharagpur

194

Regression Analysis

Income category

Mid-point of Income Category

NUMBER OF 0

NUMBER OF 1

ODDS

LN(ODDS)

9.5

13

1

-2.56395

10.5

3

4

0.287432

11.5

6

6

0

12.5

4

10

0.916291

13.5

0

3

1.386294

 

195 of 204

 

DSamanta@IIT Kharagpur

195

Regression Analysis

 

 

196 of 204

 

Statement: Time Magazine (2006) used data from the USA to compare whites and blacks opinions of the death penalty. The data consisted of responses from 32,937 participants collected between 1972 and 1996. The outcome variable was whether the respondent did or did not support the death penalty. The survey provided a table of the percentage of whites and blacks each year that supported the death penalty.

DSamanta@IIT Kharagpur

196

Regression Analysis

Year

White (%)

Black (%)

1972

57.4

28.8

1973

63.6

35.8

1974

66.3

36.3

1975

63.2

31.9

1976

67.5

41.1

1977

70

41.6

1978

69.4

43

1980

70.3

39.1

1982

76.9

48.4

1983

76.2

45

1984

74.5

43.5

1985

79

49.7

1986

75.3

42.7

Year

White (%)

Black (%)

1987

73.7

42.9

1988

76

42.5

1989

76.5

56.1

1990

77.7

52.3

1991

71.4

42.7

1993

75.4

51.5

1994

78.3

50.7

1996

75.5

50.3

197 of 204

 

DSamanta@IIT Kharagpur

197

Regression Analysis

Year

White (%)

Black (%)

1972

57.4

28.8

1973

63.6

35.8

1974

66.3

36.3

1975

63.2

31.9

1976

67.5

41.1

1977

70

41.6

1978

69.4

43

1980

70.3

39.1

1982

76.9

48.4

1983

76.2

45

1984

74.5

43.5

1985

79

49.7

1986

75.3

42.7

Year

White (%)

Black (%)

1987

73.7

42.9

1988

76

42.5

1989

76.5

56.1

1990

77.7

52.3

1991

71.4

42.7

1993

75.4

51.5

1994

78.3

50.7

1996

75.5

50.3

Convert the percentages given in the table to the In(odds) within each race and year, and plot ln(odds) versus year. Comment on any patterns you see. If there is a trend in time, does it appear linear or quadratic?

198 of 204

 

DSamanta@IIT Kharagpur

198

Regression Analysis

Year

White (%)

Odds

ln(odds)

1972

57.4

1.347418

0.29819005

1973

63.6

1.747253

0.558044696

1974

66.3

1.967359

0.67669206

1975

63.2

1.717391

0.540806456

1976

67.5

2.076923

0.730887509

1977

70

2.333333

0.84729786

1978

69.4

2.267974

0.818886859

1980

70.3

2.367003

0.861624753

1982

76.9

3.329004

1.202673259

1983

76.2

3.201681

1.163675882

1984

74.5

2.921569

1.072120673

1985

79

3.761905

1.324925415

1986

75.3

3.048583

1.114676891

Year

White (%)

Odds

ln(odds)

1987

73.7

2.802281

1.03043386

1988

76

3.166667

1.15267951

1989

76.5

3.255319

1.18029032

1990

77.7

3.484305

1.248268579

1991

71.4

2.496503

0.914891152

1993

75.4

3.065041

1.120060832

1994

78.3

3.608295

1.283235342

1996

75.5

3.081633

1.125459539

 

199 of 204

 

DSamanta@IIT Kharagpur

199

Regression Analysis

 

200 of 204

 

DSamanta@IIT Kharagpur

200

Regression Analysis

Table

Year

Black (%)

Odds

ln(odds)

1972

28.8

0.404494

-0.90512

1973

35.8

0.557632

-0.58406

1974

36.3

0.569859

-0.56237

1975

31.9

0.468429

-0.75837

1976

41.1

0.697793

-0.35983

1977

41.6

0.712329

-0.33922

1978

43

0.754386

-0.28185

1980

39.1

0.642036

-0.44311

1982

48.4

0.937984

-0.06402

1983

45

0.818182

-0.20067

1984

43.5

0.769912

-0.26148

1985

49.7

0.988072

-0.012

1986

42.7

0.745201

-0.2941

Year

Black (%)

Odds

ln(odds)

1987

42.9

0.751313

-0.28593

1988

42.5

0.73913

-0.30228

1989

56.1

1.277904

0.245221

1990

52.3

1.096436

0.092065

1991

42.7

0.745201

-0.2941

1993

51.5

1.061856

0.060018

1994

50.7

1.028398

0.028002

1996

50.3

1.012072

0.012

201 of 204

 

DSamanta@IIT Kharagpur

201

Regression Analysis

 

202 of 204

DSamanta@IIT Kharagpur

202

Regression Analysis

REFERENCES

  • The Elements of Statistical Learning, Data Mining, Inference, and Prediction (2nd Edn.), Trevor Hastie, Robert Tibshirani, Jerome Friedman, Springer, 2014.

203 of 204

Reference

DSamanta@IIT Kharagpur

203

Regression Analysis

  • The detail material related to this lecture can be found in

Web: https://en.wikipedia.org/wiki/Logistic_regression

204 of 204

Any question?

DSamanta@IIT Kharagpur

204

Regression Analysis

dsamanta@iitkgp.ac.in