1 of 23

2 of 23

Chi Square test

  • Chi square test is"non-parametric" which means that the chi‑square test does not require assumptions about population parameters nor do they test hypotheses about population parameters.
  • Previous examples of hypothesis tests, such as the t tests and analysis of variance, are parametric tests and they do include assumptions about parameters and hypotheses about parameters.

3 of 23

Chi Square test

  • The most obvious difference between the chi‑square tests and the other hypothesis tests we have considered (t and ANOVA) is the nature of the data.
  • For chi‑square, the data are frequencies rather than numerical scores.

​

4 of 23

The Chi Square Test

  • A statistical method used to determine goodness of fit
    • Goodness of fit refers to how close the observed data are to those predicted from a hypothesis

​

  • Note:
    • The chi square test does not prove that a hypothesis is correct
      • It evaluates to what extent the data and the hypothesis have a good fit

5 of 23

The Chi-Square Test for Goodness-of-Fit

  • The chi-square test for goodness-of-fit uses frequency data from a sample to test hypotheses about the proportions of a population.
  • Each individual in the sample is classified into one category on the scale of measurement.
  • The data, called observed frequencies, simply count how many individuals from the sample are in each category.

5

6 of 23

The Chi-Square Test for Goodness-of-Fit (cont.)

  • The null hypothesis specifies the proportion of the population that should be in each category.
  • The proportions from the null hypothesis are used to compute expected frequencies that describe how the sample would appear if it were in perfect agreement with the null hypothesis.

6

7 of 23

Chi-Square as a Test for independence

  • Chi-square test: an inferential statistics technique designed to test for significant relationships between two variables organized in a bivariate table.

​

  • Chi-square requires no assumptions about the shape of the population distribution from which a sample is drawn.

​

8 of 23

Limitations of the Chi-Square Test

  • The chi-square test does not give us much information about the strength of the relationship

​

  • The chi-square test is sensitive to sample size. The size of the calculated chi-square is directly proportional to the size of the sample, independent of the strength of the relationship between the variables.

​

  • The chi-square test is also sensitive to small expected frequencies in one or more of the cells in the table.

​

9 of 23

Statistical Independence

  • Independence (statistical): the absence of association between two cross-tabulated variables.

10 of 23

The Chi-Square Test for Independence

  • the chi-square test for independence, can be used and interpreted in two different ways:

1. Testing hypotheses about the relationship between two variables in a population, or

2. Testing hypotheses about differences between proportions for two or more populations.

10

11 of 23

The Chi-Square Test for Independence (cont.)

  • The first version of the chi-square test for independence views the data as one sample in which each individual is classified on two different variables.
  • The data are usually presented in a matrix with the categories for one variable defining the rows and the categories of the second variable defining the columns.

11

12 of 23

The Chi-Square Test for Independence (cont.)

  • The second version of the test for independence views the data as two (or more) separate samples representing the different populations being compared.
  • The same variable is measured for each sample by classifying individual subjects into categories of the variable.
  • The data are presented in a matrix with the different samples defining the rows and the categories of the variable defining the columns..

12

13 of 23

Hypothesis Testing with Chi-Square

Chi-square follows five steps:

  1. Making assumptions (random sampling)

​

  • Stating null hypotheses and the alternative

​

  • Selecting the sampling distribution and specifying the test statistic

​

  • Computing the test statistic

​

  • Making a decision and interpreting the results

​

14 of 23

The Assumptions

  • The chi-square test requires no assumptions about the shape of the population distribution from which the sample was drawn.

​

  • However, like all inferential techniques it assumes random sampling.

​

​

15 of 23

Stating Null and alternative Hypotheses

​

  • The null hypothesis (H0) states that no association exists between the two cross-tabulated variables in the population, and therefore the variables are statistically independent.
  • The alternate hypothesis (H1) proposes that the two variables are related in the population.

​

16 of 23

The Concept of Expected Frequencies

Expected frequencies fe : the cell frequencies that would be expected in a bivariate table if the two tables were statistically independent.

​

Observed frequencies fo: the cell frequencies actually observed in a bivariate table.

​

​

17 of 23

Calculating Expected Frequencies

To obtain the expected frequencies for any cell in any cross-tabulation in which the two variables are assumed independent, multiply the row and column totals for that cell and divide the product by the total number of cases in the table.

fe = (column total)(row total)

N

18 of 23

Calculating the Chi-Square

fe = expected frequencies

fo = observed frequencies

​

​

19 of 23

The Sampling Distribution of Chi-Square

  • The sampling distribution of chi-square tells the probability of getting values of chi-square, assuming no relationship exists in the population.

​

  • The chi-square sampling distributions depend on the degrees of freedom.

​

  • The χ2 sampling distribution is not one distribution, but is a family of distributions.

​

20 of 23

The Sampling Distribution of Chi-Square

  • The distributions are positively skewed. The research hypothesis for the chi-square is always a one-tailed test.

​

  • Chi-square values are always positive. The minimum possible value is zero, with no upper limit to its maximum value.

​

  • As the number of degrees of freedom increases, the χ2 distribution becomes more symmetrical.

​

21 of 23

22 of 23

Determining the Degrees of Freedom

​

df = (r – 1)(c – 1)

​

where

r = the number of rows

c = the number of columns

​

​

23 of 23

Calculating Degrees of Freedom

How many degrees of freedom would a table with 3 rows and 2 columns have?

​

(3 – 1)(2 – 1) =

2

2 degrees of freedom