1 of 33

Discriminant Function�Analysis

MAR 536 Biological Statistics II

April 8 2025

2 of 33

Lecture Outline

  • Eigenanalysis (again!)
    • PCA
    • Discriminant analysis
  • 14.1 Introduction to discriminant analysis (DA)
  • 14.2 Assumptions
  • 14.3 Sparrow example

Matrix Algebra

2

Advanced Stats

3 of 33

Variance-Covariance Matrix

  • Variance is the average sum-of-squares of a variable (corrected for low sample size, n-1).
  • Covariance is the average cross-product of two variables.
  • The covariance matrix is symmetric, with variance of each variable on the diagonal and paired covariances on the off diagonals.
  • The trace of C is the ‘total variance”

Matrix Algebra

3

Advanced Stats

4 of 33

Principal Component Analysis

  • PCA involves eigenanalysis of either the covariance matrix (C, variables of similar unit and scale, e.g., morphometrics) or the correlation matrix (R, variables of different unit or scale).

  • The sum of L and trace of C and is the total variance, but C has covariance, and L has zero covariance.

Matrix Algebra

4

Advanced Stats

CV = LV

5 of 33

Principal Components Analysis

  • Eigenanalysis transforms the covariance matrix, C, to a series of uncorrelated components.
    • The vector of eigenvalues (L) has elements that correspond to the amount of total variance explained by each component (usually ordered largest to smallest value).
    • The matrix of eigenvectors (V) can be used to transform the data matrix to principal component scores (P).

Matrix Algebra

5

Advanced Stats

CV = LV

P = YV

6 of 33

Discriminant Analysis

  • Principal component analysis finds best low-dimensional representation of the variation in a multivariate data set. Each principal component is a particular linear combination of variables.

  • Linear discriminant analysis (LDA) finds the linear combinations of the original variables that gives the best possible separation between known groups.

7 of 33

PCA vs Discriminant Analysis

8 of 33

Discriminant Analysis

  • Main objective of LDA is to find a projection matrix that maximizes the ratio of the determinant of the between-groups covariance matrix to the determinant of the within-groups covariance matrix.

  • Recall the determinant tells us how much variance a matrix has.

Matrix Algebra

8

Advanced Stats

9 of 33

Discriminant Analysis

  • Discriminant function analysis (DFA) involves eigenanalysis of the pooled, within-groups variance-covariance matrix (CW) to maximize differences in discriminant function scores (F) for k groups.
  • Assumes variance-covariance is homogeneous among groups (i).

Matrix Algebra

9

Advanced Stats

10 of 33

Discriminant Analysis

  • DFA of two groups.

Matrix Algebra

10

Advanced Stats

11 of 33

Discriminant Analysis

  • DFA of two groups.

Matrix Algebra

11

Advanced Stats

12 of 33

Discriminant Analysis

  • Eigenanalysis of pooled, within-groups variance-covariance matrix.

Matrix Algebra

12

Advanced Stats

13 of 33

Discriminant Analysis

Matrix Algebra

13

14 of 33

Various Terminology

Discriminant Analysis

Discriminant Function Analysis

Multiple Discriminant Analysis

Simple Discriminant Analysis

Canonical variate Analysis

Fisher’s Discriminant Analysis

Fisher’s Analysis

Discrimination Analysis

Linear Discriminant Analysis

Linear Classifier Analysis

General Discriminant Analysis

Local Discriminant Analysis

15 of 33

Uses of Discriminant Analysis

  • Classify observations into groups
    • Determine the most parsimonious way to classify groups
    • Test classification accuracy
  • Investigate differences among groups
    • Identify variables that are different among groups
    • Identify variables that are not as different among groups

Matrix Algebra

15

Advanced Stats

16 of 33

Discriminant Function

  • Zi1 is the discriminant score for discriminant function 1 and observation I with variables 1-p.
  • c1p is the discriminant coefficient of the first discriminant function and variable p.

Matrix Algebra

16

Advanced Stats

17 of 33

Discriminant Analysis Assumptions

  • Normal Distribution
    • Assumes a multivariate normal distribution
  • Homogeneity of variance/covariance
    • Assumes homogeneous variances across groups
  • Collinearity
    • Relies on some correlation among variables but not complete correlation (r~1)
  • Independence
    • All observations are independent

Matrix Algebra

17

Advanced Stats

18 of 33

Discriminant Analysis Assumptions

  • Observations can be divided a priori into groups
    • Each observation can only be in one group
  • Variables are on a continuous scale
  • Linear relationship between variables
  • Adequate sample size
    • Minimum of two observations per group
    • Recommended to have 4-5x the number of variables
    • Smallest group Observations > number of variables

Matrix Algebra

18

Advanced Stats

19 of 33

14.3 Sparrow Data

  • 7 body measurements
  • 10 observers (0 and 9 removed because of low sample size)
  • >1000 birds sampled
  • Month and year

Matrix Algebra

19

Advanced Stats

Saltmarsh sharp-tailed sparrow

20 of 33

Similar spread among observers indicates homogeneity

21 of 33

Histograms show univariate normal distribution

22 of 33

14.3 Sparrow Data

  • Cleveland dotplot also indicates homogeneity

23 of 33

14.3

  • Pairplot indicates linearity and correlation

24 of 33

14.3 Sparrow Code

Matrix Algebra

24

Advanced Stats

25 of 33

14.3 Sparrow Data

Coefficients of linear discriminants:

LD1 LD2 LD3 LD4 LD5 LD6

Flatwing -0.03785084 -0.2737932 0.11491239 -0.42823580 0.2914571 -0.00853429

tarsus 0.24588046 1.6722403 -0.29390262 -0.32513644 0.4096320 0.23987305

head -0.68101337 -0.6089259 -1.48925091 0.92558503 0.7035545 -0.27392299

culmen -1.61462537 0.3820697 1.47757092 0.41349515 0.2853035 0.30302032

nalospi 2.91308820 -0.1256085 0.73034603 0.04576332 0.4617898 -0.03379820

wtnew 0.04825864 -0.1474904 -0.04105511 0.11254511 -0.5204707 0.67584015

Proportion of trace:

LD1 LD2 LD3 LD4 LD5 LD6

0.6776 0.2034 0.0923 0.0158 0.0087 0.0022

LD1 and LD2 account for 88.1% of variance

26 of 33

14.3 Sparrow Data

Coefficients of linear discriminants:

LD1 LD2 LD3 LD4 LD5 LD6

Flatwing -0.03785084 -0.2737932 0.11491239 -0.42823580 0.2914571 -0.00853429

tarsus 0.24588046 1.6722403 -0.29390262 -0.32513644 0.4096320 0.23987305

head -0.68101337 -0.6089259 -1.48925091 0.92558503 0.7035545 -0.27392299

culmen -1.61462537 0.3820697 1.47757092 0.41349515 0.2853035 0.30302032

nalospi 2.91308820 -0.1256085 0.73034603 0.04576332 0.4617898 -0.03379820

wtnew 0.04825864 -0.1474904 -0.04105511 0.11254511 -0.5204707 0.67584015

Proportion of trace:

LD1 LD2 LD3 LD4 LD5 LD6

0.6776 0.2034 0.0923 0.0158 0.0087 0.0022

Unstandardized discrimination coefficients

27 of 33

14.3 Sparrow Data

28 of 33

14.3 Sparrow Data

Nalospi: measure of distance between the front of the nostril and the tip of the beak

29 of 33

14.3 Sparrow Data

Df Wilks approx F num Df den Df Pr(>F)

observer 1 0.9743 4.8138 6 1095 7.366e-05 ***

Residuals 1100

Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1

Df Pillai approx F num Df den Df Pr(>F)

observer 1 0.0257 4.8138 6 1095 7.366e-05 ***

Residuals 1100

Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1

Df Hotel-Law approx. F num Df den Df Pr(>F)

observer 1 0.0264 4.8138 6 1095 7.366e-05 ***

Residuals 1100

Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1

All indicate that there is a significant group effect (observer)

30 of 33

14.3 Sparrow Data

Classification

  • observer 1 2 3 4 5 6 7 8 %
  • 1 7 29 9 0 0 9 1 1 12.50
  • 2 1 284 34 0 1 3 6 3 85.54
  • 3 0 70 172 0 2 27 0 0 63.47
  • 4 0 64 3 2 1 3 0 0 2.74
  • 5 0 21 23 0 7 3 0 0 12.96
  • 6 0 29 57 0 0 47 0 2 34.81
  • 7 0 47 12 0 3 0 5 0 7.46
  • 8 0 61 41 0 1 5 0 6 5.26

31 of 33

Palmer penguins

31

https://allisonhorst.github.io/palmerpenguins/articles/intro.html

32 of 33

Summary

  • Discriminant Analysis uses eigenanalysis of the pooled within-groups covariance matrix to identify linear functions of multiple variables that maximize group differences.
  • The most discriminating variables are indicated by discriminant coefficients (standardized for variable scale).
  • Classification accuracy can be evaluated from cross-validation.
  • Significance is less important than classification accuracy.

Matrix Algebra

32

Advanced Stats

33 of 33

Additional Reading

  • Huberty, CJ. 1994. Applied Discriminant Analysis. New York: Wiley
  • Gotelli, N.J. and Ellison, A.M. 2004. A Primer of Ecological Statistics. Sunderland, MA: Sinauer Associates, Inc. Publisher
  • Klecka, W.R. 1980. Discriminant Analysis. Quantitative Applications in the Social Sciences Series, No. 19. Thousand Oaks, CA: Sage Publications.
  • Legendre, P., and Legendre, L. 1998. Numerical ecology. 2nd English ed. Developments in environmental modelling 20.
  • Tabachnick, B.G., and Fidell, L.S. 2007. Using multivariate statistics Boston. MA: Allyn and Bacon.

Matrix Algebra

33

Advanced Stats