1 of 26

Data Mining_Anoop Chaturvedi

1

Swayam Prabha

Course Title

Multivariate Data Mining- Methods and Applications

Lecture 28

Hierarchical Clustering Techniques

By

Anoop Chaturvedi

Department of Statistics, University of Allahabad

Prayagraj (India)

Slides can be downloaded from https://sites.google.com/view/anoopchaturvedi/swayam-prabha

2 of 26

  •  

Data Mining_Anoop Chaturvedi

2

Individuals

Scores

1

1

2

3

4

5

6

2

1

2

3

4

5

7

3

5

10

15

20

25

30

3 of 26

  •  

Data Mining_Anoop Chaturvedi

3

4 of 26

  •  

Data Mining_Anoop Chaturvedi

4

5 of 26

Example: The following table gives the proportions of two populations of red Campion:

Data Mining_Anoop Chaturvedi

5

Character State

Population

Character State

Population

A

B

A

B

Corolla Color

Not as petals, pink

0.01

0.15

Pink

0.95

0.80

Not as petals, white

0.14

0.10

White

0.05

0.20

Red calyx pigment

Coronal Scale

Present

0.80

0.60

As petals, pink

0.85

0.75

Absent

0.20

0.40

6 of 26

  •  

Data Mining_Anoop Chaturvedi

6

7 of 26

  •  

Data Mining_Anoop Chaturvedi

7

8 of 26

  •  

Data Mining_Anoop Chaturvedi

8

Iris Data: Cluster Dendrogram

9 of 26

  •  

Data Mining_Anoop Chaturvedi

9

10 of 26

  •  

Data Mining_Anoop Chaturvedi

10

11 of 26

  •  

Data Mining_Anoop Chaturvedi

11

12 of 26

  •  

Data Mining_Anoop Chaturvedi

12

13 of 26

  •  

Data Mining_Anoop Chaturvedi

13

14 of 26

  •  

Data Mining_Anoop Chaturvedi

14

15 of 26

Data Mining_Anoop Chaturvedi

15

Stage

Groups

[1],[2],[3],

[4],[5]

[1,2],[3],[4],

[5]

[1,2],[3],

[4,5]

[1,2],

[3,4,5]

[1,2,3,4,5]

Smallest distance

2.0

3.0

4.0

5.0

16 of 26

Data Mining_Anoop Chaturvedi

16

17 of 26

Shortcomings of single linkage clustering

  • In order to merge two groups, only need one pair of points to be close. Clusters can be too spread out. Can’t discern poorly separated clusters.
  • Suffers from the Chaining Effect where clusters are extended by connecting individual points that are close to each other but may not necessarily belong to the same cluster.
  • Sensitive to outliers. Even a single outlier can significantly affect the clustering results by connecting distant points into the same cluster.

Data Mining_Anoop Chaturvedi

17

18 of 26

  • If the density of points within the cluster may vary or there may be empty spaces between clusters, it struggles with identifying clusters.
  • Can lead to fragmented clusters, especially in datasets with varying densities or non-convex shapes.
  • Choice of distance metric can greatly influence the clustering results.
  • Does not provide a clear criterion for determining the optimal number of clusters.

Data Mining_Anoop Chaturvedi

18

19 of 26

  •  

Data Mining_Anoop Chaturvedi

19

20 of 26

Shortcomings of complete linkage clustering

  • Avoids chaining, but suffers from crowding. Its score is based on the worst-case dissimilarity between pairs, a point can be closer to points in other clusters than to points in its own cluster. Clusters are compact, but not far enough apart.
  • Tends to produce clusters of approximately equal diameter. Struggle with clusters that have significantly different sizes or densities.
  • Tends to produce compact clusters, but is not suitable for detecting irregularly shaped clusters.

Data Mining_Anoop Chaturvedi

20

21 of 26

  • Sensitivity to Outliers.
  • It may merge clusters that are connected by a few points even if the majority of points within each cluster are far apart.
  • Tends to suffer from the "cluster fusion problem," where small, well-separated clusters may be merged into larger clusters due to the presence of a few points that are relatively close to each other.
  • Subjectivity in Determining Number of Clusters.

Data Mining_Anoop Chaturvedi

21

22 of 26

  •  

Data Mining_Anoop Chaturvedi

22

23 of 26

Shortcomings of average linkage

  • Not has simple interpretations like Single and complete linkage trees.
  • Sensitivity to Outliers.
  • Tendency to produce Clusters of approximately equal sizes, resulting in merging of distinct clusters or the fragmentation of single clusters.
  • Difficulty in handling Non-Globular Clusters, where the distance between points within a cluster is not uniform.
  • Dependency on Distance Metric.
  • Inability to Detect Complex Structures such as hierarchical relationships or nested clusters within larger clusters.
  • Subjectivity in Determining Number of Clusters

Data Mining_Anoop Chaturvedi

23

24 of 26

  •  

Data Mining_Anoop Chaturvedi

24

Individual

1

2

3

4

5

Variable 1

1

1

6

8

8

Variable 2

1

2

3

2

0

25 of 26

  •  

Data Mining_Anoop Chaturvedi

25

26 of 26

Single linkage Average linkage Complete linkage

Data Mining_Anoop Chaturvedi

26