1 of 22

(K-Medoids) Using Partitional Clustering Methods to Describe Data

  • Focus: Experiments in WEKA (K-Medoids)
  • Course: Databases and Data Mining
  • Instructor: Jamolbek Mattiev

2 of 22

Learning Objectives

  • • Understand K-Medoids clustering
  • • Perform experiments in WEKA
  • • Analyze clustering cost
  • • Compare with K-Means

3 of 22

What is Partitional Clustering?

  • • Divides dataset into K clusters
  • • Each object belongs to one cluster
  • • Optimizes objective function

4 of 22

K-Medoids Overview

  • • Similar to K-Means
  • • Uses real data points as medoids
  • • More robust to noise and outliers

5 of 22

Medoid Definition

  • • Central representative object
  • • Minimizes total distance within cluster
  • • Must be actual dataset instance

6 of 22

Objective Function

  • Minimize total dissimilarity:
  • Cost = Σ distance(point, medoid)
  • Lower cost → better clustering

7 of 22

Algorithm Steps (PAM)

  • 1. Select K initial medoids
  • 2. Assign points to nearest medoid
  • 3. Try swapping medoids
  • 4. Keep swap if cost decreases
  • 5. Repeat until stable

8 of 22

Distance Measures

  • • Manhattan distance (common)
  • • Euclidean distance
  • • Any valid dissimilarity measure

9 of 22

Step 1: Open WEKA Explorer

  • • Launch WEKA
  • • Click Explorer
  • • Load dataset (.arff)
  • • Go to Cluster tab

10 of 22

Step 2: Select K-Medoids

  • • Click Choose
  • • Select SimpleKMedoids (if available)
  • • Or use PAM implementation

11 of 22

Step 3: Configure Parameters

  • • Set number of clusters (K)
  • • Choose distance function
  • • Set seed value
  • • Configure max iterations

12 of 22

Step 4: Run Clustering

  • • Click Start
  • • Observe medoids
  • • Check cluster sizes
  • • Record total cost

13 of 22

Experiment Design

  • • Run K from 2 to 10
  • • Record total cost
  • • Identify optimal K
  • • Compare cluster compactness

14 of 22

Interpreting Output

  • • Selected medoids
  • • Cluster assignments
  • • Total within-cluster cost
  • • Distribution of instances

15 of 22

Comparison: K-Means vs K-Medoids

  • K-Means → mean centroid
  • K-Medoids → actual data point
  • K-Medoids more robust to outliers

16 of 22

Advantages of K-Medoids

  • • Robust to noise
  • • Works with arbitrary distances
  • • Interpretable representatives

17 of 22

Limitations

  • • Higher computational cost
  • • Less scalable than K-Means
  • • Sensitive to initial medoids

18 of 22

Advanced Variants

  • • CLARA (large datasets)
  • • CLARANS (randomized search)

19 of 22

Discussion Questions

  • • When prefer K-Medoids over K-Means?
  • • How does distance metric affect results?

20 of 22

Summary

  • • K-Medoids minimizes total dissimilarity
  • • Uses real data points as centers
  • • WEKA allows experimental comparison
  • • Suitable for noisy datasets

21 of 22

Cost vs Number of Clusters (Example)

22 of 22

Cost Reduction Over Iterations (Example)