1 of 21

Using Partitional Clustering Methods to Describe Data

  • Focus: Experiments in WEKA (K-Means)
  • Course: Databases and Data Mining
  • Instructor: Jamolbek Mattiev

2 of 21

Learning Objectives

  • • Understand partitional clustering
  • • Apply K-Means in WEKA
  • • Interpret clustering results
  • • Compare clustering configurations

3 of 21

What is Clustering?

  • • Unsupervised learning
  • • Groups similar data objects
  • • No predefined class labels
  • • Data exploration tool

4 of 21

Partitional Clustering Methods

  • • K-Means
  • • K-Medoids
  • • CLARA
  • • CLARANS

5 of 21

K-Means Algorithm Overview

  • • Select K initial centroids
  • • Assign points to nearest centroid
  • • Update centroids (mean)
  • • Repeat until convergence

6 of 21

Objective Function

  • Minimize Within-Cluster Sum of Squares (WCSS)
  • WCSS = Σ ||x - μ||²
  • Goal: Compact and well-separated clusters

7 of 21

Distance Metric

  • • Euclidean distance (default)
  • • Important for numeric data
  • • Sensitive to scale → normalization required

8 of 21

Step 1: Open WEKA Explorer

  • • Launch WEKA
  • • Click Explorer
  • • Load dataset (.arff)
  • • Go to Cluster tab

9 of 21

Step 2: Select K-Means

  • • Click Choose
  • • Select SimpleKMeans
  • • Set number of clusters (K)

10 of 21

Step 3: Configure Parameters

  • • Set number of clusters (e.g., 3)
  • • Distance function: Euclidean
  • • Initialization method
  • • Seed value

11 of 21

Step 4: Run Clustering

  • • Click Start
  • • Observe cluster centroids
  • • Check cluster sizes
  • • Analyze assignments

12 of 21

Interpreting Output

  • • Centroid values
  • • Cluster distribution
  • • Within-cluster sum of squared errors
  • • Visualization option

13 of 21

Choosing Optimal K

  • • Elbow Method
  • • Silhouette Score
  • • Domain knowledge
  • • Try multiple K values

14 of 21

Experiment Design

  • • Run K = 2 to 10
  • • Record WCSS values
  • • Identify elbow point
  • • Compare cluster compactness

15 of 21

Comparing Partitional Methods

  • K-Means: Fast, centroid-based
  • K-Medoids: Robust to outliers
  • CLARA: Large datasets
  • Choose based on data characteristics

16 of 21

Advantages of K-Means

  • • Simple and efficient
  • • Scales to large datasets
  • • Easy implementation in WEKA

17 of 21

Limitations

  • • Sensitive to outliers
  • • Requires predefined K
  • • Sensitive to initialization

18 of 21

Practical Exercise

  • • Apply K-Means on dataset
  • • Test different K values
  • • Compare cluster sizes
  • • Discuss interpretation

19 of 21

Discussion Questions

  • • How does K affect clusters?
  • • What happens with poor initialization?
  • • When prefer K-Medoids?

20 of 21

Summary

  • • K-Means is core partitional method
  • • WEKA provides easy experimentation
  • • Proper K selection is crucial
  • • Compare methods for better understanding

21 of 21

Elbow Method Visualization (Example)