1 of 22

Evaluating and Comparing ML Algorithms in WEKA

  • Course: Databases and Data Mining
  • Topic: Model Evaluation & Statistical Testing
  • Instructor: Jamolbek Mattiev

2 of 22

Learning Objectives

  • • Compare multiple algorithms
  • • Evaluate performance on different datasets
  • • Perform statistical significance testing
  • • Interpret p-values correctly

3 of 22

Why Algorithm Comparison Matters

  • • Different algorithms perform differently
  • • Dataset characteristics matter
  • • Avoid biased conclusions
  • • Select best generalizable model

4 of 22

Evaluation Metrics

  • • Accuracy
  • • Precision
  • • Recall
  • • F1-score
  • • ROC-AUC

5 of 22

Validation Methods in WEKA

  • • Percentage split
  • • 10-fold cross-validation
  • • Repeated cross-validation

6 of 22

Experimental Design

  • • Select multiple datasets
  • • Apply same validation method
  • • Use identical preprocessing
  • • Record metrics for each algorithm

7 of 22

Algorithms to Compare (Example)

  • • J48 (Decision Tree)
  • • Naive Bayes
  • • k-Nearest Neighbors
  • • Support Vector Machine

8 of 22

Running Experiments in WEKA

  • • Open Explorer
  • • Load dataset
  • • Select Classify tab
  • • Choose algorithm and validation

9 of 22

Recording Results

  • Create comparison table:
  • Dataset | Algorithm | Accuracy | F1 | AUC
  • Repeat for multiple datasets

10 of 22

Need for Statistical Testing

  • • Differences may occur by chance
  • • Must test statistical significance
  • • Avoid misleading comparisons

11 of 22

Paired t-Test Concept

  • • Compare two algorithms
  • • Uses paired results across folds/datasets
  • • Null hypothesis: no difference

12 of 22

p-value Interpretation

  • • p < 0.05 → statistically significant
  • • p ≥ 0.05 → not significant
  • • Smaller p → stronger evidence

13 of 22

Using WEKA Experimenter

  • • Open Experimenter
  • • Add datasets
  • • Add algorithms
  • • Select cross-validation
  • • Run experiment

14 of 22

Performing Significance Test in WEKA

  • • Go to Analyse tab
  • • Select baseline algorithm
  • • Choose statistical test (paired t-test)
  • • Interpret output table

15 of 22

Multiple Dataset Comparison

  • • Evaluate across many datasets
  • • Use average ranks
  • • Consider non-parametric tests

16 of 22

Wilcoxon Signed-Rank Test

  • • Non-parametric alternative
  • • No normality assumption
  • • Suitable for multiple datasets

17 of 22

Friedman Test

  • • Compare multiple algorithms
  • • Rank-based method
  • • Followed by post-hoc tests

18 of 22

Common Mistakes

  • • Comparing single run results
  • • Ignoring variance
  • • Misinterpreting p-values

19 of 22

Best Practices

  • • Use cross-validation
  • • Report mean ± standard deviation
  • • Perform statistical testing
  • • Compare across multiple datasets

20 of 22

Summary

  • • Proper evaluation is essential
  • • Statistical testing validates conclusions
  • • WEKA Experimenter simplifies comparison
  • • Always report reproducible experiments

21 of 22

Algorithm Accuracy Comparison (Example)

22 of 22

Statistical Significance Results (Example)