1 of 11

Downloading Datasets from the UCI Machine Learning Repository

Step-by-step guide for machine learning experiments

Dr. Jamolbek Mattiev

2 of 11

What is the UCI ML Repository?

  • A public collection of datasets for ML and data mining
  • Maintained by the University of California, Irvine
  • Widely used in education and research

3 of 11

Why Use UCI Datasets?

  • Free and publicly available
  • Well-documented datasets
  • Suitable for WEKA, Python, R
  • Standard benchmarks for ML algorithms

4 of 11

Step 1: Access the Repository

  • Open a web browser
  • Search for “UCI Machine Learning Repository”
  • Open the official website

5 of 11

6 of 11

Step 2: Browse Datasets

  • Browse datasets by name
  • Filter by:
    • Task (classification, regression)
    • Data type
    • Domain
  • View dataset summaries

7 of 11

8 of 11

Step 3: Select a Dataset

Dataset page includes:

  • Abstract and description
  • Attribute information
  • Number of instances
  • Data source

9 of 11

Step 4: Download the Dataset

  • Scroll to Data Folder / Download section
  • Download files:
    • CSV
    • TXT
    • ZIP
  • Save to local computer

Common File Formats:

CSV – easy to import into WEKA; TXT / DAT – text-based datasets; ZIP – compressed files with documentation

10 of 11

Preparing Data for WEKA

  • Before using in WEKA:
  • Extract ZIP files
  • Check missing values
  • Convert CSV → ARFF if needed
  • Set class attribute correctly

Popular UCI Datasets: Iris, Breast Cancer, Wisconsin, Wine, Diabetes

11 of 11

Summary

  • Visit UCI repository
  • Choose a dataset
  • Download dataset files
  • Prepare data for WEKA analysis