1 of 22

Data Processing with Python Libraries NumPy and Pandas

REU Site: Applying Data Science on Energy-efficient Cluster Systems and Applications

Training Material

​

2 of 22

Python Libraries - NumPy

REU Site: Applying Data Science on Energy-efficient Cluster Systems and Applications

  • NumPy (Numerical Python) is a fundamental library for scientific computing in Python, providing support for large, multi-dimensional arrays and matrices and a collection of mathematical functions to operate on them.
  • With its high-performance capabilities, NumPy is instrumental for data analysis, machine learning, and scientific computations.
  • The official NumPy website is a comprehensive resource, providing information on installation, documentation and various tutorials.
    • https://numpy.org/

3 of 22

NumPy - Examples (1)

REU Site: Applying Data Science on Energy-efficient Cluster Systems and Applications

  • numpy.array()
    • function is used to create a new array object that produces a specific result.
  • numpy.add(), numpy.subtract(), numpy.multiply(), numpy.divide()
    • functions use for arithmetic operations (addition, subtraction, multiplication, division, etc.).
  • numpy.mean(), numpy.median(), and numpy.std()
    • statistical functions calculating the mean, median, and standard deviation of an array.

4 of 22

NumPy - Examples (2)

REU Site: Applying Data Science on Energy-efficient Cluster Systems and Applications

  • numpy.random
    • module is crucial for generating random numbers and offers several functions for generating random data.
    • np.random.rand(n, m)
      • Generate a nxm array of random floats between 0 and 1
    • np.random.randint(n)
      • Generate a random integer between 0 and n

5 of 22

Python Libraries - Pandas

REU Site: Applying Data Science on Energy-efficient Cluster Systems and Applications

  • Pandas is a high-level data manipulation tool, providing data structures and functions needed to manipulate structured data.
  • Its key data structure, the DataFrame, enables the handling of tabular data in rows of observations and columns of variables.
  • The official Pandas website offers extensive resources for users including a getting started guide, detailed API reference, and various tutorials for specific tasks.
    • https://pandas.pydata.org/

6 of 22

Pandas - Examples

REU Site: Applying Data Science on Energy-efficient Cluster Systems and Applications

  • pandas.DataFrame()
    • Construct a DataFrame from lists, dictionaries, or other dataframes
  • Data Selection:
    • Pandas provides various ways to slice, dice, and select your data.
    • For instance, df['Name'] will return the 'Name' column, and df[df['Age'] > 25] will select all rows where 'Age' is greater than 25.
  • Data Aggregation:
    • For example, functions like groupby() allows for complex data aggregation tasks.

7 of 22

Pandas - Other Functionality

REU Site: Applying Data Science on Energy-efficient Cluster Systems and Applications

  • Handling Missing Data
    • Pandas provides the isnull(), notnull(), and dropna() methods for detecting, removing, or replacing null (NaN) values.
  • Data Merging and Concatenation:
    • You can merge multiple datasets using the concat(), merge(), and join() functions.

8 of 22

Pandas - Other Functionality (2)

REU Site: Applying Data Science on Energy-efficient Cluster Systems and Applications

  • Date Functionality:
    • Pandas provides robust tools for working with dates, times, and time-indexed data.
  • Setting Index:
    • Function such as set_index() can be used to set the DataFrame index using one or more existing columns.
  • Reading/Writing to Files:
    • Pandas provides functionality for reading and writing data to a variety of formats like CSV, Excel, SQL databases, and more.

9 of 22

Pandas - Other Functionality (3)

REU Site: Applying Data Science on Energy-efficient Cluster Systems and Applications

  • Slicing, Dicing, and Selecting Data
    • Using loc[] and iloc[]: You can select data using labels (loc[]) or integer-based location (iloc[]).
    • For example:
      • df.loc[:, 'sepal length in cm'] will select all rows for the column 'sepal length in cm', and df.iloc[0:5, :] will select the first five rows.
    • More information can be found here: https://pandas.pydata.org/docs/user_guide/indexing.html

10 of 22

Alternative to installing Jupyter Notebook (1)

REU Site: Applying Data Science on Energy-efficient Cluster Systems and Applications

  • For the following practice exercises, if Anaconda and Jupyter Notebook are not install in your local machine, another option is to use JupyterLab through the website
    • https://jupyter.org/
  • You may use it for free without installation by selecting ‘Try it in your browser’

​

11 of 22

Alternative to installing Jupyter Notebook (2)

REU Site: Applying Data Science on Energy-efficient Cluster Systems and Applications

  • Select JupyterLab from the options

​

  • Select ‘+’

12 of 22

Alternative to installing Jupyter Notebook (3)

REU Site: Applying Data Science on Energy-efficient Cluster Systems and Applications

  • Select Python
  • Drag and drop the data set .csv file provided into the notebooks section.

​

13 of 22

Alternative to installing Jupyter Notebook (4)

REU Site: Applying Data Science on Energy-efficient Cluster Systems and Applications

  • This will allow you to practice the following exercises without the Jupyter Notebook installation

​

14 of 22

Practice using Jupyter Notebook - 1

REU Site: Applying Data Science on Energy-efficient Cluster Systems and Applications

  • Download Dataset from Training Materials

​

  • Create a new Python Notebook and name your Notebook

​

  • Add the dataset into the same location

​

15 of 22

Practice using Jupyter Notebook - 2

REU Site: Applying Data Science on Energy-efficient Cluster Systems and Applications

  • Import the NumPy and Pandas Libraries

​

  • use the Pandas read_csv() function to import the dataset into a variable called data_set

​

  • Convert the dataset into a dataframe object named df using the Pandas function DataFrame()

​

16 of 22

Practice using Jupyter Notebook - 3

REU Site: Applying Data Science on Energy-efficient Cluster Systems and Applications

  • Display the dataframe object
    • Beginning portion and ending portion of the dataframe will display
  • Call df.head()to display only the first five rows in the dataframe

​

  • Call df.tail()to display only the last five rows of the dataframe

17 of 22

Practice using Jupyter Notebook - 4

REU Site: Applying Data Science on Energy-efficient Cluster Systems and Applications

  • Call df.describe()to display additional information including count, mean, standard deviation, minimum and maximum, 25th percentile, median (50th percentile), and 75th percentile.

​

  • Call the Pandas mean() function to display the mean for the column ‘sepal length in cm’

18 of 22

Practice using Jupyter Notebook - 5

REU Site: Applying Data Science on Energy-efficient Cluster Systems and Applications

Give it a try:

​

​

  • Use the NumPy library to calculate and display the mean for the column ‘sepal length in cm’
  • Hint

19 of 22

Practice using Jupyter Notebook - 6

REU Site: Applying Data Science on Energy-efficient Cluster Systems and Applications

Solution:

​

​

  • Use the NumPy library to calculate and display the mean for the column ‘sepal length in cm’

​

​

​

20 of 22

Practice using Jupyter Notebook - 7

REU Site: Applying Data Science on Energy-efficient Cluster Systems and Applications

​

​

  • Call df['class'].value_counts() to count the number of rows per ‘class’ feature.

21 of 22

Resources - NumPy and Pandas

REU Site: Applying Data Science on Energy-efficient Cluster Systems and Applications

  • The official NumPy website is a comprehensive resource, providing information on installation, documentation and various tutorials.
    • https://numpy.org/

​

  • The official Pandas website offers extensive resources for users including a getting started guide, detailed API reference, and various tutorials for specific tasks.
    • https://pandas.pydata.org/

22 of 22

Progress Presentation/Report

REU Site: Applying Data Science on Energy-efficient Cluster Systems and Applications

  • For the rest of the time, break 10:10am - 10:15am and after please continue to work on finishing up the practice exercise and then the weekly progress presentation or report (however your grad mentor recommends). 10:15am - 11am.