1 of 47

COVID-19 Statistical Analysis & Visualisation of Data

Harshitha T S

2 of 47

COVID-19 Pandemic

2

3 of 47

COVID-19 Patients Data Analysis & Surveillance Map

3

4 of 47

INTRODUCTION

4

  • COVID-19 Pandemic is growing day by day.
  • Here, we are analyzing the data from 30/01/2020 to 30/04/2020.
  • We are trying to find out which group gets more affected by this disease.
  • We also analyze district and city wise diseased patients data.
  • Cases are arising day by day which makes it impossible to make proper conclusions from them.
  • Here, we are resolving this issue.

5 of 47

How can we do this?

5

  • Maps allow us to analyze the data across several dimensions in a single visualization.
  • Based on the collected data, this Python Folium code produces a google map showing the location of patients and their important details can be viewed.
  • We get the details about the patients in a particular region by zooming the map.
  • We can analyse the data based on the various statistical measures.
  • Visualisations using Tableau help us to visually analyse the various aspects of data.

6 of 47

Major procedures

6

  • Data Collection
  • Data Pre Processing
  • Data Analysis using ‘Statistics’ library in Python
  • Geo Spatial Data Visualization using Python Folium
  • Visualization of data in Tableau
  • Creating Interactive Dashboard
  • Gain useful inferences

7 of 47

DETAILED DESCRIPTION

7

  • This project views the data from three different angles:
  • Statistical Analysis using different functions in Python.
  • Data Visualisation using Tableau
  • Data Visualisation using Python Folium

8 of 47

Before Pre-Processing

8

  • The dataset has 295 rows and 15 columns.
  • There were null and missing values in the dataset.

9 of 47

DATA PRE-PROCESSING

9

Data Collection

RAW DATA collected from government websites

Data Cleansing

Replacing the null or missing values with mean value by using Excel software

Data Transformation

Adding the Latitude and Longitude column to transform the data into appropriate forms suitable for mining.

Data Editing

Adjustment of collected data by adjusting their column names

Data Wrangling

Converting the xlsx to csv format

10 of 47

After Pre-Processing

10

  • The dataset has 295 rows and 17 columns.
  • Two columns namely latitude and longitude were added to the dataset . By using Google Maps, we identified the latitudinal and longitudinal values based on the detected city values.

11 of 47

Column analysis

11

ID

Ordinal data

Govt ID

Ordinal data

Diagnosed date

Ordinal data

Age

Quantitative data

Gender

Nominal data

Detected city

Nominal data

Detected District

Nominal data

Latitude

Interval data

Longitude

Interval data

Detected State

Nominal data

Hospital name

Nominal data

Nationality

Nominal data

Status

Ordinal data

Curing Rate

Ordinal data

12 of 47

Statistical Measures

12

13 of 47

‘Statistics’ Library

13

  • It is a built-in statistics library for descriptive statistics.
  • They provide the functions to mathematical statistics of numeric data.
  • The Statistics module is new in Python 3.4
  • There are some popular statistical functions introduced in this module.

14 of 47

IMPLEMENTING LIBRARIES AND IMPORTING DATA

14

15 of 47

IMPLEMENTING LIBRARIES AND IMPORTING DATA

15

16 of 47

Measures of

Central Tendency

16

17 of 47

MEAN

MEDIAN

17

  • statistics.mean()
  • Used to calculate the arithmetic mean of the numbers in the list.
  • We can find the average age value for whom the disease gets affected.
  • Example:

Mean_value = statistics.mean(x)

  • statistics.median()
  • Used to return the middle value of the numeric data in the list.
  • We can find the middle age value for whom the disease gets affected.
  • Example:

Median_value = statistics.median(x)

18 of 47

MODE

18

  • statistics.mode()
  • Used to the most common that that occurs in the list.
  • We can find the most common age value for whom the disease gets affected.
  • If there is more than on modal value, multimode() is used.
  • Example:

Mode_value = statistics.mode(x)

Mode_value = statistics.multimode(x)

19 of 47

Measures of

Dispersion

19

20 of 47

RANGE

PERCENTILE

20

  • min() and max()
  • Used to identify the range in which the values lie.
  • We can find the age group getting affected to this COVID-19 virus using Range.
  • Example:

Range = max(x) - min(x)

  • statistics.quantiles()
  • Used to divide the data into 3 equal parts.
  • We can find the semi variation between the age group getting affected to this COVID-19 virus using Quartile Deviation.
  • Example:

Percentiless = statistics.quantiles(x, n = 3)

21 of 47

SKEWNESS

STANDARD DEVIATION

21

  • skew()
  • Measures the asymmetry of the data sample.
  • We can identify the pattern the data follows.
  • Example:

.skew(age)

  • statistics.stdev()
  • How far are th data points from the center of the data?
  • How spread out are values from the middle getting affected to the virus? This can be found out using Standard Deviation.
  • Example:

Stdev = statistics.stdev(x)

22 of 47

VARIANCE

22

  • statistics.variance()
  • How far are the data points from the mean?
  • To identify the measure of variability of age values.
  • Average degree to which each point differs from the mean can be analysed using variance here by considering the column age.
  • Example:

Var = statistics.variance(age)

23 of 47

DATA ANALYSIS

23

24 of 47

DATA ANALYSIS

24

25 of 47

Visualisation using Tableau

25

26 of 47

Analysis using Tableau

26

  • It is based on Business Intelligence
  • It can connect to spatial information.
  • We can create an interactive dashboard which depicts the trends, variations and density of the data in the form of graphs and charts.
  • We can create bar charts, pie charts, line charts based on the detected state, district or city.
  • We can create waterfall charts to analyse the flow of the pandemic in certain regions.

27 of 47

Dashboard 1

27

28 of 47

Dashboard 1

28

29 of 47

Line Graph

29

30 of 47

30

31 of 47

31

32 of 47

32

33 of 47

33

34 of 47

34

35 of 47

35

36 of 47

Visualisation using Python Folium

36

37 of 47

Python visualisation

37

  • Folium library is used for analysing geo spatial data.
  • We can plot an interactive map using this library.
  • We can take a glimpse of the map to get useful insights from them.

38 of 47

Building up the application

38

  • First, we collect the updated COVID-19 patient's data from the government agencies or other available resources.
  • Then we convert the data to CSV format.
  • This Python Folium code will convert this CSV data set into a map object and that will be rendered and visualized in a google map.
  • From that file, latitude and longitude values of each patient are taken, and using the marker method is marked in the map.
  • Marker method has many attributes, using the color attribute we can give color according to the patient’s condition so that we can easily identify whether he/she is hospitalized, recovered, or isolated.
  • By using other statistical measures we can analyse which age groups are getting more affected to this disease.

39 of 47

39

40 of 47

40

41 of 47

41

42 of 47

42

43 of 47

43

44 of 47

44

45 of 47

INSIGHTS

45

  • Maximum age value to whom the disease is affected is 99.
  • Minimum age value to the disease is affected is 20.
  • Range of the disease getting affected is 35 to 45.
  • From the dashboard, we can gain useful insights.
  • By using Packed plots; Line Maps, we can see that District Hospital Chemmatamvayal has the most recorded patients and Kasargod is the most affected District. Secondly stand Kuttyadi. Third stands Kochi.
  • From that dashboard we have options to choose a particular district so that we can view its summary details.

46 of 47

What to do next?

46

  • Focus on those age groups.
  • Focus on those regions.
  • To reduce the spread of the pandemic.
  • Thus we can lead a better future.

Corona has become a part of our lives.

47 of 47

THANK YOU!

47