1 of 179

EE 629 �Internet of Things �Lesson 8: Data Analysis

Kevin W. Lu

2022-10-24

2 of 179

Lesson 8: Data Analysis

  • Overview of data science
    • Characteristics of big (fast) data
    • Data analysis and visualization
  • Python for data analysis
    • NumPy array
    • Matplotlib pyplot, histogram, and boxplot
    • SciPy stats linear regression
    • SciPy interpolation
    • Scikit-learn cross-validation prediction
    • Load files of comma-separated values (CSV) into Pandas
    • Pandas groupby
    • Write Pandas dataframe to CSV files
    • TensorFlow and Keras

2

3 of 179

Lab 8A — Examples

  • Install Python packages on laptop or Raspberry Pi
    • numpy, scipy, matplotlib, pandas, scikit-learn, tensorflow, keras
  • Go to ~/iot, run git pull, and run the following examples in ~/iot/lesson8
    • ticklabels_demo_rotation.py
    • legend_demo.py
    • histogram_demo_features.py, histogram_demo_extended.py
    • boxplot_demo.py
    • linreg.py, interpolation.py
    • plot_lda.py, plot_lda_qda.py
    • plot_cv_predict.py, plot_cv_diabetes.py
    • keras_diabetes.py, keras_first_network.py
  • Copy train.csv, test.csv, titanic_1.py, and titanic_2.py to ~/demo
    • Run both and view predictions at result.csv and result2.csv

3

4 of 179

Lab 8B — Data Analysis

  • Open the Google sheet from Lab 7B that collected Raspberry Pi date/time and CPU usage/temperature
  • Insert four charts including time series, two histograms, and a scatter plot with a linear trendline
  • Save the sheet in CSV format
  • Edit ~/iot/lesson8/plt_final.py and plt_cv2.py to read the CSV file and show seven figures including time series, two histograms, two box plots, a scatter plot with a linear regression line, and cross-validation prediction with temperature as target
  • Include required title/labels, and add legend or adjust ticks as needed

4

5 of 179

IoT Value Chain

5

Components,

Sensors,

Semiconductors

Things

Connectivity,

Infrastructure,

Gateway

Software,

Platforms,

Analytics

Services

  • IoT is not about adding connectivity to all things
  • IoT is about how sensors, devices, things, and services can be integrated to create value
  • Value is derived from making sense of data, turning it into knowledge and meaningful action

6 of 179

Complexity Levels of IoT Systems

6

Level

Node

Analysis

Storage

Example

1

Single

Local

Local

Home Automation

2

Single

Local

Cloud

Smart Irrigation

3

Single

Cloud

Cloud

Vibration Monitoring

4

Multiple

Local

Cloud

Noise Monitoring

5

Multiple + Coordinator

Cloud

Cloud

Forest Fire detection

6

Multiple +

Centralized

Controller

Cloud

Cloud

Weather Monitoring

7 of 179

Characteristics of Big (Fast) Data

  • Volume: size of data (MB, GB, TB, PB, EB, ZB, and YB)
  • Velocity: speed of data collection, modeling and analysis, monitoring, predictive analytics, and decision making (batch, periodic, near real time, and real time)
  • Variety: form of source data (structured, semi-structured, and unstructured)
  • Variability: difference in data context and meaning
  • Veracity: trustworthiness in accuracy of data and analysis
  • Visualization: readable and understandable graphs
  • Value: actionable information and knowledge from data analysis with insight into events, trends, and correlations*

* correlation is not causation

7

8 of 179

Decimal and Binary Prefixes

8

Decimal (Hard Drive and Bit Rate)

Binary (RAM)

Value

SI Prefix

Value

IEC Prefix

1000

k

kilo

1024

Ki

kibi

10002

M

mega

10242

Mi

mebi

10003

G

giga

10243

Gi

gibi

10004

T

tera

10244

Ti

tebi

10005

P

peta

10245

Pi

pebi

10006

E

exa

10246

Ei

exbi

10007

Z

zetta

10247

Zi

zebi

10008

Y

yotta

10248

Yi

yobi

SI: International System of Units IEC: International Electrotechnical Commission

9 of 179

Power of 10

9

Power of 10

Short Scale

Long Scale

Power of 10

Short Scale

Long Scale

million

106

106

duodecillion

1039

1072

billion

109

1012

tredecillion

1042

1078

trillion

1012

1018

quattuordecillion

1045

1084

quadrillion

1015

1024

quindecillion

1048

1090

quintillion

1018

1030

sexdecillion

1051

1096

sextillion

1021

1036

septendecillion

1054

10102

septillion

1024

1042

octillion

1027

1048

centillion

10303

10600

nonillion

1030

1054

decillion

1033

1060

103003

106000

undecillion

1036

1066

In long scale, 109 is milliard, 1015 is billiard, 1021 is trilliard, etc.

10 of 179

Short and Long Scale Usage

10

11 of 179

Exponential vs. Factorial

11

12 of 179

Hyperoperation

  • The hyperoperation sequence is an infinite sequence of arithmetic operations with Knuth's up-arrow notation introduced by Donald Knuth in 1976

12

Hyper0

1 + b

Hyper1

a + b

Hyper2

a × b

a ∙ b

Hyper3

ab

a ↑ b

Hyper4

ba

a ↑↑ b

Hyper5

a ↑3 b

a ↑↑↑ b

Hyper6

hexation

a ↑4 b

a ↑↑↑↑ b

13 of 179

Other Large Numbers

  • Avogadro constant named after Amedeo Avogadro 1776—1856

6.02214076 x 1023

where 3↑↑↑↑3=3↑↑↑(3↑↑↑3) and 3↑↑↑3=7,625,597,484,987

13

14 of 179

Googol and Googolplex

1920: Prof. Edward Kasner 1878—1955 of Columbia University sought a name for 10100 and one of his nephews, nine-year-old Milton Sirotta 1911—1981 who might have read The Google Book (1913) and/or Barney Google (1919), suggested googol (perhaps, three o's resemble 10100) and googolplex "for writing zeros until one gets tired"

1940: Edward Kasner and James Newman 1909—1966 coauthored Mathematics and the Imagination that introduced googol for 10100 and googolplex for 10googol

1996: Larry Page and Sergey Brin called their search engine "BackRub" for its analysis of the web's backlinks

1997: Larry's officemate, Sean Anderson, suggested a new name Googolplex for BackRub; Larry shortened it to Googol; Sean misspelled Googol as Google in search of the domain name; Larry liked it, and registered Google.com for himself and Sergey on 1997-09-15

14

15 of 179

Google and Googleplex

1913: Vincent Cartwright Vickers 1879—1939 wrote and illustrated The Google Book, a children's book about a strange creature Google and imaginary birds living in Google Land

1919: Billy DeBeck 1890—1942 created a comic strip, Take Barney Google, F'rinstance

1923: Billy Rose 1899—1966 wrote the lyrics for "Barney Google (with the Goo-Goo-Googly Eyes)"

1934: The strip became Barney Google and Snuffy Smith with Snuffy as the main character

1997: Larry Page registered the domain name Google.com instead of Googol.com

2002: The first book that Google Books scanned was Vickers' The Google Book

2004: Google moved headquarters to the Googleplex

15

16 of 179

Baidu

东风夜放花千树,更吹落,星如雨。宝马雕车香满路。凤箫声动,玉壶光转,一夜鱼龙舞。

蛾儿雪柳黄金缕,笑语盈盈暗香去。众里寻他千百度,蓦然回首,那人却在,灯火阑珊处。

  • Baidu, a Chinese company incorporated in 2000, has the world's second largest web search engine
  • The Chinese name Baidu (百度) means a hundred (or countless) times
  • It is a quote from the last line of the classical poem "Green Jade Table in the Lantern Festival" (青玉案·元夕) by Xin Qiji (辛弃疾 1140—1207)
  • "Having searched hundreds of times in the crowd, suddenly turning back, she is there in the dimmest candlelight."

16

17 of 179

Regret Minimization Framework

Jeff Bezos left D. E. Shaw & Co, L.P. and founded Amazon.com, Inc. on 1994-07-05

17

Yes

When I am 80 years old, will I regret not doing this?

Do it

Let it go

No

Idea

18 of 179

Apache Hadoop

The name Hadoop is not an acronym; it's a made-up name. The project's creator, Doug Cutting, explains how the name came about: “The name my kid gave a stuffed yellow elephant. Short, relatively easy to spell and pronounce, meaningless, and not used elsewhere: those are my naming criteria.”

18

19 of 179

ACID Database Transaction Properties

  • Atomicity: Transactions are all or nothing — if one part of the transaction fails, then the entire transaction fails, and the database state is left unchanged
  • Consistency: Only valid data are saved — any transaction will bring the database from one valid state to another
  • Isolation: Transactions do not affect each other — the concurrent execution of transactions results in a system state that would be obtained if transactions were executed serially
  • Durability: Written data will not be lost — once a transaction has been committed, it will remain so, even in the event of power loss, crashes, or errors

Jim Gray, “The transaction Concept: Virtues and Limitations,” June 1981 http://research.microsoft.com/en-us/um/people/gray/papers/theTransactionConcept.pdf

Andreas Reuter and Theo Härder, “Principles of transaction-oriented database recovery,” December 1983

http://web.stanford.edu/class/cs340v/papers/recovery.pdf

19

20 of 179

Data Warehouse vs. Data Lake

20

Data Warehouse

Data Lake or Data Reservoir

Data

Structured, processed

Structured, semi-structured, unstructured

Processing

Schema-on-write

Schema-on-read

Storage

Expensive for large data volumes

Designed for low-cost storage

Agility

Less agile, fixed configuration

Highly agile, configure and reconfigure as needed

Security

Mature

Maturing

Users

Business professionals

Data scientists

21 of 179

Sharding

  • The word shard means a small part of a whole
  • Sharding is a type of database partitioning that separates very large databases into smaller, faster, more easily managed parts called data shards
  • Each shard is held on a separate, independent, self-sufficient database server instance to spread load via a shared nothing (SN) distributed computing architecture

21

22 of 179

Edwin A. Stevens Hall Est. 1870

22

What does this symbol mean?

23 of 179

Pendentive Dome

23

D. Vaccari, “The Barbed Quatrefoil Is a 1500-Year-Old Symbol of Architectural Advance,” 2018. [Online]. Available.

24 of 179

Data Analysis

Data analysis is the process of

  • Inspecting, cleaning, transforming, and modeling data with the goal of discovering useful information, suggesting conclusions, and supporting decision-making
  • Systematically applying statistical and/or logical techniques to describe and illustrate, condense and recap, and evaluate data

"Responsible Conduct in Data Management," Northern Illinois University and the Office of Research Integrity (ORI) of the U.S. Department of Health and Human Services (HHS)

24

25 of 179

Types of Data Analysis

25

Descriptive

Diagnostic

Predictive

Hindsight:

What has happened?

Insight:

Why did it happen?

Foresight:

What could happen?

Prognosis

Diagnosis

Prescriptive

Oversight:

What needs to happen?

"Knowledge is telling the past. Wisdom is predicting the future."

W. Timothy Garvey, M.D.

26 of 179

Black Swan and Gray Rhino

  • "A black swan is an event, positive or negative, that is deemed improbable yet causes massive consequences"
  • "A gray rhino is a highly probable, high impact yet neglected threat: kin to both the elephant in the room and the improbable and unforeseeable black swan"
  • "Gray rhinos are not random surprises, but occur after a series of warnings and visible evidence"

26

27 of 179

Falsifiability vs. Verifiability

  • The law "All swans are white" is falsifiable because "Here is a black swan" contradicts it
  • Karl Popper 1902—1994 noticed that while it is impossible to verify that every swan is white, finding a single black swan shows that not every swan is white
  • Even if it is impossible to both observe a swan and see that its color is black, it would still make the law falsifiable, because observing that a bird is a swan and seeing that a bird is black would still be separately possible
  • See "Dragon in my garage" in The Demon-Haunted World: Science as a Candle in the Dark (1995) book by Carl Sagan 1934—1996

27

28 of 179

Left of Bang

  • "Left of Bang reflects the moments before something bad happens"
  • "It's better to detect sinister intentions early than respond to violent actions late"
  • "There are observable indicators for assessing and baselining individuals, groups, environment, and collective mood"

28

29 of 179

Machine Learning and Deep Learning

  • Machine learning uses statistical techniques to construct a model from observed data rather than users enter specific set of instructions that define the model for observed data
  • Deep learning is a set of techniques that parameterize deep neural network structures with a significant number of layers and parameters
  • Michael Copeland, "What’s the Difference Between Artificial Intelligence, Machine Learning, and Deep Learning?" NVIDIA Blogs, 2016-07-29
  • Forbes 2019 AI 50, e.g., Nuro, Aurora Innovation, Uptake

29

30 of 179

Kaggle

  • Kaggle is an online community of data scientists and machine learners since 2010
  • Google announced its acquisition of Kaggle on 2017-03-08
  • Kaggle provides
    • Competitions: featured, research, getting started, playground, recruitment, annual, and limited participation
    • Datasets: CSV, JSON, SQLite, and BigQuery
    • Kernels: scripts in R or Python, RMarkdown scripts, and Jupyter notebooks
    • Forums: Kaggle Forum, Getting Started, Product Feedback, Questions and Answers, Datasets, and Learn
    • Learn: Python, machine learning, Pandas, data visualization, SQL, R, deep learning, embeddings, and machine learning explainability
  • Google Dataset Search

30

31 of 179

COVID-19 Data

"Are Countries Flattening the Curve for the Coronavirus?" | GitHub repository

31

32 of 179

IEEE DataPort

  • IEEE DataPort is an online data platform created and supported by IEEE
    • Enables users to store, search, access, and manage data
    • Is designed to accept datasets up to 2TB in size and in formats including CSV, TXT, ORC, Avro, RC, XML, SQL, and JSON
    • Provides both downloading capabilities and access to datasets in the Cloud
    • Supports the IEEE overall mission of Advancing Technology for Humanity
  • IEEE DataPort is a universally accessible web-based portal that serves four primary purposes:
    • Enable individuals and institutions to indefinitely store and make datasets easily accessible to a broad set of researchers, engineers, and industry
    • Enable researchers, engineers, and industry to gain access to datasets that can be analyzed to advance technology
    • Facilitate data analysis by enabling access to data in the AWS S3 and by enabling the downloading of datasets
    • Supports reproducible research
  • IEEE Society Members automatically receive a free IEEE DataPort Subscription

32

33 of 179

IEEE DataPort Dataset Categories

  • Artificial Intelligence
  • Astronomy
  • Biomedical and Health Sciences
  • Biophysiological Signals
  • Cloud Computing
  • Communications
  • Computational Intelligence
  • Computer Vision
  • Demographic
  • Ecology
  • Environmental
  • Financial
  • Geoscience and Remote Sensing
  • Image Fusion
  • Image Processing
  • IoT
  • Machine Learning
  • North and South Poles
  • Other
  • Power and Energy
  • Reliability
  • Security
  • Sensors
  • Signal Processing
  • Social Sciences

33

34 of 179

Additional Data Sources

34

35 of 179

Data Visualization

Edward Tufte, The Visual Display of Quantitative Information

Excellence in statistical graphics consists of complex ideas communicated with clarity, precision and efficiency. Graphical displays should:

  • Show the data
  • Induce the viewer to think about the substance rather than about methodology, graphic design, the technology of graphic production or something else
  • Avoid distorting what the data has to say
  • Present many numbers in a small space
  • Make large data sets coherent
  • Encourage the eye to compare different pieces of data
  • Reveal the data at several levels of detail, from a broad overview to the fine structure
  • Serve a reasonably clear purpose: description, exploration, tabulation or decoration
  • Be closely integrated with the statistical and verbal descriptions of a data set

35

36 of 179

A Picture Is Worth 500 Billion Words

36

37 of 179

Popularity

  • Charlie Chaplin 1889—1977 wrote in October 1933 about his latest journey in Germany visiting Albert Einstein 1879—1955
  • “We sat down to delicious home-baked tarts made by Mrs. Einstein. During the course of conversation, his son remarked on the psychology of the popularity of Einstein and myself."
  • “You are popular,” he said, “because you are understood by the masses. On the other hand, the professor’s popularity with the masses is because he is not understood.”
  • Left: Einstein and Chaplin at the City Lights movie premiere in January 1931

37

38 of 179

Prometheus and Frankenstein

  • In Greek mythology, Prometheus is a Titan, who created human from clay and gave fire to humanity, embodying the lone genius whose efforts to improve human existence could also result in overreaching or unintended consequences
  • Prometheus was sentenced and bound to a rock, where each day an eagle, the emblem of Zeus, was sent to feed on his liver
  • On 1818-01-01, Mary Shelley 1797—1851 published Frankenstein; or, The Modern Prometheus, a novel that tells the story of Victor Frankenstein who creates a monster in an scientific experiment, inadvertently endangers his own life and the lives of his family and friends

38

39 of 179

Data Visualization Examples

39

40 of 179

Related Coronavirus Genomes

40

41 of 179

SARS-CoV-2 Phylogenetic Network

41

42 of 179

Coronavirus Infectivility

42

43 of 179

Install Packages on Raspberry Pi

Install SciPy, Matplotlib, pandas, and dependencies on Raspberry Pi

pi@piot4:~ $ sudo apt update

pi@piot4:~ $ sudo apt install python3-scipy

pi@piot4:~ $ sudo apt install python3-matplotlib

pi@piot4:~ $ sudo apt install python3-pandas

pi@piot4:~ $ sudo apt install libopenblas-dev

pi@piot4:~ $ sudo apt install libatlas-base-dev

Install NumPy, scikit-learn, TensorFlow, and Keras on Raspberry Pi

pi@piot4:~ $ pip3 -V

pip 18.1 from /usr/lib/python3/dist-packages/pip (python 3.7)

pi@piot4:~ $ sudo pip3 install -U numpy

pi@piot4:~ $ sudo pip3 install --only-binary :all: -U scikit-learn

pi@piot4:~ $ sudo pip3 install -U tensorflow

pi@piot4:~ $ sudo pip3 install -U keras==2.3.1

43

BLAS: Basic Linear Algebra Subprograms

ATLAS: Automatically Tuned Linear Algebra Software

Contribution by Scott Maslin, 2018 Fall

44 of 179

Install Packages on Laptop

All examples in Lesson 8 can run on a laptop; use conda on Windows, or pip3 on Linux and macOS to install or update Python3 data analysis packages

$ pip3 -V

pip 19.3.1 from /Library/Frameworks/Python.framework/Versions/3.7/lib/python3.7/site-packages/pip (python 3.7)

$ sudo pip3 install -U numpy scipy scikit-learn matplotlib pandas

$ sudo pip3 install -U tensorflow keras

$ sudo pip3 list

44

45 of 179

Numpy Array 0/4

pi@piot4:~ $ python3

>>> import numpy as np

>>> a = np.arange(6)

>>> a

array([0, 1, 2, 3, 4, 5])

>>> print(a)

[0 1 2 3 4 5]

>>>

>>> b = np.arange(12).reshape(4,3)

>>> print(b)

[[ 0 1 2]

[ 3 4 5]

[ 6 7 8]

[ 9 10 11]]

>>> c = np.arange(24).reshape(2,3,4)

>>> print(c)

[[[ 0 1 2 3]

[ 4 5 6 7]

[ 8 9 10 11]]

[[12 13 14 15]

[16 17 18 19]

[20 21 22 23]]]

45

  • An array of rank 1 with one axis
  • An array of rank 2 with two axes: Axis 0 has a length of 4 Axis 1 has a length of 3

  • An array of rank 3 with three axes: Axis 0 has a length of 2 Axis 1 has a length of 3 Axis 2 has a length of 4

46 of 179

Numpy Array 1/4

>>> b.shape

(4, 3)

>>> b.reshape(-1)

array([ 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11])

>>> b.reshape(-1, 1)

array([[ 0],

[ 1],

[ 2],

[ 3],

[ 4],

[ 5],

[ 6],

[ 7],

[ 8],

[ 9],

[10],

[11]])

>>> b.reshape(2, -1)

array([[ 0, 1, 2, 3, 4, 5],

[ 6, 7, 8, 9, 10, 11]])

>>>

46

  • -1 for the unknown remaining dimension

47 of 179

Numpy Array 2/4

>>> d = np.array([20, 30, 40, 50])

>>> e = np.arange(4)

>>> e

array([0, 1, 2, 3])

>>> f = d-e

>>> f

array([20, 29, 38, 47])

>>> e**2

array([0, 1, 4, 9], dtype=int32)

>>> A = np.array([[1, 1], [0, 1]])

>>> B = np.array([[2, 0], [3, 4]])

>>> A*B

array([[2, 0],

[0, 4]])

>>> A.dot(B)

array([[5, 4],

[3, 4]])

>>> np.dot(A, B)

array([[5, 4],

[3, 4]])

47

  • Element-wise product
  • Matrix product
  • Another matrix product

48 of 179

Numpy Array 3/4

>>> g = np.ones((2, 3), dtype=int)

>>> h = np.random.random((2, 3))

>>> g *= 3

>>> g

array([[3, 3, 3],

[3, 3, 3]])

>>> h += g

>>> h

array([[3.98175695, 3.86494818, 3.52676445],

[3.31238832, 3.97164736, 3.22392774]])

>>> k = np.random.random((2, 3))

>>> k

array([[0.56497847, 0.13008329, 0.41460729],

[0.56262637, 0.16092589, 0.63772776]])

>>> k.sum()

2.4709490742230726

>>> k.min()

0.1300832909029529

>>> k.max()

0.6377277578978647

>>>

48

49 of 179

Numpy Array 4/4

>>> m = np.arange(12).reshape(3, 4)

>>> m

array([[ 0, 1, 2, 3],

[ 4, 5, 6, 7],

[ 8, 9, 10, 11]])

>>> m.sum(axis=0)

array([12, 15, 18, 21])

>>> m.min(axis=1)

array([ 0, 4, 8])

>>> m.cumsum(axis=1)

array([[ 0, 1, 3, 6],

[ 4, 9, 15, 22],

[ 8, 17, 27, 38]], dtype=int32)

>>> n = np.arange(5)

>>> n

array([0, 1, 2, 3, 4])

>>> n[[1, 3, 4]] = 0

>>> n

array([0, 0, 2, 0, 0])

>>> exit()

pi@piot4:~ $

49

  • Cumulative sum

50 of 179

Matplotlib Examples

50

51 of 179

John D. Hunter 1968—2012

  • American neurobiologist and the original author of Matplotlib
  • The idea was originally conceived in 2002 to visualize electrocorticography (ECoG) data of epilepsy patients at the University of Chicago
  • The Python Software Foundation (PSF) awarded him posthumously the first installment of its highest distinction, the 2012 Distinguished Service Award

51

52 of 179

Bokeh and Seaborn

  • Other Python visualization libraries include Bokeh and Seaborn
  • Photographers use the Japanese word bokeh to describe the blurring of the out-of-focus parts of an image
  • "The bokeh library was so named because it allows users the flexibility to focus on the most important data without losing track of the rich context that allows it to be understood"

52

53 of 179

Learning How to Learn

  • Prof. Barbara Oakley, Oakland University, Rochester, Michigan
  • Use both focused and diffused modes of thinking while learning
  • Use the Pomodoro Technique—25 minutes of focused concentration followed by mental relaxation
  • Practice the Feynman Technique: concept-explain-gap-analogy-simplify
  • Tackle procrastination by focusing on the process instead of the product
  • Avoid the Einstellung effect of predisposition to solve a given problem in a specific manner even though better or more appropriate methods exist

53

54 of 179

X Window System

  • The X Window System originated at the Project Athena at MIT in 1984
  • The X protocol has been at version 11 (hence "X11") since September 1987
  • X11 is a software package and network protocol for using a local networked computer to interact with the graphical user interface (GUI) of an application running on a remote networked computer
  • For X forwarding in "ssh -Y" to work, a laptop computer must be running an X server program that manages the interaction between the remote application (the X client) and the laptop graphics hardware and input devices
    • For Windows, download and install Xming, open Git Bash

$ export DISPLAY=localhost:0

$ ssh -Y pi@xxx.xxx.xxx.xxx

    • For macOS, download and install XQuartz, open a terminal

$ ssh -Y pi@xxx.xxx.xxx.xxx

    • For Linux (most distributions have the X server installed), open a terminal

$ ssh -Y pi@xxx.xxx.xxx.xxx

54

55 of 179

Matplotlib Pyplot

Either enable X11 forwarding with ssh -Y or use VNC Viewer

$ ssh -Y pi@155.246.200.x

pi@piot4:~ $ cd iot

pi@piot4:~/iot $ git pull

pi@piot4:~/iot $ cd lesson8

pi@piot4:~/iot/lesson8 $ ls

boxplot_demo.py plot_lda.py README.md

histogram_demo_extended.py plot_lda_qda.py result2.csv

histogram_demo_features.py plt_cv2.py result.csv

interpolation.py plt_final.py rpidata1.csv

keras_diabetes.py pyplot_annotate.py scatter_demo.py

legend_demo.py pyplot_formatstr.py simple_plot.py

linreg.py pyplot_scales.py test.csv

major_minor_demo1.py pyplot_simple.py ticklabels_demo_rotation.py

pima-indians-diabetes.csv pyplot_text.py titanic_1.py

plot_cv_diabetes.py pyplot_three.py titanic_2.py

plot_cv_predict.py pyplot_two_subplots.py train.csv

pi@piot4:~/iot/lesson8 $ cat pyplot_simple.py

import matplotlib.pyplot as plt

plt.plot([1,2,3,4])

plt.ylabel('some numbers')

plt.show()

pi@piot4:~/iot/lesson8 $ python3 pyplot_simple.py

pi@piot4:~/iot/lesson8 $

55

  • Update repository

56 of 179

IPython and Jupyter

  • IPython (Interactive Python) provides the following features
    • Interactive shells (terminal and Qt-based)
    • A browser-based notebook interface with support for code, text, mathematical expressions, inline plots, and other media
    • Support for interactive data visualization and use of GUI toolkits
    • Flexible, embeddable interpreters to load into one's own projects
    • Tools for parallel computing
  • Project Jupyter is a spin-off project from IPython in 2014 by Fernando Pérez
  • IPython continues to exist as a Python shell and a kernel for Jupyter, while the notebook and other language-agnostic parts of IPython moved under Jupyter
  • A particularly interesting backend, provided by IPython, is the inline backend available only for the Jupyter Notebook and the Jupyter QtConsole that can be invoked as follows

%matplotlib inline

56

57 of 179

pyplot_simple.py

57

58 of 179

simple_plot.py

58

59 of 179

pyplot_formatstr.py

59

60 of 179

ticklabels_demo_rotation.py

60

61 of 179

pyplot_three.py

61

62 of 179

pyplot_two_subplots.py

62

63 of 179

pyplot_scales.py

63

symlog: symmetric log, logit or logistic unit: the logarithm of the odds p/(1-p)

64 of 179

pyplot_annotate.py

64

65 of 179

major_minor_demo1.py

65

66 of 179

legend_demo.py

66

67 of 179

scatter_demo.py

67

68 of 179

histogram_demo_features.py

pi@piot4:~/iot/lesson8 $ cat histogram_demo_features.py

import numpy as np

import scipy.stats

import matplotlib.pyplot as plt

# Example data

mu = 100 # mean of distribution

sigma = 15 # standard deviation of distribution

x = mu + sigma * np.random.randn(10000)

num_bins = 50

# The histogram of the data

n, bins, patches = plt.hist(x, num_bins, density=1, facecolor='green', alpha=0.5)

# Add a 'best fit' line

y = scipy.stats.norm.pdf(bins, mu, sigma)

plt.plot(bins, y, 'r--')

plt.xlabel('Smarts')

plt.ylabel('Probability')

plt.title(r'Histogram of IQ: $\mu=100$, $\sigma=15$')

# Tweak spacing to prevent clipping of ylabel

plt.subplots_adjust(left=0.15)

plt.show()

pi@piot4:~/iot/lesson8 $

68

  • density normalizes bin heights so that the integral of the histogram is 1
  • alpha for opacity

69 of 179

Save Figure

pi@piot4:~/iot/lesson8 $ python3 histogram_demo_features.py

69

Step 2: close this window to exit Python

Step 1: Save as figure_1.png

70 of 179

pyplot_text.py

70

71 of 179

Fig. 1. Filled Steps and Line

pi@piot4:~/iot/lesson8 $ python3 histogram_demo_extended.py

71

  • Shows seven figures

72 of 179

Fig. 2. Unequally Spaced Bars

72

73 of 179

Fig. 3. Cumulative Histogram

73

74 of 179

Fig. 4. Normalized Bars

74

75 of 179

Fig. 5. Normalized Stacked Bars

75

76 of 179

Fig. 6. Stacked Filled Steps

76

77 of 179

Fig. 7. Multiple Histograms

77

78 of 179

John Tukey 1915—2000

78

79 of 179

Box Plot

  • A box plot drawn either vertically or horizontally depicts groups of numerical data through their upper and lower quartiles
  • Whiskers extending from the box indicate variability outside the interquartile range (IQR), the difference between the upper and lower quartiles
  • The maximum is the highest datum within 1.5 IQR of the upper quartile, and the minimum is the lowest datum within 1.5 IQR of the lower quartile
  • Outliers beyond the range within the maximum and minimum are plotted as individual points
  • The spacings between the different parts of the box plot indicate the degree of dispersion (spread) and skewness in the data, allowing one to visually estimate various linear estimators (L-estimators): IQR, midhinge (the average of upper and lower quartiles), range, mid-range, and trimean (the average of median and midhinge)

79

  • Outliers

  • Maximum
  • Minimum

  • Outliers
  • Upper Quartile
  • Median
  • Lower Quartile

Contribution by Nagrajan Chandrasekaran, 2017 Spring

80 of 179

Fig. 1. Basic Box Plot

pi@piot4:~/iot/lesson8 $ python3 boxplot_demo.py

80

  • Shows seven figures

81 of 179

Fig. 2. Notched Box Plot

81

82 of 179

Fig. 3. Outlier Point Symbols

82

83 of 179

Fig. 4. Without Outlier Points

83

84 of 179

Fig. 5. Horizontal Box Plot

84

85 of 179

Fig. 6. Change Whisker Length

85

86 of 179

Fig. 7. Multiple Box Plots

86

87 of 179

Linear Regression 0/1

  • Linear regression attempts to model the relationship between two variables by fitting a linear equation to observed data
  • One variable is considered to be an explanatory variable, and the other is considered to be a dependent variable
  • This does not necessarily imply that one variable causes the other but that there is some significant association between the two variables
  • A linear regression line has an equation of the form Y = a + bX, where X is the explanatory variable, Y is the dependent variable, the slope of the line is b, and a is the intercept (the value of y when x = 0)

87

88 of 179

Linear Regression 1/1

  • The most common method for fitting a regression line is the method of least-squares
  • This method calculates the best-fitting line for the observed data by minimizing the sum of the squares of the vertical deviations from each data point to the line
  • An outlier lies far from the line and thus has a large residual value
  • A lurking variable exists when the relationship between two variables is significantly affected by the presence of a third variable which has not been included in the modeling effort
  • Attempting to extrapolate a regression equation to predict values outside of the range is often inappropriate

88

89 of 179

How Regression Got Its Name

  • Regression toward the mean in biology is a phenomenon that the heights of descendants tend to regress towards a normal average
  • Francis Galton 1822—1911 compared the heights of 205 sets of parents with adult children in "Regression towards Mediocrity in Hereditary Stature," the Journal of the Anthropological Institute of Great Britain and Ireland, Vol. 15, pp. 246-263, 1886
  • Galton concluded that as heights of the parents deviated from the average height, their children tended to be less extreme in height
  • That is, the heights of the children regressed to the average height of an adult — not about the fit lines

89

90 of 179

linreg.py 0/1

pi@piot4:~/iot/lesson8 $ cat linreg.py

import numpy as np

from scipy import stats

import matplotlib.pyplot as plt

x = np.random.random(10)

y = np.random.random(10)

slope, intercept, r_value, p_value, std_err = stats.linregress(x,y)

plt.xlabel('x')

plt.ylabel('y')

plt.plot(x,y,'ro')

plt.plot([intercept,intercept+slope])

plt.show()

pi@piot4:~/iot/lesson8 $ python3 linreg.py

90

  • Red circles, default is a solid blue line 'b-'
  • Scipy builds on Numpy, and Scipy modules need to be imported separately

91 of 179

linreg.py 1/1

91

92 of 179

Interpolation.py 0/1

pi@piot4:~/iot/lesson8 $ cat interpolation.py

import numpy as np

from scipy.interpolate import interp1d

import matplotlib.pyplot as plt

x = np.linspace(0, 10, num=11, endpoint=True)

y = np.cos(-x**2/9.0)

f = interp1d(x, y)

f2 = interp1d(x, y, kind='cubic')

xnew = np.linspace(0, 10, num=41, endpoint=True)

plt.plot(x, y, 'o', xnew, f(xnew), '-', xnew, f2(xnew), '--')

plt.legend(['data', 'linear', 'cubic'], loc='best')

plt.xlabel('x')

plt.ylabel('y')

plt.show()

pi@piot4:~/iot/lesson8 $ python3 interpolation.py

92

93 of 179

Interpolation.py 1/1

93

94 of 179

Overfitting and Underfitting

  • Overfitting is the production of an analysis that corresponds too closely or exactly to a particular set of data, and may therefore fail to fit additional data or predict future observations reliably, e.g., an overfitted model that contains more parameters than can be justified by the data
  • Underfitting occurs when a statistical model cannot adequately capture the underlying structure of the data, e.g., fitting a linear model to non-linear data
  • In machine learning, the phenomena are sometimes called overtraining and undertraining
  • To lessen the chance of, or amount of, overfitting, several techniques are available, e.g., model comparison, cross-validation, regularization, early stopping, pruning, Bayesian priors, or dropout

94

95 of 179

Choose the Right Estimator

95

96 of 179

Classification

  • Normal and shrinkage linear discriminant analysis (LDA) for classification

pi@piot4:~/iot/lesson8 $ python3 plot_lda.py

pi@piot4:~/iot/lesson8 $ python3 plot_lda_qda.py

96

97 of 179

plot_lda.py

97

98 of 179

plot_lda_qda.py

98

99 of 179

Play Audio/Video on Raspberry Pi

99

100 of 179

Piano Concerto No. 21 — Andante

100

101 of 179

Browsing History on Raspberry Pi

101

102 of 179

Google Sheet Chart — Time Series

102

  • Video Ad
  • Mouse Movement

103 of 179

Google Sheet Chart — Histogram

103

104 of 179

Google Sheet Chart — Histogram

104

105 of 179

Google Sheet Chart — Scatter Plot

105

Time is a lurking variable; Customize > Series > Linear Trendline*

* Contribution by Gina Schnecker, 2017 Fall

106 of 179

Time Series

106

107 of 179

Histogram of CPU Usage

107

108 of 179

Histogram of Temperature

108

109 of 179

Horizontal Box Plot of CPU Usage

109

110 of 179

Vertical Box Plot of Temperature

110

111 of 179

Linear Regression

111

112 of 179

Cross-Validation (CV)

  • A polynomial regression can keep adding higher order terms and get better and better fits to the data.
  • But the predictions from the model on new data will usually get worse as higher order terms are added.
  • Suppose there are n independent observations, y1,...,yn.
  • Let observation yi form the test set, and fit the model using the remaining data as the training set. Then compute the error, ei, for the omitted observation. This is a predicted residual to distinguish it from an ordinary residual.
  • Repeat step 1 for i=1,...,n, then compute cross-validation as the mean squared error from e1,...,en.

112

113 of 179

K-Fold Cross-Validation

The complete data set is partitioned into K folds, one is left out as the test set in each of the K iterations of training, and the trained model is validated by the test set.

113

Test

Training

Training

Training

Training

Training

Test

Training

Training

Training

Training

Training

Test

Training

Training

Training

Training

Training

Test

Training

Training

Training

Training

Training

Test

Test

Test

Test

Test

Test

Complete Data

Iteration 1

Iteration 2

Iteration 3

Iteration 4

Iteration 5

Predicted

Measured

114 of 179

Cross-Validation Prediction

114

Data include CPU usage only

115 of 179

Cross-Validation Prediction

115

Data include both time and CPU usage

116 of 179

plot_cv_predict.py 0/1

pi@piot4:~/iot/lesson8 $ cat plot_cv_predict.py

from sklearn import datasets

from sklearn.model_selection import cross_val_predict

from sklearn import linear_model

import matplotlib.pyplot as plt

lr = linear_model.LinearRegression()

boston = datasets.load_boston()

y = boston.target

print('Number of instances: %d' % (boston.data.shape[0]))

predicted = cross_val_predict(lr, boston.data, y, cv=10)

fig, ax = plt.subplots()

ax.scatter(y, predicted)

ax.plot([y.min(), y.max()], [y.min(), y.max()], lw=2)

ax.set_xlabel('Measured')

ax.set_ylabel('Predicted')

plt.show()

pi@piot4:~/iot/lesson8 $ python3 plot_cv_predict.py

Number of instances: 506

116

  • 10 folds

print(boston.DESCR) for the full description of the dataset

UC Irvine Machine Learning Repository

  • Cross_validation module deprecated in version 0.18 in favor of model_selection
  • Scikit-learn dataset 15.3: Boston house prices

117 of 179

plot_cv_predict.py 1/1

117

118 of 179

plot_cv_diabetes.py

118

119 of 179

Lasso Regression

  • A lasso is a rope with a noose designed as a restraint to be thrown around a target (e.g., a calf's neck) and tightened when pulled
  • In statistics and machine learning, least absolute shrinkage and selection operator (lasso) is a regression analysis method that performs both variable selection and regularization to enhance the prediction accuracy and interpretability of the statistical model it produces
  • Lasso was introduced by altering the model fitting process to select only a subset of the provided covariates for use in the final model rather than using all of them

119

120 of 179

Precision and Recall

  • The precision (or positive predictive value) is the number of correctly identified positive results divided by the number of all positive results, including those not identified correctly
  • The recall (or sensitivity) is the number of correctly identified positive results divided by the number of all samples that should have been identified as positive
  • The F-score is a measure of a test's accuracy
  • The F1 score is the harmonic mean of precision and recall

120

121 of 179

Support Vector Machines (SVMs)

In machine learning, support vector machines (SVMs) are supervised learning models with associated learning algorithms that analyze data for classification, regression, and outliers detection

pi@piot4:~ $ cd ~/iot/lesson8

pi@piot4:~/iot/lesson8 $ python3 traffic.py

[[22 4]

[ 1 40]]

precision recall f1-score support

0 0.96 0.85 0.90 26

1 0.91 0.98 0.94 41

accuracy 0.93 67

macro avg 0.93 0.91 0.92 67

weighted avg 0.93 0.93 0.92 67

121

122 of 179

TensorFlow

  • TensorFlow is a symbolic math library, and is also used for machine learning applications such as neural networks
  • TensorFlow was developed by the Google Brain team in 2011, and was released under the Apache 2.0 open source license on 2015-11-09
  • Get started with TensorFlow tutorials with examples in the Google interactive notebook
  • A tensor may be represented as a multidimensional array as in NumPy
    • A scalar is a rank-0 tensor with no axes (or dimensions)
    • A vector is a rank-1 tensor with one axis
    • A matrix is a rank-2 tensor with two axes
    • Tensors may have more than two axes
    • Ragged tensors and sparse tensors can handle different shapes

122

123 of 179

Keras

  • Keras is an open source neural network library written in Python, and is capable of running on top of TensorFlow, Microsoft Cognitive Toolkit (CNTK), or Theano until version 2.3
  • As of version 2.4, only TensorFlow is supported
  • Keras was developed as part of the research effort of project ONEIROS (Open-ended Neuro-Electronic Intelligent Robot Operating System)
  • A tutorial to develop a neural network with Keras
  • A comparison of deep learning software

123

124 of 179

The Gates of Horn and Ivory

  • In the Odyssey by Homer, book 19, lines 560-569, the literary image of the gates of horn and ivory was used to distinguish true dreams (corresponding to factual occurrences) from false
  • Keras, the Greek word for "horn," is similar to that for "fulfill," and the Greek word for "ivory" is similar to that for "deceive"
  • A true dream is spoken of as coming through the gate of horn, a false dream as coming through the gate of ivory
  • The House of Sleep (1731) by Bernard Picart 1673—1733 depicts the facade of the dwelling place of the Oneiros (Dream)
  • The Oneiroi (Dreams) are the sons of Nyx (Night), and brothers of Hypnos (Sleep)

124

125 of 179

Keras/TensorFlow on Raspberry Pi

pi@piot4:~ $ python3

Python 3.7.3 (default, Apr 3 2019, 05:39:12)

[GCC 8.2.0] on linux

Type "help", "copyright", "credits" or "license" for more information.

>>> import tensorflow as tf

/usr/local/lib/python3.7/dist-packages/tensorflow_core/python/framework/dtypes.py:516: FutureWarning: Passing (type, 1) or '1type' as a synonym of type is deprecated; in a future version of numpy, it will be understood as (type, (1,)) / '(1,)type'.

...

>>> hello = tf.constant('hello, world')

>>> sess = tf.Session()

>>> print(sess.run(hello).decode())

hello, world

>>> exit()

pi@piot4:~ $ cd iot/lesson8

pi@piot4:~/iot/lesson8 $ python3 keras_diabetes.py

More Keras examples at GitHub

125

126 of 179

Google Colab and AI Hub

  • Google Colaboratory is a Google research project created to help disseminate machine learning education and research
  • Colab is a Jupyter notebook environment that requires no installation and setup to use, and runs the TensorFlow tutorials directly in the browser and entirely in the cloud
  • Google Cloud AI Hub is a hosted repository of plug-and-play AI components, including end-to-end AI pipelines and out-of-the-box algorithms

126

Contribution by Alhussain Almarhabi, 2017 Fall

127 of 179

RMS (Royal Mail Ship) Titanic

  • Titanic carried 20 lifeboats only enough for 1,178 people
  • Approximately 2,224 people aboard—1,316 passengers and 908 crew members
  • 710 people rescued by RMS Carpathia—498 passengers and 212 crew members

127

128 of 179

Maiden Voyage of Titanic

128

41°43′ N, 49°56′ W

129 of 179

Thermal Inversion

The Titanic was sailing from Gulf Stream waters into the frigid Labrador Current, where the air column was cooling from the bottom up, creating a thermal inversion: layers of cold air below layers of warmer air. Extraordinarily high air pressure kept the air free of fog.

129

130 of 179

Superior Mirage

A thermal inversion refracts light abnormally and can create a superior mirage: Objects appear higher (and therefore nearer) than they actually are, before a false horizon. The area between the false horizon and the true one may appear as haze.

130

131 of 179

Iceberg Camouflage

The Californian’s radio operator warned the Titanic of ice. But the moonless night provided little contrast, and a calm sea masked the line between the true and false horizons, camouflaging the iceberg. A Titanic lookout sounded the alarm when the berg was about a mile away—too late.

131

132 of 179

Mistaken Identity

Shortly before the collision, the Titanic sailed into the Californian’s view—but it appeared too near and small to be the great ocean liner. Californian captain Stanley Lord knew the Titanic was the only other ship in the area with a radio, and so concluded this ship did not have one.

132

133 of 179

Morse Lamp

Californian captain said he repeatedly had someone signal the ship by Morse lamp “and she did not take the slightest notice of it.” The Titanic, now in trouble, signaled the Californian by Morse lamp, also to no avail. The abnormally stratified air was distorting and disrupting the signals.

133

134 of 179

Distress Rockets Ignored

The Titanic fired distress rockets some 600 feet into the air—but they appeared to be much lower relative to the ship. Those aboard the Californian, unsure of what they saw, ignored the signals. When the Titanic sank at 2:20 am on Monday, 1912-04-15, they thought she might be simply sailing away.

134

135 of 179

titanic_1.py 0/2

pi@piot4:~/iot/lesson8 $ cat titanic_1.py

import numpy as np

import matplotlib.pyplot as plt

import pdb

from pandas import *

data = read_csv('train.csv')

cols = data.columns

print('Data Columns:')

print(cols)

survivors = data.groupby('Sex').Survived.mean()

print('Survived:')

print(survivors)

plt.hist(data.Pclass, bins=[1,2,3,4], align='left', rwidth=0.5)

plt.xticks([0,1,2,3,4])

plt.xlabel('Passenger Class')

plt.ylabel('Number of Passengers')

plt.show()

135

  • The module pdb defines an interactive source code debugger for Python programs

136 of 179

titanic_1.py 1/2

pi@piot4:~/iot/lesson8 $ cp train.csv ~/demo

pi@piot4:~/iot/lesson8 $ cp test.csv ~/demo

pi@piot4:~/iot/lesson8 $ cp titanic_1.py ~/demo

pi@piot4:~/iot/lesson8 $ cp titanic_2.py ~/demo

pi@piot4:~/demo $ python3 titanic_1.py

Data Columns:

Index(['PassengerId', 'Survived', 'Pclass', 'Name', 'Sex', 'Age', 'SibSp',

'Parch', 'Ticket', 'Fare', 'Cabin', 'Embarked'],

dtype='object')

Survived:

Sex

female 0.742038

male 0.188908

Name: Survived, dtype: float64

136

  • Captain's order to evacuate women and children first

Contribution by Abhinandan Nuli, 2020 Spring

  • SibSp: Number of siblings/spouse aboard
  • Parch: Number of parents/children aboard

137 of 179

titanic_1.py 2/2

137

138 of 179

titanic_2.py 0/3

pi@piot4:~/demo $ cat titanic_2.py

import numpy as np

from pandas import *

options.mode.chained_assignment = None

data=read_csv('train.csv')

columns=data.columns

look_up=(data.groupby(['Sex', 'Pclass']).Survived.mean())

test=read_csv('test.csv')

test['Prediction']=0

for i in range(len(test)):

test.Prediction[i]=round(look_up[test.Sex[i]][test.Pclass[i]])

test.to_csv('result.csv', index=False)

data['Fare_Bracket']=0

test['Fare_Bracket']=0

na=sum(test.Fare != test.Fare)

wh_badfare=np.flatnonzero(test.Fare != test.Fare)

sv = test.Fare[test.Fare != test.Fare]

test.Fare[test.Fare != test.Fare]=(3-test.Pclass)*11

138

  • Add a column to the right
  • Number and index of missing value where np.nan != np.nan is True
  • NaN (Not a Number): default missing value marker

139 of 179

titanic_2.py 1/3

test.Fare_Bracket=np.array([min(int(price/10),3) for price in test.Fare])

test.Fare[wh_badfare]=sv

data.Fare_Bracket=np.array([min(int(price/10),3) for price in data.Fare])

look_up2=(data.groupby(['Sex', 'Pclass', 'Fare_Bracket']).Survived.mean())

for i in range(len(test)):

test.Prediction[i] =

round(look_up2[test.Sex[i]][test.Pclass[i]][test.Fare_Bracket[i]])

test.to_csv('result2.csv', index=False)

print('Survived:')

print(look_up)

print('Number of Missing Fare: %d' % (na))

print('Indices of Missing Fare: %d' % (wh_badfare))

print('Survived:')

print(look_up2)

pi@piot4:~/demo $ python3 titanic_2.py

139

Four fare brackets: ≤£9, £10~19, £20~29, and ≥£30

140 of 179

titanic_2.py 2/3

Survived:

Sex Pclass

female 1 0.968085

2 0.921053

3 0.500000

male 1 0.368852

2 0.157407

3 0.135447

Name: Survived, dtype: float64

Number of Missing Fare: 1

Indices of Missing Fare: 152

140

  • Third-class passengers primarily emigrants traversed unfamiliar areas to reach lifeboats

141 of 179

titanic_2.py 3/3

Survived:

Sex Pclass Fare_Bracket

female 1 2 0.833333

3 0.977273

2 1 0.914286

2 0.900000

3 1.000000

3 0 0.593750

1 0.581395

2 0.333333

3 0.125000

male 1 0 0.000000

2 0.400000

3 0.383721

2 0 0.000000

1 0.158730

2 0.160000

3 0.214286

3 0 0.111538

1 0.236842

2 0.125000

3 0.240000

Name: Survived, dtype: float64

pi@piot4:~/demo$

141

  • Female members from large families—the Anderssons (7), Goodwins (8), Lefebvres (5), Pålssons (5), Panulas (6), Rices (6), Sages (11), and Skoogs (6)

}

142 of 179

Titanic Passengers

142

143 of 179

Decision Trees or Neural Networks

  • Scikit-Learn can create decision trees [site, repo] for the Titanic data
    • The smaller tree is better at classifying the test data
    • The larger one is better with the training data
    • Overall, the smaller one is the better choice
  • Scikit-Learn can also create neural networks [site, repo] for the Titanic data
  • Graphviz provides graph visualization

143

Contribution by Noah McDermott, 2020 Spring

144 of 179

Titanic Passengers and Crew

https://upload.wikimedia.org/wikipedia/commons/6/69/Titanic_casualties.svg

  • Apathy — feeling no concern for the other person

  • Sympathy — feeling sorrow or concern for the other person

  • Empathy — feeling the same emotions as the other person

144

145 of 179

The New Colossus

  • "The New Colossus" is a sonnet by American poet Emma Lazarus 1849—1887 who wrote the poem in 1883 to raise money for the construction of a pedestal for the Statue of Liberty dedicated on 1886-10-28
  • The raised right foot of the statue depicts that she is walking forward
  • In 1903, the poem was cast onto a bronze plaque and mounted inside the lower level of the pedestal

145

146 of 179

Charles Minard, Napoleon's March

146

Napoleon’s troops size to Moscow in brown and the way back in black with temperature

Historical Data Visualization: Minard’s Map Vectorized and Revisited

147 of 179

War and Peace

  • Lev Nikolayevich Tolstoy 1828—1910 was a Russian writer best known for the novels War and Peace (1869) and Anna Karenina (1877)
  • War and Peace chronicles the French invasion of Russia and the impact of the Napoleonic era on Tsarist society through the stories of five Russian aristocratic families
  • "Tolstoy's use of visual detail is often comparable to cinema, using literary techniques that resemble panning, wide shots, and close-ups"
  • "These devices, while not exclusive to Tolstoy, are part of the new style of the novel that arose in the mid-19th century and of which Tolstoy proved himself a master," Caryl Emerson, "The Tolstoy Connection in Bakhtin"

147

148 of 179

Life Expectancy vs. GDP per Capita

148

149 of 179

Hans Rosling 1948—2017

  • Hans Rosling was a Swedish medical doctor, academic, statistician, and public speaker
  • He was the Professor of International Health at Karolinska Institutet
  • He co-founded the Gapminder Foundation together with his son Ola and daughter-in-law Anna Rosling Rönnlund to develop Trendalyzer software to animate data compiled by the United Nations and the World Bank that helped explain the world with graphics
  • Hans Rosling, "The best stat you've ever seen," TED 2006

149

150 of 179

GSMA Mobile Connectivity Index

150

151 of 179

Global Human Journey

151

152 of 179

Most Recent Common Ancestor

152

Last Universal Common Ancestor (LUCA) and Most Recent Common Ancestor (MRCA)

153 of 179

Country Population

153

154 of 179

154

155 of 179

US Wind Map

155

156 of 179

New York City 311 Calls

156

34,522 complaints called into 311 between September 8 and 15, 2010

http://www.wired.com/2010/11/ff_311_new_york

157 of 179

Nanocubes

157

158 of 179

Super Bowl and Driving Behavior

158

Hard accelerations and breaks from a subset of Automatic drivers

http://blog.automatic.com/super-bowl-50

159 of 179

ThingSpeak

159

160 of 179

Kibana

160

161 of 179

Freeboard

161

162 of 179

IBM Bluemix

162

163 of 179

DGLogik

163

164 of 179

Qlik

164

165 of 179

Data Analysis and Report

165

166 of 179

Natural Language Generation (NLG)

166

167 of 179

Wikipedia Editing Bots

167

168 of 179

Parabola

Parabola is a visual coding tool that enables anyone to work with large data sets, use APIs, and automate workflows, all without writing a single line of code

168

Contribution by Kyra DiFrancesco, 2019 Spring

169 of 179

Robotic Process Automation

  • Robotic process automation (or RPA) is a form of business process automation technology based on metaphorical software robots (bots) or AI workers
  • In traditional workflow automation tools, a software developer produces a list of actions to automate a task and interface to the back-end system using internal application programming interfaces (APIs) or dedicated scripting language
  • In contrast, RPA systems develop the action list by watching the user perform that task in the application's graphical user interface (GUI), and then perform the automation by repeating those tasks directly in the GUI
  • Utilizing RPA 2.0 or unassisted RPA, a process can be run on a computer without needing input from a user, freeing up that user to do other work
  • Hyperautomation is the application of advanced technologies like RPA, AI, machine learning (ML), and process mining to augment workers and automate processes with a combination of tools to help support replicating pieces where the human is involved in a task

169

170 of 179

Sentiment Analysis

  • Sentiment analysis or opinion mining refers to the use of natural language processing, text analysis, computational linguistics, and biometrics to systematically identify, extract, quantify, and study affective states and subjective information
  • Online opinion has turned into a kind of virtual currency for businesses looking to market products, identify new opportunities, and manage reputations

170

171 of 179

Apache MXNet

171

172 of 179

deeplearning.ai

  • Andrew Ng cofounded Google Brain in 2011, Coursera in 2012, and deeplearning.ai in 2017
  • The deeplearning.ai courses include
    • Deep learning specialization
      • Neural Networks and Deep Learning
      • Improving Deep Neural Networks
      • Structuring Machine Learning Projects
      • Convolutional Neural Networks
      • Sequence Models
    • TensorFlow specialization
      • Introduction to TensorFlow for AI, ML and DL
      • Convolutional Neural Networks in TensorFlow
      • Natural Language Processing in TensorFlow
      • Sequences, Time Series, and Prediction
    • AI for everyone
      • What is AI
      • Building AI Projects
      • AI in Your Company
      • AI and Society
  • Workera provides skill-based self-assessments, career advice, and access to job offers with AI companies
  • The Batch presents the most important AI events and perspective for engineers and business leaders

172

173 of 179

Object Detection

  • Object detection applies computer vision and digital image processing to detect instances of semantic objects of a certain class such as humans, buildings, or cars in digital images and videos
  • Deep learning approaches to object detection include
    • Region-based convolutional neural networks (R-CNN), Fast R-CNN, Faster R-CNN, Mask R-CNN, and Mesh R-CNN
    • Single Shot MultiBox Detector (SSD)
    • You Only Look Once (YOLO)
    • Single-shot refinement neural network for object detection (RefineDet)
    • Retina-Net
    • Deformable convolutional networks

173

174 of 179

Intel Neural Compute Stick (NCS)

  • NCS 2, a fanless deep learning development kit, features the Intel Movidius Myriad X vision processing unit (VPU) for deep neural network (DNN) applications
  • Intel OpenVINO (Visual Inference and Neural network Optimization) toolkit enables deep learning inference and heterogeneous execution across cloud architectures to edge devices

174

Contribution by Ziran Gong, 2018 Fall

175 of 179

Acceleration-as-a-Service (AaaS)

Soft DNN processing units (DPUs) based on field-programmable gate arrays (FPGAs) enable Acceleration-as-a-Service such as Accelize, Amazon EC2 F1, and Microsoft Brainwave with access to ready-to-use accelerators in the Cloud

175

176 of 179

CUDA Processing Flow

  • Compute Unified Device Architecture (CUDA) is a parallel computing platform and application programming interface (API) model created by Nvidia
  • The CUDA platform is a software layer that gives direct access to the virtual instruction set and parallel computational elements of the graphics processing unit (GPU), for the execution of compute kernels

176

177 of 179

Tensor Processing Unit (TPU)

  • A tensor processing unit (TPU) mounted in a heat sink assembly is an application-specific integrated circuit (ASIC) developed by Google for machine learning:
    • TensorFlow
    • AlphaGo
    • Street View text processing
    • Google Photos
    • RankBrain
  • Compared to a graphics processing unit (GPU), TPU is designed explicitly for a higher volume of reduced precision computation with higher input/output operations per second (IOPS) per watt

177

178 of 179

Cerebras Wafer Scale Engine (WSE)

  • Cerebras is dedicated to accelerating deep learning
  • The Cerebras WSE is 46,225 mm2 with 1.2 Trillion transistors, 400,000 Sparse Linear Algebra (SLA) cores, 18 GB on-chip SRAM, 100 Pb/s interconnect
  • The Nvidia GV100 GPU is 815 mm2 with 21.1 Billion transistors
  • See Graphcore Intelligence Processing Unit (IPU), SambaNova

178

179 of 179

Lesson 8 Summary

  • Data scientists
    • Use data to solve problems
    • Understand data and extract value from data
    • Work on data solutions that would have an immediate and massive impact on the business
  • Pay attention to
    • Missing data, outliers, and lurking variables
  • Most importantly
    • Correlation is not causation
    • Prudent extrapolation out of data range

179