Introduction to Apache Spark
Big Data Analytics with Python, AIMS2023
Dunstan Matekenya, PhD
Contents
Whats Apache Spark
Spark’s Development History
Spark Vs. Hadoop MapReduce
Spark Vs. Hadoop MapReduce
Spark Vs. Hadoop MapReduce Con’t
Source: IBM site
READ PLEASE
In Mining Massive Datasets, they provide a nice overview of differences between Spark and MapReduce. See page 41-48
Misconceptions About Spark and MapReduce
Source: IBM site
Why Choose Spark
Getting, Installing and Running Spark
Spark Language APIs
Spark’s language APIs make it possible for you to run Spark code using various programming languages.
Installing Spark for Single Use in Python
Installing Spark for a Cluster
Spark’s Basic Architecture
Overview of Spark Components
Apache Spark components and architecture
Spark Driver Program
Another view of Spark’s Architecture
The SparkSession
The Cluster Manager
Spark Executor
The Spark Web UI
Deploying Spark Applications
Because the cluster manager is agnostic to where it runs (as long as it can manage Spark’s executors and fulfill resource requests), Spark can be deployed in some of the most popular environments—such as Apache Hadoop YARN and Kubernetes.
Spark Deployment modes
Deployment mode distinguishes where the driver process runs. In "cluster" mode, the framework launches the driver inside of the cluster. In "client" mode, the submitter launches the driver outside of the cluster.
Other Spark Concepts Related to Spark Application
Spark Language APIs-Revisited
Spark Application Concepts
Distributed Data and Partitions
Partitions
Each executor’s core gets a partition of data to work on
With DataFrames, we do not (for the most part) manipulate partitions manually(on an individual basis), however, we can still (and need to) control partitioning. For instance, specify number of partitions.
This code snippet will break up the physical data stored across clusters into eight partitions, and each executor will get one or more partitions to read into its memory
Summary of Key Concepts-Again
Spark components communicate through the Spark driver in Spark’s distributed architecture
Spark Jobs
Spark driver creating one or more Spark jobs
Spark Stages
Spark driver creating one or more Spark jobs
Spark Stages
Spark job creating one or more stages
Spark Tasks
Spark stage creating one or more tasks to be distributed to executors
Transformations and Actions
Spark operations on distributed data can be classified into two types: transformations and actions.
Transformations and Actions
Transformation | Action |
Sample() | Count() |
OrderBy() | Show() |
filter() | take() |
select() | collect() |
Transformation vs actions in Spark
Narrow Vs. Wide Transformations
Narrow versus wide transformations
Narrow Vs. Wide Transformations
Lazy Evaluation
Lazy Evaluation Illustrated
Lazy transformations and eager actions
The Spark UI
A Tour of Spark APIs and Toolset
Spark APIs
Apache Spark Ecosystem of APIs and Libraries
Apache Spark components and API stack
SparkSQL
SparkMLlib
Spark Structured Streaming
GraphX
Further Reading
First Spark Program
Spark in Local Mode
Spark Through Pyspark Shell
Spark Folder Contents
Running Spark Shell
We used the high-level Structured APIs to read a text file into a Spark DataFrame rather than an RDD.
From now on, we will run Spark in