1 of 21

Introduction to Big Data

Hadoop and MapReduce

Thejdeep G

12CO99

2 of 21

What is Big Data ?

Volume | Velocity | Variety

Moore’s Law

3 of 21

Why is it important ?

Efficiency | Save Lives | Technology Advancement

4 of 21

Preparing for the ‘Data Flood’

5 of 21

Problems Much ?

Slow Read/Writes | HW failure | Merge multiple reads

Distributed Processing

6 of 21

What is Hadoop ?

Open Source | Distributed | Massive Storage Faster Processing

7 of 21

How did it get there ?

This one for the history buffs

8 of 21

Why Hadoop ?

Cost | Power | Scalability | Protection

9 of 21

A Typical Hadoop Cluster

10 of 21

Core Components of Cluster

11 of 21

Core Components of Cluster

12 of 21

Core Components of Cluster

13 of 21

The Hadoop Workflow

  • How a file gets loaded into a cluster ?
  • To which data nodes to load the blocks ?
  • Who does the block replication ?

14 of 21

Architecture of Hadoop

  • Hadoop Common
  • HDFS
  • MapReduce
  • YARN

15 of 21

What is HDFS ?

Overview & Design

16 of 21

HDFS Architecture

17 of 21

HDFS Failure Techniques

  • DataNode Failure & Recovery
  • NameNode Failure
  • Secondary NameNode
  • Block Placement
  • Balancing Cluster

18 of 21

What is MapReduce ?

The Basics

19 of 21

Two Phases

Mapping Reducing

20 of 21

MapReduce Workflow

21 of 21

Thanks

Questions ?