1 of 13

Sri Raghavendra Educational Institutions Society (R)

(Approved by AICTE, Accredited by NAAC, Affiliated to VTU, Karnataka)

Sri Krishna Institute of Technology

www.skit.org.in

Title: HDFS (Hadoop Distributed File System)

CO addressed: CO2

Course: Big Data Analytics

Presented by: Mr. P. Kiran Kumar

Department: ISE

2 of 13

2

14/08/2026

/skit.org.in

(Approved by AICTE, Accredited by NAAC, Affiliated to VTU, Karnataka)

Sri Krishna Institute of Technology

HDFS (Hadoop Distributed File System)

  • HDFS is the storage system of Hadoop.
  • It is designed to store and manage large volumes of data.
  • Data is distributed across multiple machines in a Hadoop cluster.
  • HDFS provides reliable and fault-tolerant storage.

Large File → Split into Blocks → Distributed across DataNodes

​

HDFS = Distributed Storage + Fault Tolerance + Scalability

3 of 13

3

14/08/2026

/skit.org.in

(Approved by AICTE, Accredited by NAAC, Affiliated to VTU, Karnataka)

Sri Krishna Institute of Technology

Main Components:

​

Client Application

  • Interacts with HDFS through the Hadoop File System Client
  • Performs read/write operations.

NameNode

  • Stores metadata about files.
  • Maintains the file-to-block mapping.
  • Assigns blocks to different DataNodes.

DataNodes

  • Store the actual data blocks.
  • Blocks are replicated across multiple DataNodes for fault tolerance.

​

Example

Sample.txt → Block A + Block B + Block C

These blocks are distributed and replicated across DataNode A, DataNode B and DataNode C.

4 of 13

4

14/08/2026

/skit.org.in

(Approved by AICTE, Accredited by NAAC, Affiliated to VTU, Karnataka)

Sri Krishna Institute of Technology

Key Features of HDFS

Distributed Storage

  • Stores data across multiple machines in a Hadoop cluster.

Fault Tolerance

  • Maintains multiple copies of data blocks to protect against node failures.

High Scalability

  • Storage capacity can be increased by adding more DataNodes.

High Throughput

  • Designed for fast processing of large volumes of data.

Large File Support

  • Optimized for storing and processing very large files.

Data Replication

  • Replicates blocks across DataNodes to improve reliability.

5 of 13

5

14/08/2026

/skit.org.in

(Approved by AICTE, Accredited by NAAC, Affiliated to VTU, Karnataka)

Sri Krishna Institute of Technology

NameNode

  • HDFS breaks a large file into smaller pieces called blocks.
  • NameNode uses a Rack ID to identify DataNodes in a rack.
  • A rack is a collection of DataNodes within the cluster.
  • NameNode keeps track of the blocks of a file placed on various DataNodes.
  • Manages file-related operations such as:
    • Read
    • Write
    • Create
    • Delete
  • Its main job is managing the File System Namespace.
  • The file system namespace is a collection of files in the cluster.
  • NameNode stores the HDFS namespace, including block-to-file mapping and file properties, in FsImage.
  • EditLog records every transaction that occurs in the file system metadata.
  • The material states that there is a single NameNode per cluster.

6 of 13

6

14/08/2026

/skit.org.in

(Approved by AICTE, Accredited by NAAC, Affiliated to VTU, Karnataka)

Sri Krishna Institute of Technology

DataNode

  • There are multiple DataNodes per cluster.
  • DataNodes communicate with each other during pipeline read and write operations.
  • A DataNode continuously sends a heartbeat message to the NameNode.
  • The heartbeat ensures connectivity between the NameNode and DataNode.
  • If there is no heartbeat from a DataNode:
    • NameNode assumes that the DataNode is unavailable.
    • NameNode triggers replication of its data to another DataNode.
  • This mechanism provides fault tolerance and high availability.

​

Heartbeat & Replication

​

DataNode → Heartbeat → NameNode

​

DataNode Failure → No Heartbeat → NameNode → Replicate Data

7 of 13

7

14/08/2026

/skit.org.in

(Approved by AICTE, Accredited by NAAC, Affiliated to VTU, Karnataka)

Sri Krishna Institute of Technology

Secondary NameNode

  • Takes a snapshot of HDFS metadata at intervals specified in the Hadoop configuration.
  • Its memory requirements are similar to those of the NameNode.
  • Therefore, it is better to run the NameNode and Secondary NameNode on different machines.
  • In case of NameNode failure, the Secondary NameNode can be configured manually to bring up the cluster.
  • The Secondary NameNode does not record real-time changes that happen to HDFS metadata.

8 of 13

8

14/08/2026

/skit.org.in

(Approved by AICTE, Accredited by NAAC, Affiliated to VTU, Karnataka)

Sri Krishna Institute of Technology

Anatomy of File Read

File Read Operation

​

  1. Client calls open() on DistributedFileSystem.
  2. NameNode provides the locations of the data blocks.
  3. Client receives an FSDataInputStream.
  4. Client connects to the closest DataNode

and reads the block.

  • After one block, it connects to the best

DataNode for the next block.

  • After reading all blocks, client calls close().

​

Flow:

Client → NameNode → DataNode → Data

​

9 of 13

9

14/08/2026

/skit.org.in

(Approved by AICTE, Accredited by NAAC, Affiliated to VTU, Karnataka)

Sri Krishna Institute of Technology

HDFS File Write

File Write Operation

​

  1. Client calls create() on DistributedFileSystem.
  2. NameNode checks and creates the file.
  3. Data is divided into packets.
  4. NameNode selects DataNodes to form a pipeline.
  5. Packets flow through the DataNodes and are replicated.
  6. DataNodes send acknowledgements.
  7. Client calls close() after writing.

​

Flow

Client → NameNode → DataNode 1 → DataNode 2 → DataNode 3

​

10 of 13

10

14/08/2026

/skit.org.in

(Approved by AICTE, Accredited by NAAC, Affiliated to VTU, Karnataka)

Sri Krishna Institute of Technology

HDFS Replica Placement Strategy

Hadoop Default Replica Placement

  • 1st Replica: Same node as the client.
  • 2nd Replica: A node on a different rack.
  • 3rd Replica: A different node on the same rack as the 2nd replica.

Once replica locations are selected, a pipeline is created.

This strategy provides good reliability.

​

Replica Placement

Client → Rack 1 → Rack 2

​

Replica 1 → Replica 2 → Replica 3

​

11 of 13

11

14/08/2026

/skit.org.in

(Approved by AICTE, Accredited by NAAC, Affiliated to VTU, Karnataka)

Sri Krishna Institute of Technology

HDFS Commands

1. List files and directories

  • hadoop fs -ls /

Lists contents of the HDFS root directory.

​

2. List recursively

  • hadoop fs -ls -R /

Lists all files and directories, including subdirectories.

​

3. Create a directory

  • hadoop fs -mkdir /sample

Creates a new directory named sample in HDFS.

​

4. Copy local file to HDFS

  • hadoop fs -put /root/sample/test.txt /sample/test.txt

Uploads a file from the local file system to HDFS.

​

12 of 13

12

14/08/2026

/skit.org.in

(Approved by AICTE, Accredited by NAAC, Affiliated to VTU, Karnataka)

Sri Krishna Institute of Technology

Copying Files in HDFS

​

1. Copy Local → HDFS

  • hadoop fs -put /root/sample/test.txt /sample/test.txt

Uploads a file from the local file system to HDFS.

​

2. Copy HDFS → Local

  • hadoop fs -get /sample/test.txt /local/path/testsample.txt

Downloads a file from HDFS to the local file system.

​

3. Copy Local → HDFS using copyFromLocal

  • hadoop fs -copyFromLocal /root/sample/test.txt /sample/testsample.txt

Copies a file from the local file system to HDFS.

13 of 13

13

14/08/2026

/skit.org.in

(Approved by AICTE, Accredited by NAAC, Affiliated to VTU, Karnataka)

Sri Krishna Institute of Technology

Special Features of HDFS

1. Data Replication

  • Client application does not need to track all blocks.
  • Client is directed to the nearest replica.
  • This helps ensure high performance.

​

2. Data Pipeline

  • Client writes a block to the first DataNode.
  • The first DataNode forwards the data to the next DataNode.
  • The process continues through the pipeline.
  • All replicas are eventually written to disk.

​

Data Pipeline

Client → DataNode 1 → DataNode 2 → DataNode 3