DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
1
III B.Tech II Sem
Subject: NoSQL Databases Code: 23AD08
UNIT-4
Topic: Column Family Data store
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
2
UNIT-IV: Column-oriented NoSQL databases using Apache HBASE, Column-oriented NoSQL databases using Apache Cassandra, Architecture of HBASE, Column-Family Data Store Features, Consistency, Transactions, Availability, Query Features, Scaling, Suitable Use Cases, Event Logging, Content Management Systems, Blogging Platforms, Counters, Expiring Usage.
Course Outcomes: At the end of the Course the student will be able to
CO1: Explain and compare different types of NoSQL Databases
CO2: Compare and contrast RDBMS with different NoSQL databases.
CO3: Demonstrate the detailed architecture and performance tune of Document-oriented NoSQL databases.
CO4 : Explain performance tune of Key-Value Pair NoSQL databases.
CO5: Apply NoSQL development tools on different types of NoSQL Databases.
NOSQL DATABASES
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
3
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
4
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
5
Row vs Column Oriented Databases
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
6
APACHE HBASE
Limitations of Hadoop
Hadoop can perform only batch processing, and data will be accessed only in a sequential manner. That means one has to search the entire dataset even for the simplest of jobs.
A huge dataset when processed results in another huge data set, which should also be processed sequentially. At this point, a new solution is needed to access any point of data in a single unit of time (random access).
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
7
APACHE HBASE
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
8
HBase evolution flowchart:
Year | Event |
Nov 2006 | Google released the paper on BigTable. |
Feb 2007 | Initial HBase prototype was created as a Hadoop contribution. |
Oct 2007 | The first usable HBase along with Hadoop 0.15.0 was released. |
Jan 2008 | HBase became the sub project of Hadoop. |
Oct 2008 | HBase 0.18.1 was released. |
Jan 2009 | HBase 0.19.0 was released. |
Sept 2009 | HBase 0.20.0 was released. |
May 2010 | HBase became Apache top-level project. |
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
9
Key Features of Apache Hbase
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
10
Architecture of Apache HBase
Apache HBase follows a master-slave architecture and is built on top of Hadoop HDFS. Here's how its major components work together:
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
11
Here’s how each part of the Apache HBase architecture works in detail:
1. HMaster
Acts as the master node of the HBase cluster.
Its main responsibilities include:
If the HMaster fails, a backup HMaster can take over to ensure high availability.
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
12
Here’s how each part of the Apache HBase architecture works in detail:
2. RegionServer
A worker node in HBase, serving client requests.
Each RegionServer manages multiple regions, meaning chunks of tables. Internally, RegionServers have:
RegionServers also manage WAL (Write Ahead Log) for crash recovery.If a RegionServer fails, the HMaster reassigns its regions to other RegionServers.
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
13
Here’s how each part of the Apache HBase architecture works in detail:
3. Region
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
14
Here’s how each part of the Apache HBase architecture works in detail:
4. ZooKeeper
It's an external, reliable coordination service.
HBase uses ZooKeeper to:
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
15
Here’s how each part of the Apache HBase architecture works in detail:
5. HDFS (Hadoop Distributed File System)
HBase uses HDFS to store actual data on disk. It stores HFiles which are compressed files with the actual data and also WAL (Write Ahead Logs) files for durability. It Provides fault-tolerant, distributed storage so data is protected even if hardware fails.
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
16
Difference Between Hadoop and HBase
Hadoop: Hadoop is an open source framework from Apache that is used to store and process large datasets distributed across a cluster of servers. Four main components of Hadoop are Hadoop Distributed File System(HDFS), Yarn, MapReduce, and libraries. It involves not only large data but a mixture of structured, semi-structured, and unstructured information. Amazon, IBM, Microsoft, Cloudera, ScienceSoft, Pivotal, Hortonworks are some of the companies using Hadoop technology.
HBase: HBase is an open source database from Apache that runs on Hadoop cluster. It falls under the non-relational database management system. Three important components of HBase are HMaster, Region server, Zookeeper. CapitalOne, JPMorganchase, apple, MTB, AT& T, Lockheed Martin are some of the companies using HBase.
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
17
S.No. | Hadoop | HBase |
1 | Hadoop is a collection of software tools | HBase is a part of hadoop eco-system |
2 | Stores data sets in a distributed environment | Stores data in a column-oriented manner |
3 | Hadoop is a framework | HBase is a NOSQL database |
4 | Data are stored in form of chunks | Data are stored in form of key/value pair |
5 | Hadoop does not allow run time changes | HBase allows run time changes |
6 | File can be written only once, can be read many times | File can be read and write multiple times |
7 | Hadoop has low latency operations | HBase has high latency operations |
8 | HDFS can be accessed through MapReduce | HBase can be accessed through shell commands, Java API, REST |
Below is a table of differences between Hadoop and HBase:
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
18
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
19
Column-Family Stores
Column-family stores, such as Cassandra [Cassandra], HBase [Hbase], Hypertable [Hypertable], and Amazon SimpleDB [Amazon SimpleDB], allow you to store data with keys mapped to values and the values grouped into multiple column families, each column family being a map of data.
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
20
What Is a Column-Family Data Store?
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
21
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
22
Features
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
23
Features
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
24
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
25
When we use super columns to create a column family, we get a super column family.
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
26
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
27
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
28
Consistency
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
29
Consistency
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
30
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
31
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
32
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
33
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
34
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
35
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
36
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
37
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
38
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
39
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
40
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
41
DEPT.OF AI&DS V.SOWJANYA, SR.ASST.PROF.,
42