1 of 137

Module 1

Introduction to Big Data

Analytics of Big Data could enable discovery of new facts, knowledge and strategy in a number of fields, such as manufacturing, business, finance, healthcare, medicine and education.

Need of Big Data

  • The rise in technology has led to the production and storage of voluminous amounts of data.
  • Earlier megabytes (106 B) were used but nowadays petabytes (1015 B) are used for processing, analysis, discovering new facts and generating new knowledge.
  • Conventional systems for storage, processing and analysis pose challenges in large growth in volume of data, variety of data, various forms and formats, increasing complexity, faster generation of data and need of quickly processing, analyzing and usage.

2 of 137

What is Big Data?

According to Gartner, the definition of Big Data –“Big data” is high-volume, velocity, and variety information assets that demand cost-effective, innovative forms of information processing for enhanced insight and decision making.”

 

This definition clearly answers the “What is Big Data?” question – Big Data refers to complex and large data sets that have to be processed and analyzed to uncover valuable information that can benefit businesses and organizations.

What is Big Data:

It refers to a massive amount of data that keeps on growing exponentially with time.

It is so voluminous that it cannot be processed or analyzed using conventional data processing techniques.

It includes data mining, data storage, data analysis, data sharing, and data visualization.

The term is an all-comprehensive one including data, data frameworks, along with the tools and techniques used to process and analyze the data.

3 of 137

The History of Big Data

Although the concept of big data itself is relatively new, the origins of large data sets go back to the 1960s and '70s when the world of data was just getting started with the first data centers and the development of the relational database.

Around 2005, people began to realize just how much data users generated through Facebook, YouTube, and other online services. Hadoop (an open-source framework created specifically to store and analyze big data sets) was developed that same year. NoSQL also began to gain popularity during this time.

4 of 137

The development of open-source frameworks, such as Hadoop (and more recently, Spark) was essential for the growth of big data because they make big data easier to work with and cheaper to store. In the years since then, the volume of big data has skyrocketed. Users are still generating huge amounts of data—but it’s not just humans who are doing it.

With the advent of the Internet of Things (IoT), more objects and devices are connected to the internet, gathering data on customer usage patterns and product performance. The emergence of machine learning has produced still more data.

While big data has come far, its usefulness is only just beginning. Cloud computing has expanded big data possibilities even further. The cloud offers truly elastic scalability, where developers can simply spin up ad hoc clusters to test a subset of data.

5 of 137

Benefits of Big Data and Data Analytics

  • Big data makes it possible for you to gain more complete answers because you have more information.
  • More complete answers mean more confidence in the data—which means a completely different approach to tackling problems.

 

6 of 137

Figure 1.1 shows data usage and growth. As size and complexity increase, the proportion of unstructured data types also increase.

Figure 1.1 Evolution of Big Data and their characteristics

7 of 137

  • An example of a traditional tool for structured data storage and querying is RDBMS.
  • Volume, velocity and variety (3Vs) of data need the usage of number of programs and tools for analyzing and processing at a very high speed.

BIG DATA

  • Data is information, usually in the form of facts or statistics that one can analyze or use for further calculations.
  • Data is information that can be stored and used by a computer program.
  • Data is information presented in numbers, letters, or other form.
  • Data is information from series of observations, measurements or facts.

8 of 137

Web Data

  • Web data is the data present on web servers (or enterprise servers) in the form of text, images, videos, audios and multimedia files for web users.
  • A user (client software) interacts with this data.
  • A client can access (pull) data of responses from a server. The data can also publish (push) or post (after registering subscription) from a server.
  • Internet applications including web sites, web services, web portals, online business applications, emails, chats, tweets and social networks provide and consume the web data.

Examples of Web data

      • Wikipedia
      • Google Maps
      • YouTube
      • Face Book

9 of 137

Classification of Digital Data

10 of 137

Classification of Digital Data

11 of 137

Classification of Digital Data

12 of 137

Classification of Digital Data

13 of 137

Classification of Dataa

1.1.1 Structured Data

Structured is one of the types of big data and By structured data, we mean data that can be processed, stored, and retrieved in a fixed format. It refers to highly organized information that can be readily and seamlessly stored and accessed from a database by simple search engine algorithms. For instance, the employee table in a company database will be structured as the employee details, their job positions, their salaries, etc., will be present in an organized manner.

 

14 of 137

15 of 137

16 of 137

17 of 137

18 of 137

19 of 137

20 of 137

21 of 137

22 of 137

23 of 137

24 of 137

25 of 137

26 of 137

Classification of Data

1.1.3 Unstructured

Unstructured data refers to the data that lacks any specific form or structure whatsoever. This makes it very difficult and time-consuming to process and analyze unstructured data. Email is an example of unstructured data. Structured and unstructured are two important types of big data.

OR

27 of 137

28 of 137

29 of 137

30 of 137

31 of 137

Dealing with unstructured Data

32 of 137

33 of 137

34 of 137

35 of 137

36 of 137

Summary

37 of 137

38 of 137

39 of 137

40 of 137

41 of 137

42 of 137

43 of 137

44 of 137

45 of 137

46 of 137

47 of 137

48 of 137

49 of 137

50 of 137

51 of 137

52 of 137

53 of 137

54 of 137

55 of 137

56 of 137

57 of 137

2.10 A Typical Data Warehouse Environment

Data Sources in a DW Environment

- Data gathered from:

• ERP systems

• CRM systems

• Legacy systems

• Third-party applications

  • Formats: RDBMS (Oracle, SQL Server, MySQL, etc.),

spreadsheets (.xls, .xlsx), CSV, TXT

  • Sources may be local or geographically distributed

Refer to Figure 2.9 for a visual representation of the DW environment

58 of 137

2.10 A Typical Data Warehouse Environment

ETL: Extraction, Transformation, Loading

Extraction: Gather data from multiple sources

• Transformation: Clean, integrate, and standardize data

• Loading: Insert into enterprise data warehouse or data marts

Usage of Data Warehouse

- Data stored in:

• Enterprise Data Warehouse (EDW)

• Data Marts (functional/unit-specific)

- Tools used:

• Business intelligence & analytics tools

• SQL queries, dashboards, data mining

59 of 137

60 of 137

2.11 A Typical Hadoop Environment

61 of 137

62 of 137

63 of 137

64 of 137

65 of 137

66 of 137

67 of 137

68 of 137

69 of 137

70 of 137

71 of 137

72 of 137

73 of 137

74 of 137

75 of 137

76 of 137

77 of 137

78 of 137

79 of 137

80 of 137

81 of 137

82 of 137

83 of 137

84 of 137

85 of 137

86 of 137

87 of 137

88 of 137

89 of 137

90 of 137

91 of 137

92 of 137

93 of 137

94 of 137

95 of 137

96 of 137

97 of 137

98 of 137

99 of 137

100 of 137

101 of 137

102 of 137

103 of 137

104 of 137

105 of 137

106 of 137

107 of 137

108 of 137

109 of 137

110 of 137

111 of 137

112 of 137

113 of 137

114 of 137

Limitations of Hadoop 1.0

115 of 137

116 of 137

117 of 137

118 of 137

119 of 137

120 of 137

121 of 137

122 of 137

123 of 137

124 of 137

125 of 137

126 of 137

127 of 137

128 of 137

129 of 137

130 of 137

131 of 137

132 of 137

133 of 137

134 of 137

135 of 137

136 of 137

Questions:

  1. Define Big Data and discuss structured, semi-structured and unstructured data with examples.
  2. Explain the characteristics and challenges of Big Data.
  3. Compare traditional BI versus Big Data.
  4. What is Big Data Analytics and their classification?
  5. Discuss any three terminologies used in Big Data Environments.
  6. Differences between parallel system versus distributed system.
  7. Explain CAP theorem with different databases that follow one of the possible three combinations(AP,CP,CA).
  8. What is NoSQL and explain its features.
  9. Explain different types of NoSQL databases.
  10. Discuss the advantages of NoSQL.

137 of 137

Questions:

11. Compare SQL versus NoSQL.

12. Discuss the comparison of SQL, NoSQL and NewSQL.

13.Define Hadoop and explain features of Hadoop.

14. Discuss the key advantages of Hadoop.

15. Explain Hadoop component ecosystem with neat diagram.

16. Differentiate between Hbase and HDFS.

17. Compare Hadoop versus RDBMS.

18. Compare Hadoop versus SQL.