In Big Data, Hadoop components such as Hive (SQL construct), Pig ( Scripting construct), and MapReduce (Java programming) are used to perform all the data transformations and aggregation. Now, with Apache Spark, the same is being achieved with many more advantages like unified API performance, support for multiple languages, and 10X-100X faster than MapReduce. Spark provides a single platform with SQL, Scripting, as well as the programming construct.

Big Data (Setting up the context)

The amount of data has grown considerably in recent years due to the growth of social networking, education, surveillance cameras, healthcare, business, satellite images, manufacturing, online purchasing, research analysis, banking, bioinformatics, Internet of Things, criminal investigation, media, information technology, etc. This huge volume of data in the world has created a new field of data processing which is called Big Data.

Data can be private or public

So, to do something meaningful with the data, we have to convert the unstructured data which is messy and semantically complex into structured data which is clean and easy to consume. This is called Data Processing.

Data Processing Tasks

The complexity of the data can be measured by the messiness and speed of scaling of data as explained in the below points.

Tools for Data Processing

Apache Spark

Apache Spark

Apache Spark is an open-source, lightning, cluster-computing framework. It is an engine for data processing and analytics.

Features

Characteristics

In order to work with Spark we have to use Spark APIs like,

Almost all the data is processed using specific data structures called RDDs (Resilient Distributed Datasets).

Components of Spark

Spark

Storage System and Cluster Manager, are both plug-and-play components.

Apache Spark Ecosystem

Apache Spark

In the next article, we will discuss more regarding RDDs and will learn how to load a data set.