Introduction
There is a lot of buzz in the industry regarding Big Data and naturally many questions and confusion. In this series of articles, I will attempt to help ease the understanding.
Big Data is a set of technologies that allows users to store data and compute leveraging multiple machines as a single entity. Think of it as a poor man's supercomputer.
One of the earliest examples of this technology is “Seti @ Home”, where radio signals were captured and a virtual supercomputer was leveraged using hundreds of thousands of individual personal computers, to analyze radio signals in search of intelligent life in our galaxy. Later on, similar technology was used to analyze human DNA for cancer research.
In early 2000, Google developed its own version, to analyze all the world's websites to power its search engine. Today, it's used behind all the major products powered by Google, such as Google Maps.
Later on, Google published a series of papers that were chosen by engineers at Yahoo, who eventually made the technology open source and hence here we are with the Big Data movement
Let's run through an example and see how this technology works and hopefully it will help answer some questions.
Let's calculate an average of numbers.
Today's Technology
Suppose we have 1 thousand numbers, all random positive integers.
Big Data
- Formula for Average = Sum(Numbers) / Count(Numbers)
- Formula for Average = Sum(Numbers) / Count(Numbers)
- On Machine M1, we will have X1 S1, C1
- On Machine M2, we will have X2 S2, C2
- On Machine M3, we will have X3 S3, C3
- On Machine M4, we will have X4 S4, C4
- On Machine M5, we will have X5 S5, C4
- Sum_Master = Sum (S1, S2, S3, S4, S5)
- Count_Master = Sum (C1, C2, C3, C4, C5)
- Average = Sum_Master / Count_Master
Observations
- Data needs to be stored on worker machines.
- Each worker machine works on its own chunk of data.
- A worker machine can work in Map mode or Reduced mode.
- Map mode: worker creates an intermediate variable.
- The Reduced mode collects data from Map workers and further reduces it to get the results.
- As you can see, it uses the divide and conquers approach, working in Batch mode.
- Highly Parallel works very well with this approach, so long as we can divide the task in smaller chunks of work and process it as a batch.
- Highly recursive tasks, such as a graph search and a recursion based algorithm do not perform well, we need something different and I will talk about it later.
- Most important, we need to break the simple Average formula to work in parallel and but not all the algorithms can be broken in such fashion.
- However, with the large quantity of data, we generally do not need very complex algorithms, simple algorithms work as well.
I hope this example was helpful, I will post a new article, with more details on the framework and the inner workings of data storage, coordination, and algorithms.
Resources

Hadshana KamalanathanPosted Jul 19, 2018, 12:36 AM
Thanks for sharing
Reshwanth SPosted Sep 15, 2017, 7:31 AM
Thank you sir for sharing your resource
Gowtham RajamanickamPosted Apr 13, 2016, 7:25 AM
expecting more from you
Gowtham RajamanickamPosted Apr 13, 2016, 7:25 AM
good article
Ankur MistryPosted Nov 29, 2015, 11:33 AM
Nice
Kamlesh BhorPosted Oct 16, 2015, 8:15 AM
Nice...
Harshad PansuriyaPosted Oct 16, 2015, 4:24 AM
Nice one
Nilesh JadavPosted Oct 16, 2015, 3:31 AM
Good Article sir !
Mukesh KumarPosted Oct 16, 2015, 3:29 AM
Good One, Thanks for sharing
Chiheb ChebbiPosted Feb 24, 2015, 6:06 PM
Great article and we need more posts about Big data
Prasham SabadraPosted Feb 23, 2015, 5:13 AM
Thanks for sharing. Nice Article.
Mohammad KhalidPosted Jan 21, 2015, 4:47 AM
Its a good article to learn this emerging technology. Willing to know about the technical part of Big Data.
Tapan PalPosted Jan 14, 2015, 6:13 AM
nice
Menish GuptaPosted Jan 8, 2015, 11:21 AM
Thank you for nice comments, I have posted one more article on HDFS, please let me know if you guys have any questions.
Pramod LawatePosted Dec 30, 2014, 11:46 AM
want join this programme
Vithal WadjePosted Dec 29, 2014, 10:00 PM
nice to know big data,thanks for introducing
Gaurav KumarPosted Dec 28, 2014, 11:04 PM
This topic is cool and interesting ....thanks@Menish Gupta for sharing ..
Gaurav Kumar AroraPosted Dec 28, 2014, 2:54 PM
I did not find many resources/articles on BigData but my search ends here. As a beginner, its good for me. Waiting for next article.
Shweta LodhaPosted Dec 27, 2014, 1:21 PM
Looking forward for more write-ups on this topic. Thanks for sharing
Mahesh ChandPosted Dec 27, 2014, 1:07 PM
Welcome aboard, Menish. Good to have expert trainers like yourself on the platform. Look forward to read the series. Cheers!
Manish Kumar ChoudharyPosted Dec 27, 2014, 12:31 PM
nice..
Michal HabalcikPosted Dec 27, 2014, 12:12 PM
looking forward to more resources on this topic, thanks for sharing
Menish GuptaPosted Dec 27, 2014, 12:03 PM
Jasminder, Yes you are correct, Big data technologies, leverage multiple machines (typically as low as 5 to hundred & thousands of machine), to store data and process data.
Jasminder SinghPosted Dec 27, 2014, 8:59 AM
Interesting example to start with. Looking forward to more such explanations. Just one thing to confirm, worker machines here are different physical servers and big data will always involve the use of distribution of the data on different servers ?