The Disconnect The Need What is BigData? Characteristics Approach Architectural Requirements Techniques Challenges Solutions Issues Deep Dive – Practical Approaches to Big Data Hadoop Aster Data
to fill 60,000 Library of Congress (LoC collected 235TB in Apr/2011) YouTube receives 24hours of video, every minute 5 Billion mobile phones in use in 2010 Tesco (British Retailer) collects 1.5 billion pieces of information to adjust prices and promotions Amazon.com: 30% of sales is out of its recommendation engine Planecast, Mobclix : Track & Target systems promotes contextual promotions A Boeing Jet Engine produces 20TB/Hour for engineers to examine in real time to make improvements Sources: Forrester, The Economist, McKinsey Global Institute
Gateways Customer Information CRM Product Information Barcodes RFID Web Pages Web Repositories Unstructured Information Social Media Signals Mobile GPS, GeoSpatial
Divide and Access Data Consolidation Broader Tables Access all as a row Fine Grain Access Security Rules and Policies Problems Data Growth When storage cost is not an issue Scalability Issues Performance Issues New types of requirements Deciding what to analyze, when and how? Cost of a change in the subject-area to analyze
Old DBMS vs. New volume Old DBMS vs. New Analysis Old DBMS vs. Data Retention Old DBMS vs. Data Element Striping Old DBMS vs. Data Infrastructure Old DBMS vs. One DB Platform for all
data, in high volume, in high velocity with varied requirements to mine them” Characteristics Size Scale up and scale out: Terabyte, Petabyte … Structure Structured Unstructured : Audio, Video, Text, GeoSpatial Schema Less Structures Stream Torrent of real-time information Operation Massively Parallel Processing (MPP)
Framework - Java Inspired by Google’s white papers on Map/Reduce (MR) Google File System (GFS) Big Table Originally developed to support Apache Nutch Designed Large scale data processing For batch processing For sophisticated analysis To deal with structured and unstructured data DB Architect’s Hadoop : "Heck Another Darn Obscure Open-source Project"
and software platforms Shared-nothing architecture Scale hardware when ever you want System compensates for hardware scaling and issues (if any) Run large-scale, high volume data processes Scales well with complex analysis jobs (Hardware) “Failure is an option” Ideal to consolidate data from both new and legacy data sources Highly Integrable Value to the business Why Hadoop?
for Clustered, Distributed data processing ZooKeeper Scheduler Avro Data Serialization Chukwa Data Collection System to monitor Distributed Systems HBase Data storage for distributed large tables Hive Data warehouse Pig High-Level Query Language Scribe Log Collection UDF User Defined Functions Hadoop Ecosystem
Tolerant Handle large volumes of data Provides High Throughput Streaming data-access Simple file coherency model Portable to heterogeneous hardware and software Robust Handles disk failures, replication (& re-replication) Performs cluster rebalancing, data integrity checks Hadoop Distributed File System HDFS
separate chunk’s Processed by map tasks, in parallel Sorts the output of the maps Processed by reduce tasks, in parallel Typically stored and processed in a file system Framework takes care of Scheduling tasks Monitoring Re-executing failed tasks Infrastructure issues Load-balancing, Load-redistribution Replication, Failover Hadoop M/R
takes you or you die" Jim Starkey – Founder and CTO NimbusDB "Our philosophy is to build infrastructure using the best tools available for the job and we are constantly evaluating better ways to do things when and where it matters." Facebook "In any year we probably generate more data than the Walt Disney Co. did in the first 80 years of existence" Bud Albers - Disney