UK MapR SE & Benelux MapR DACH MapR Nordics MapR Japan MapR Hyderbad Company Profile § Founded in 2009 § Came out of stealth in 2011 § Deep management bench with extensive analyLc, storage, virtualizaLon and open source experience – Google, EMC, MicrosoN, InformaLca, Cisco, VMWare, NetApp, IBM, MicrosoN, Apache FoundaLon, Aster Data, Brio § Worldwide presence – Engineering and support in California and Hyderabad – Sales and field engineering in US, UK, France, Germany, Sweden, Singapore, Japan, Korea, Australia § 1000s of deployments including: – 10+ of Fortune 100 companies in producLon
Automated stateful failover § Automated re-‐replicaLon § Self-‐healing from HW and SW failures § Load balancing § Rolling upgrades § No lost jobs or data § 99999’s of upLme Reliable Compute Dependable Storage § Business conLnuity with snapshots and mirrors § Recover to a point in Lme § End-‐to-‐end check summing § Strong consistency § Data safe § Mirror across sites to meet Recovery Time ObjecLves
POSIX compliant – Random reads/writes – Simultaneous reading and wriLng to a file – Compression is automaLc and transparent § Industry-‐standard NFS interface (in addiLon to HDFS API) – Stream data into the cluster – Leverage thousands of tools and applicaLons – Easier to use non-‐Java programming languages – No need for most proprietary Hadoop connectors
management suite for Hadoop – Health monitoring – Cluster administraLon – ApplicaLon resource provisioning – Job monitoring and management – Job and data placement control – Security § MulLple interfaces: – GUI – REST API – CLI
self-‐heal • No pracLcal limit on # of files No-‐NameNode architecture • Jobs are not impacted by failures • Meet your data processing SLAs JobTracker HA • High throughput and resilience for NFS-‐based data ingesLon, import/export and mulL-‐client access NFS HA • Files and tables are accessible within seconds of a node failure or cluster restart Instant recovery • Upgrade the soNware with no downLme Rolling upgrades • No special configuraLon to enable HA • All MapR customers operate with HA HA is built-‐in
DataNode DataNode DataNode DataNode DataNode DataNode DataNode No NameNode Architecture Other DistribuLons (HDFS FederaLon) MapR § Single point of failure § Limited to 50M files per NameNode § Performance bocleneck § Metadata must fit in memory § HA w/ automaLc failover and re-‐replicaLon § Up to 1T files (> 5000x advantage) § Higher performance § Metadata is persisted to disk A F C D E D B C E B C F B F A B A D E DataNode DataNode DataNode
• Protect from hardware failures • File chunks, table regions and metadata are automaLcally replicated (3x by default) • At least one replica on a different rack Snapshots • Protect from user and applicaLon errors • Point-‐in-‐Lme recovery • No data duplicaLon • No performance or scale impact • Read files and tables directly from snapshot C1 C2 C3 C1 C2 C4 C1 C4 C4 C2 C5 C5 C6 C3 C5 C6 C3 C6 C7 C7 C7 Ac#ve&Volume Snapshot 13505505.09500 A B C D D₁
Engine § ElasLc Resource AllocaLon § Launch your first cluster in minutes § Only pay for what you use § No upfront expenses or long-‐term commitments § Launch parallel clusters for simultaneous access by different users § If your needs change ... no problem! It’s easy to change cluster size, node types, etc. § No need to worry about launching and managing Hadoop clusters