1 2 3 Cloudera: The Hybrid Data Platform Company Enabling Analytics at Scale Q&A Streaming Data Lifecycle & Machine learning for Generative AI value creation
THE HYBRID DATA PLATFORM COMPANY Manage and secure the data lifecycle in any cloud or datacenter SECURITY | GOVERNANCE | CATALOG | METADATA | INTELLIGENCE 01 03 04 05 STREAMING & DATA FLOW DATA ENGINEERING DATA WAREHOUSE OPERATIONAL DATABASE MACHINE LEARNING & AI 02 Collect Enrich Report Serve Predict
AND INSIGHTS ANYWHERE Driving enterprise business value REAL-TIME STREAMING ENGINE ANALYTICS & DATA WAREHOUSE DATA SCIENCE/ MACHINE LEARNING CENTRALIZED DATA PLATFORM STORAGE & PROCESSING ANALYTICS & INSIGHTS Stream Ingest Ingest – Data at Rest Deploy Models BI Solutions SQL Predictive Analytics • Model Building • Model Training • Model Scoring Actions & Alerts [SQL] Real-Time Apps STREAMING DATA SOURCES Clickstream Market data Machine logs Social ENTERPRISE DATA SOURCES CRM Customer history Research Compliance Data Risk Data Lending
• Highly reliable distributed messaging system • Decouple applications, enables many-to-many patterns • Publish-Subscribe semantics • Horizontal scalable • Efficient implementation to operate at speed with big data volumes • Organized by topic to support several use cases Source System Source System Source System Kafka Hadoop Security Systems Real-Time Monitoring Source System Source System Source System Hadoop Security Systems Real-Time Monitoring Many-To-Many Publish-Subscribe Point-To-Point Request-Response
USE-CASES Disaster Recovery Geo-Locality Data Movement / Deployment Centralized Analytics Workload Isolation Legal / Compliance In an event of a partial or complete datacenter disaster, providing failover/failback to a secondary cluster in a different region / DC Active-active geo-localized deployments allows users to access a near-by data center to optimize their architecture for low latency and high performance. Use Kafka to synchronize data between on-prem applications and cloud deployments Aggregate data from multiple Kafka clusters into one location for organization-wide analytics Creation of different envs for SDLC: Dev, Test, Prod. Clusters for specific use case cases (ETL, ingestion, analytics, etc) Different data storage and security policies require clusters to be created in region but data still needs to be shared.
Bounded and Unbounded join(s) Write Streaming Result to Kudu Join 2 Streaming User Event Topics Enrich Stream from Warehouse HR Table Enrich Stream from RT Mart Timesheet Table Filter & Transform
Warehouse with Apache Kudu CDF IOT Devices Applications Metrics Logs & Files HDFS/ Object Storage Hot Storage Cold Storage SQL ◦ s u p p o r t Real-Time Analytics Alerting Event Driven Applications Dashboards
a Generative AI Model DA Scope Adapt and Align model Application Integration Select Define the use case Choose an existing model Prompt Engineering Evaluate Integrate a model and build LLM- powered applications Optimise and deploy your model for inference Augmentatio n Fine Tuning Inspired by, Source Citation: Andrew Ng, DeepLearning.AI , Generative AI with LLMs course Data Collection and Preparation
LLMS DIFFERENT THAN ANYTHING BEFORE? Simplicity, speed and scale, over all your data! TECH SPARK / MAPREDUCE HIVE SQL / SQL LLMs PERSONAS/ SKILLS Programmers • Complex coding Analysts • Semantic queries Everyone • Natural language RESPONSE TIME Hours Seconds to minutes Milliseconds to seconds DATA PB Scale Most Data • Slower processing • Structured data • Semi-structured data TB scale with ETL High Value Data • Structured data • Rest is ETL-ed out PB scale All Data • Structured data • Semi-structured data • Unstructured data