Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
Druid + R
Search
Metamarkets
April 03, 2013
Technology
210
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
Druid + R
Metamarkets
April 03, 2013
More Decks by Metamarkets
See All by Metamarkets
R Workshop for Beginners
metamx
2
4.7k
Other Decks in Technology
See All in Technology
Screen Lens - 今見てる画面を翻訳する
komagata
0
260
2026-09-09 【sigma_ucj#1】Sigma を IaC 管理したい! / IaC for Sigma
civitaspo
0
130
データ界隈LT祭 第1回LT登壇
taromatsui_cccmkhd
1
970
リージョンの壁を越える、 ちょっと変わったAWSサービスの話
falken
PRO
1
310
Sigmaユーザーのための有用リソース一挙公開 & Sigmaで使えるMCP #sigma_ucj /useful-resources-for-sigma-computing-users-and-mcps-with-sigma
shinyaa31
0
200
俺の仕事は AIに奪われないし、たぶんその BIも要らない
hikaruri
0
280
ASTを使って影響範囲を特定する
nealle
0
180
絵ではじめるKubernetesセキュリティ
aoi1
3
230
事業課題から技術的負債に向き合う
sansantech
PRO
2
530
銀行勘定系システムにおける開発プロセス刷新×AIによる環境モダナイゼーション / Development Process Transformation and AI-Driven Environment Modernization
muit
0
850
GoにおけるFFIのこれまでとこれから
goccy
5
2.6k
10Xに技術的負債をもたらした「2つの境界の歪み」その構造と解消への営み
10xinc
0
700
Featured
See All Featured
Leadership Guide Workshop - DevTernity 2021
reverentgeek
1
370
I Don’t Have Time: Getting Over the Fear to Launch Your Podcast
jcasabona
35
2.8k
Dealing with People You Can't Stand - Big Design 2015
cassininazir
367
27k
RailsConf & Balkan Ruby 2019: The Past, Present, and Future of Rails at GitHub
eileencodes
141
35k
Conquering PDFs: document understanding beyond plain text
inesmontani
PRO
4
3k
The browser strikes back
jonoalderson
0
1.7k
Intergalactic Javascript Robots from Outer Space
tanoku
273
27k
Navigating Weather and Climate Data
rabernat
0
520
Raft: Consensus for Rubyists
vanstee
141
7.7k
Building Flexible Design Systems
yeseniaperezcruz
330
41k
Building Experiences: Design Systems, User Experience, and Full Site Editing
marktimemedia
0
600
The Invisible Side of Design
smashingmag
301
52k
Transcript
Druid + R aggregate all your data
agenda An Overview of Druid RDruid Lab Conclusions
motivation visualize big data existing data engines did not meet
our needs
motivation relational databases scans were too slow! NoSQL computationally intractable
pre-computations took too long! nothing existed that could solve our problems (or was cost prohibitive)
enter Druid real-time distributed column-oriented analytical data store scales horizontally
open-source
how is Druid different highly optimized fast scans & aggregations
real-time data ingestion explore events within milliseconds no pre-computation arbitrarily slice & dice data highly available
using Druid we will explore Druid architecture in future meetups
let's learn to use Druid!
RDruid slicing & dicing on steroids
what are we addressing? slicing and dicing data in R
is fun… …until you run out of memory
solution fire up a 64G EC2 machine and hope it
works or let Druid do the work for you
how we use it ad-hoc reporting analyze client data internal
metrics prototyping
metrics
let’s try it code bit.ly/YtJ1Xj
setup launch your favorite R environment install and load the
druid R package install.packages("devtools") install.packages("ggplot2") library(devtools) install_github("RDruid", "metamx") library(RDruid) library(ggplot2) druid-meetup.R
concepts Druid always computes aggregates events are based in time
Druid understands time bucketing dimensions along which to slice & dice metrics to aggregate
concepts think aggregates and group by in SQL SELECT hour(timestamp),
time page, language, dimensions sum(count) metrics GROUP BY hour(timestamp), page, language
data sources connect to our cluster druid <- druid.url("druid-meetup.mmx.io") Wikipedia
druid.query.dimensions(url = druid, dataSource = "wikipedia_editstream") druid.query.metrics(url = druid, dataSource = "wikipedia_editstream") Twitter dataSource = "twitterstream" x0-sources.R
timeseries Wikipedia page edits since January, by hour edits <-
druid.query.timeseries( url = druid, dataSource = "wikipedia_editstream", intervals = interval(ymd("2013-01-01"), ymd("2013-04-01")), aggregations = sum(metric("count")), granularity = "hour" ) qplot(data = edits, x = timestamp, y = count, geom = "line") x1-timeseries.R
filters what if I'm only interested in articles in English
and French enfr <- druid.query.timeseries( [...] granularity = "hour", filter = dimension("namespace") == "article" & ( dimension("language") == "en" | dimension("language") == "fr" ) ) x2-filters.R
group by let's break it out by language enfr <-
druid.query.groupBy( [...] filter = dimension("namespace") == "article" & ( dimension("language") == "en" | dimension("language") == "fr" ), dimensions = list("language") ) qplot(data = enfr, x = timestamp, y = count, geom = "line", color = language) x3-groupby.R
granularity arbitrary time slices granularity = granularity( "PT6H", timeZone =
"America/Los_Angeles" ) try out a few more P1D · P1W · P1M x4-timeslices.R
aggregations sum, min, max aggregations = list( count = sum(metric("count")),
total = sum(metric("added")) ) timestamp total count 1 2013-01-01 127232693 346895 2 2013-01-02 130657602 403504 3 2013-01-03 134643672 387462 x5-aggs.R
math you can do math too + - * /
constants aggregations = list( count = sum(metric("count")), added = sum(metric("added")), deleted = sum(metric("deleted")) ), postAggregations = list( average = field("added") / field("count"), pct = field("deleted") / field("added") * -100 ) x6-postaggs.R
more advanced all pages edited by users matching regex '^Bob.*'
druid.query.groupBy([...] intervals = interval(ymd("2013-03-01"), ymd("2013-04-01")), granularity = "all", single time bucket filter = dimension("user") %~% "^Bob.*", dimensions = list("user", "page") ) x7-advanced.R
academy awards stats awards <- druid.query.groupBy( url = druid, dataSource
= "twitterstream", intervals = interval(ymd("2013-02-24"), ymd("2013-02-28")), aggregations = list(tweets = sum(metric("count"))), granularity = granularity("PT1H"), filter = dimension("first_hashtag") %~% "academyawards" | dimension("first_hashtag") %~% "oscars", dimensions = list("first_hashtag")) awards <- subset(awards, tweets > 10) qplot(data=awards, x = timestamp, y = tweets, color = first_hashtag, geom="line") x8-awards.R
academy awards stats x8-awards.R
roll your own run your own Druid cluster github.com/metamx/druid/wiki/ Druid-Personal-Demo-Cluster
contribute fork us on github Druid github.com/metamx/druid RDruid github.com/metamx/RDruid
thank you