Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
InfluxDB - a distributed events and time series...
Search
Sponsored
·
Ship Features Fearlessly
Turn features on and off without deploys. Used by thousands of Ruby developers.
→
Paul Dix
April 27, 2014
Technology
2k
1
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
InfluxDB - a distributed events and time series database
Slides from my lightning talk at the GopherCon pre-party.
Paul Dix
April 27, 2014
More Decks by Paul Dix
See All by Paul Dix
InfluxDB IOx Project Update - 2021-02-10
pauldix
0
280
InfluxDB IOx data lifecycle and object store persistence
pauldix
1
720
InfluxDB 2.0 and Flux
pauldix
1
800
Flux and InfluxDB 2.0
pauldix
1
1.5k
Querying Prometheus with Flux
pauldix
1
1k
Flux (#fluxlang): a new (time series) data scripting language
pauldix
7
5.4k
At Scale, Everything is Hard
pauldix
2
780
IFQL and the future of InfluxData
pauldix
2
1.5k
Time series & monitoring with InfluxDB and the TICK stack
pauldix
0
530
Other Decks in Technology
See All in Technology
地方移住と都心キャリアの両立は「金・時間・人」のリソースをフル活用すれば実現できる!〜Snowflake女子会 vol.8
snowwmn0824
0
110
Railsのように考える: See through the Master
snoozer05
PRO
5
1.3k
GoのInterface内部構造から学ぶ!最高パフォーマンスを出すコード設計
yappli_developers
1
290
JSONataとAWS Step Functionsで目指すRuntimelessな世界
mu7889yoon
1
490
人にやさしく、AIにやさしく、書き手を選ばないIaCのガードレール再考 / Rethinking IaC Guardrails for Humans and AI Alike
kohbis
5
1k
Lambda MicroVMsは常駐サーバーの代わりに なるか? Kiro Crew を動かして検証してみた / Kiro Crew on Lambda MicroVMs
k_adachi_01
2
210
AIエージェントを安全で速い現場監督にする:Jev・Obsidian・メタハーネス
x5gtrn
PRO
0
110
品質と信頼性を地続きにする
grimoh
2
1k
AIに賢く動いてもらうためのコンテキスト〜Snowflake女子会 vol.8
snowwmn0824
0
170
AI Agent入門〜今更聞けないAgentの話〜
hiromimaganuma
0
110
AIに任せた品質は、誰が見立てるのか - AI時代のテストマネジメント
nakanao
3
3k
人間はどの意思決定を手放せるのか
kawasima
16
8.2k
Featured
See All Featured
Money Talks: Using Revenue to Get Sh*t Done
nikkihalliwell
0
500
The Curse of the Amulet
leimatthew05
3
15k
Heart Work Chapter 1 - Part 1
lfama
PRO
10
36k
Navigating the Design Leadership Dip - Product Design Week Design Leaders+ Conference 2024
apolaine
2
450
Why Your Marketing Sucks and What You Can Do About It - Sophie Logan
marketingsoph
0
410
Scaling GitHub
holman
464
140k
How to Get Subject Matter Experts Bought In and Actively Contributing to SEO & PR Initiatives.
livdayseo
0
200
The #1 spot is gone: here's how to win anyway
tamaranovitovic
4
1.2k
HU Berlin: Industrial-Strength Natural Language Processing with spaCy and Prodigy
inesmontani
PRO
0
720
Save Time (by Creating Custom Rails Generators)
garrettdimon
PRO
32
4.9k
What Being in a Rock Band Can Teach Us About Real World SEO
427marketing
0
1.1k
Fireside Chat
paigeccino
43
4k
Transcript
InfluxDB - a distributed time series, metrics, and events database
Paul Dix paul@influxdb.com @pauldix @influxdb
YC (W13), 3 people full time: Todd Persen John Shahid
Paul Dix (me)
What it’s for…
Metrics
Time Series
Analytics
Events
Can’t you just use a regular DB?
order by time?
Doesn’t Scale
Example from metrics: ! 100 measurements per host * 10
hosts * 8640 per day (once every 10s) * 365 days ! = 3,153,600,000 records per year
Have fun with that table…
But wait, we’ll just keep the summaries!
1h averages = ! 8,760,000 per year
Lose Detail and AdHoc Queryability
So let’s use Cassandra, HBase, or Scaleasaurus!
Too much application code and complexity
Application logic and scripts to compute summaries
Application level logic for balancing
No data locality for AdHoc queries
And then there’s more…
Web services
Libraries for web services
Data collection
Visualization
–Paul Dix “Building an application with an analytics component today
is like building a web application in 1998. You spend months building infrastructure before getting to the actual thing you want to build.”
Analytics should be about analyzing and interpreting data, not the
infrastructure to store and process it.
None
HTTP API Web services built in
HTTP API (writes) curl -X POST \ 'http://localhost:8086/db/mydb/series?u=paul&p=pass' \ -d
'[{"name":"foo", "columns":["val"], "points": [[3]]}]'
Data (with timestamp) [ { "name": "cpu", "columns": ["time", "value",
"host"], "points": [ [1395168540, 56.7, "foo.influxdb.com"], [1395168540, 43.9, "bar.influxdb.com"] ] } ]
HTTP API (queries) curl 'http://localhost:8086/db/mydb/series?u=paul&p=pass&q=.'
SQL-ish select * from events where time > now() -
1h
SQL-ish select * from “series with weird chars ()*@#0982#$” where
time > now() - 1h
Where Regex select line from application_logs where line =~ /.*ERROR.*/
and time > "2014-03-01" and time < "2014-03-03"
Only scans the time range Series and time are the
primary index
Work with many series…
Select from Regex select * from /stats\.cpu\..*/ limit 1
Downsampling on the fly…
Aggregates select percentile(90, value) from response_times group by time(10m) where
time > now() - 1d
Continuous Downsampling…
Continuous queries (summaries) select count(page_id) from events group by time(1h),
page_id into events.[page_id]
Series per page id select count from events.67 where time
> now() - 7d
Continuous queries (regex downsampling) select percentile(value, 90) as value from
/stats\.*/ group by time(5m) into percentile.90.:series_name
Percentile series per host select value from percentile.90.stats.cpu.host1 where time
> now() - 4h
Denormalization for performance
Range scans all user events for last hour select *
from events where user_id = 3 and time > now() - 1h
Continuous queries (fan out) select * from events into events.[user_id]
Series per user id select * from events.3 where time
> now() - 1h
Distributed Scale out, data locality, high availability
Raft for metadata We owe Ben Johnson a beer or
three…
Protobuf + TCP for queries, writes
Scalable Have billions of points in 1 series* or a
million different series
Libraries Go, Ruby, Javascript, Python, Node.js, Clojure, Java, Perl, Haskell,
R, Scala, CLI (ruby and node)
Visualization
Built-in UI
Grafana
Javascript library + D3, HighCharts, Rickshaw, NVD3, etc. Definitely more
to do here!
Data Collection CollectD Proxy, StatsD backend, Carbon ingestion, OpenTSDB (soon)
Coming Soon
ugh, Documentation
Series Metadata
Binary Protocol
Pubsub select * from some_series where host = “serverA” into
subscription() select percentile(90, value) from some_series group by time(1m) into subscription()
Custom Functions select myFunc(value) from some_series
Rack aware sharding and querying
Multi-datacenter replication Push and bi-directional
Indexes?
Ponies? Tell @jvshahid that you want your pony ;)
But it’s ready to go now. Production deployments already running.
Need help? support@influxdb.com Thanks! paul@influxdb.com @pauldix