Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
How to scale a Logging Infrastructure
Search
Paul Stack
June 03, 2015
Technology
220
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
How to scale a Logging Infrastructure
Logging infrastructure using ELK + Kafka
Paul Stack
June 03, 2015
More Decks by Paul Stack
See All by Paul Stack
Infrastructure as Software
stack72
0
110
Mirror, Mirror on the way, what is the vainest metric of them all?
stack72
1
2.4k
Continuously Delivering Infrastructure to the Cloud
stack72
0
250
DevOops 2016
stack72
0
150
The Quest for Infrastructure Management 2.0
stack72
0
180
The Biggest Trick Consultants Ever Pulled was Telling The World Continuous Delivery is Easy
stack72
1
160
The Transition from Product to Infrastructure
stack72
0
98
Continuous Delivery - the missing parts
stack72
0
1k
Windows: Having its ass kicked by puppet and powershell
stack72
0
170
Other Decks in Technology
See All in Technology
セルフサービスのオブザーバビリティ基盤をOpenTelemetryで作る / Building a Self-Service Observability Platform with OpenTelemetry
ymotongpoo
3
430
特殊変数大全
dak2
0
150
全社共通データ基盤をつくる。ソニーのDatabricks活用とデータガバナンス設計の裏側
sony
0
130
10年欲しかった音楽管理アプリを、AIと一緒に作りはじめた
judau
1
170
AWS FinOps Agent 結局何が得意なの?
siromi
0
220
【データ横丁主催】AI Agentがコンテキストを使って仕事をした後、何が残るのか― 組織の経験を次の判断に引き継ぐ「Agent Memory」
shisyu_gaku
2
170
AWS App Runnerから Cloudflare Workersへ移行した話
ryota09
0
110
Azure Copilot Resiliency Agentをいろいろ試してみる
tomokusaba
0
130
雪かき部 #7 もう怖くない!SELECT文!
foursue
0
130
VS Code × GitHub Copilot での Fabric 開発
ryomaru0825
1
120
AgentCore Runtime上にAgentic Coding基盤を構築・展開する際の設計ポイントと限界点 / Design considerations and limitations when building an agentic coding platform on AgentCore Runtime
har1101
7
810
OpenClawでAzure DevOpsのWiki更新を自動化する - クラウドAIだけでは届かない場所へ
yutakaosada
0
120
Featured
See All Featured
Facilitating Awesome Meetings
lara
57
7.2k
[RailsConf 2023] Rails as a piece of cake
palkan
59
7k
Save Time (by Creating Custom Rails Generators)
garrettdimon
PRO
33
5k
Design in an AI World
tapps
1
340
Building AI with AI
inesmontani
PRO
1
1.3k
Sam Torres - BigQuery for SEOs
techseoconnect
PRO
0
550
Conquering PDFs: document understanding beyond plain text
inesmontani
PRO
4
3.1k
Testing 201, or: Great Expectations
jmmastey
46
8.3k
Rails Girls Zürich Keynote
gr2m
96
14k
Understanding Cognitive Biases in Performance Measurement
bluesmoon
32
3k
Money Talks: Using Revenue to Get Sh*t Done
nikkihalliwell
0
510
The Art of Programming - Codeland 2020
erikaheidi
57
14k
Transcript
How do you scale a logging infrastructure to accept a
billion messages a day? Paul Stack http://twitter.com/stack72 mail:
[email protected]
About Me Infrastructure Engineer for a cool startup :) Reformed
ASP.NET / C# Developer DevOps Extremist Conference Junkie
Background Project was to replace the legacy ‘logging solution’
Iteration 0: A Developer created a single box with the
ELK all in 1 jar
Time to make it production ready now
None
Iteration 1: Using Redis as the input mechanism for LogStash
None
None
Enter Apache Kafka
“Kafka is a distributed publish- subscribe messaging system that is
designed to be fast, scalable, and durable” Source: Cloudera Blog
Introduction to Kafka • Kafka is made up of ‘topics’,
‘producers’, ‘consumers’ and ‘brokers’ • Communication is via TCP • Backed by Zookeeper
Kafka Topics Source: http://kafka.apache.org/documentation.html
Kafka Producers • Producers are responsible to chose what topic
to publish data to • The producer is responsible for choosing a partition to write to • Can be handled round robin or partition functions
Kafka Consumers • Consumption can be done via: • queuing
• pub-sub
Kafka Consumers • Kafka consumer group • Strong ordering
Kafka Consumers • Strong ordering
https://github.com/opentable/puppet-exhibitor
None
Iteration 2 Introduction of Kafka
None
None
Iteration 3 Further ‘Improvements’ to the cluster layout
None
The Numbers • Logs kept in ES for 30 days
then archived • 12 billion documents active in ES • ES space was about 25 - 30TB in EBS volumes • Average Doc Size ~ 1.2KB • V-Day 2015: ~750M docs collected without failure
What about metrics and monitoring?
Monitoring - Nagios • Alerts on • ES Cluster •
zK and Kafka Nodes • Logstash / Redis nodes
None
https://github.com/stack72/nagios-elasticsearch
Metrics - Kafka Offset Monitor
https://github.com/opentable/KafkaOffsetMonitor
Metrics - ElasticSearch
None
None
None
Visibility Rocks!
None
So what would I do differently?
Questions?
Paul Stack @stack72