Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Features
Speaker Deck
PRO
Sign in
Sign up for free
Search
Search
Large scale distributed systems patterns
Search
Ryosuke Iwanaga
September 22, 2025
Technology
110
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
Large scale distributed systems patterns
Ryosuke Iwanaga
September 22, 2025
Other Decks in Technology
See All in Technology
Kiro Crew入門 - 常駐エージェントの仕組みと使いどころ / Intro to Kiro Crew
k_adachi_01
1
800
8bit CPU 2026
koba789
6
2.7k
AI時代に、人は何を、どう学ぶのか #pbl_pub / What and How Do We Learn in the AI Era?
takaking22
1
240
平文パスワードはログに“残り” ── 肝心の侵入は“痕跡すら残らない”
kuroneko13
0
120
医療の現場を変革に挑戦した半年間の軌跡 - PythonとAIで現場を変える / From Code to Care
soudai
PRO
1
570
強化学習「理論」入門
enakai00
3
3.6k
DDDのエッセンスを取り入れたAIでの開発
ak2ie
1
210
Sets in Go
ramalho
1
1.2k
RelayerというPHPのフレームワークを作った
polidog
PRO
0
120
LLM・AIエージェントシステムベストプラクティス
shibuiwilliam
6
1.6k
dbt in Microsoft Fabric
ryomaru0825
0
160
AIに狂うスタートアップが、あえて「人との協働」に全振りした新卒エンジニア研修 / New Graduate Engineer Training at a Startup Accelerating AI Adoption
ohnoeight
0
180
Featured
See All Featured
Responsive Adventures: Dirty Tricks From The Dark Corners of Front-End
smashingmag
254
22k
Put a Button on it: Removing Barriers to Going Fast.
kastner
60
4.5k
Dominate Local Search Results - an insider guide to GBP, reviews, and Local SEO
greggifford
PRO
0
310
KATA
mclloyd
PRO
35
15k
Why Our Code Smells
bkeepers
PRO
340
58k
Navigating Team Friction
lara
192
16k
Groundhog Day: Seeking Process in Gaming for Health
codingconduct
0
290
Writing Fast Ruby
sferik
630
63k
Heart Work Chapter 1 - Part 1
lfama
PRO
8
36k
The innovator’s Mindset - Leading Through an Era of Exponential Change - McGill University 2025
jdejongh
PRO
1
280
Making the Leap to Tech Lead
cromwellryan
135
10k
Fireside Chat
paigeccino
42
4k
Transcript
Large scale distributed systems patterns Ryosuke Iwanaga
Agenda Visit architecture patterns and talk about problems * Web
app * Cloud native * Microservice in scale * Resource management * Event system * Problem 1: Cold start * Problem 1: Poison pill
Distributed systems always fail => Design for failure Takeaways One
solution introduces another problem => Design exercise
2009 Mobile browser gaming (SRE / DBA) * ~5,000 physical
servers * MySQL replications, sharding * Datacenter operations AI! 2025 Cloud (Solutions architect) Distributed datastore (Developer) 2015 2018 * Architecture * Container * Analytics * Distributed system * ~50 Microservices * Horizontal scale * Cell-based My experience in large scale distributed systems
LB ... User Web app distributed system Replication delay Deadlock
Write bottleneck LB's scalability Typical problems: App App App App DB Writer DB Reader DB Reader ... DB Writer DB Reader DB Reader ... ... Payment
Write Read Read LB ... App App App App DB
DB DB ... Cache Cache Cache ... Web app distributed system Cache invalidation Cache scalability Typical problems:
... Server 1 Resource orchestrator/manager App1 Amazon EC2, Eucalyptus, OpenStack
Hadoop, Mesos, YARN, Omega, Borg, k8s Server 2 Server N App1 App1 App1 App2 Cloud resource distributed system App3 App2 App3 Manager’s scalability Consistency Typical problems:
App Stream Speed layer Event distributed system e.g. Lambda architecture
App App Stream process 1 Object storage Stream process 2 Batch process 1 Batch process 2 Batch layer At least once At most once Stream scalability Back pressure Typical problems:
App Service A ... User 1 Metadata User 1,3,4 User
1 => DB 1 User 2 => DB 2 User 3 => DB 1 ... 💀 DB 1 DB 2 Service B App App User 2 LB Service C User 2,5 Microservice distributed system ...
App Service A ... User 1 Metadata User 1,3,4 User
1 => DB 1 User 2 => DB 2 User 3 => DB 1 ... DB 1 DB 2 Service B App App User 2 LB Service C User 2,5 Microservice distributed system ... Cache 😁?
Warm start 🔄Restart ✅ https://aws.amazon.com/message/11201/ Cold start Metadata App App
App App App App App 🔄Restart App App App App App App App 🔄Restart 🔄Restart 🔄Restart 🔄Restart 🔄Restart 💀 Cold start problem
💊 If user 1's requests trigger a bug on app
that crashes the app... 💀 Retry Retry Retry Retry Retry Retry Retry 💀 0% availability => App User 1 App App User 2 App App App App App User 3 💀 💀 💀 💀 💀 💀 💀 LB Poison pill problem 0% availability => 0% availability =>
💊 💀 ✅ App User 1 App App User 2
App App App App App User 3 💀 Naive sharding 💀 💀 0% availability => 100% availability => 0% availability =>
💊 ✅ App User 1 App App User 2 App
App App App App User 3 💀 Shuffle sharding 💀 ✅ ✅ https://aws.amazon.com/blogs/architecture/shuffle-sharding-massive-and-magical-fault-isolation/ : Server set for user 1, 2 5 overlap (k=5): 0.00000013% 4 overlap (k=4): 0.00063% 3 overlap (k=3): 0.059% 2 overlap (k=2): 1.8% 1 overlap (k=1): 21% 0 overlap (k=0): 77% 💀 50% availability => 100% availability => 0% availability => : Number of total servers : Size of each shard : Overlap between user 1 and 2
What’s next? App Service A ... User 1 Metadata DB
1 DB 2 Service B App App User 2 Service C ... App Service D ... App App Metadata LB Service E 💀
Distributed systems always fail => Design for failure Takeaways (again)
One solution introduces another problem => Design exercise
Thanks! @riywo OpsBR