U P P A V I N G T H E R O A D F O R P R O A C T I V E R E L I A B I L I T Y 35+ functional domains >26k+ apps 500+ Kubernetes clusters 40k+ nodes Multiple AWS regions Our Scale
U P P A V I N G T H E R O A D F O R P R O A C T I V E R E L I A B I L I T Y Before After Path to Paved Road Runtime Platform 1 Observability Tool 1 Service Mesh 1 Runtime Platform 2 Observability Tool 2 Service Mesh 2 Runtime Platform 3 Observability Tool 3 Service Mesh 3 Runtime Platform Observability Tool Service Mesh
U P P A V I N G T H E R O A D F O R P R O A C T I V E R E L I A B I L I T Y Change Everything Else https://www.subbu.org/articles/2019/incidents-trends-from-the-trenches/ Change Config Drift Unknown Infrastructure Failure Certificate Expiration Other Incidents grouped by theme Contributing factor that caused incidents Incident Analysis (2019)
U P P A V I N G T H E R O A D F O R P R O A C T I V E R E L I A B I L I T Y Canary Flow Blue/Green Release Safety Strategies Router Blue Green Initiate Shift % traffic Bake Judge Promote/ Rollback
U P P A V I N G T H E R O A D F O R P R O A C T I V E R E L I A B I L I T Y Progressive Deployment Primary Canary Baseline Initial State Primary Canary Baseline Traffic shifting 100% 0% 0% 50% 25% 25% V1 V2 V2
U P P A V I N G T H E R O A D F O R P R O A C T I V E R E L I A B I L I T Y Verify Autoscaling Understand Tolerated Disruptions Practice past incidents Prevent future incidents Chaos Engineering – Goals
U P P A V I N G T H E R O A D F O R P R O A C T I V E R E L I A B I L I T Y Chaos Engineering – Architecture Continuous Delivery Control Plane chaos-controller Block connectivity between Pod X and Pod Y running in Cluster Z Target Pod X in Cluster Z Cluster Z Pod X Experiment: Block connectivity Target Pod X Experiment: Block connectivity Experiment: Block connectivity Pod Y
U P P A V I N G T H E R O A D F O R P R O A C T I V E R E L I A B I L I T Y Region Failover as a Service – Use Cases Use case 1 – Cluster local (new platform) Region 1 app-A app-B app-B egress ingress Region 2 Cluster X Cluster X
U P P A V I N G T H E R O A D F O R P R O A C T I V E R E L I A B I L I T Y Region Failover as a Service – Use Cases Use case 2 – Cross Cluster (new platform) app-A app-A ingress app-C egress ingress Region 1 Cluster Y Region 1 Cluster X Region 2 Cluster X
U P P A V I N G T H E R O A D F O R P R O A C T I V E R E L I A B I L I T Y Region Failover as a Service – Use Cases Use case 3 – Legacy platform to cluster in new platform Region 1 app-A app-A ingress Region 2 Legacy Platform app-D ingress Cluster X Cluster X
U P P A V I N G T H E R O A D F O R P R O A C T I V E R E L I A B I L I T Y Byte Size Videos https://medium.com/expedia-group-tech Public Blogposts Internal Success Stories GameDays Promotion and Advocacy
U P P A V I N G T H E R O A D F O R P R O A C T I V E R E L I A B I L I T Y The internal Reliability Hub aims to: ü Provide best practices for engineers that want to design highly available and resilient systems ü Raise awareness of the Reliability Engineering solutions which are available ü Provide practical guidelines on how to verify resiliency using Chaos Engineering Reliability Hub
U P P A V I N G T H E R O A D F O R P R O A C T I V E R E L I A B I L I T Y 2021 2022 2023 Systems Design and Feedback Implementation Integrations Adoption Our Journey
O A D F O R P R O A C T I V E R E L I A B I L I T Y E X P E D I A G R O U P Key Takeaways • Analyse incident data to identify themes • Use tools to prevent and prepare for incidents • Produce data when building tools to measure adoption and drive investment decisions • Create a paved road and provide great developer experience • Integrate with existing platforms and tools • Bridge the gaps between people, processes, and tools 01 | Drive reliability through tools and data 02 | Focus on your Customers