operations as if it’s a software problem. Our mission is to protect, provide for, and progress the software and systems with an ever-watchful eye on their availability, latency, performance, and capacity.
metrics for services that not be available through APM tools or profilers. For example, the average latency of Azure Storage, or the rate for messaging within an Azure IoT Hub. Azure Services Azure helps to maximize the availability and performance of applications and services. Azure delivers a solution for collecting, analyzing, and acting on telemetry from cloud and on-premises environments. www.ingenieriadelcaos.com .
experimenting on a system in order to build confidence in the system’s capability to withstand turbulent conditions in production. https://principlesofchaos.org
templates) • ARM is Infrastructure as Code for Azure. • A template is a JavaScript Object Notation (JSON) file. • A template uses declarative syntax, which lets you state what you intend to deploy without having to write the sequence of programming commands to create it. • A template specify the resources to deploy and the properties for those resources www.ingenieriadelcaos.com .
solid understanding of source code management, compilers, build configuration languages, automated build tools, package managers, and installers. 4 principles: Self-Service Model, High Velocity, Hermetic Builds, Enforcement Policies and Procedures. You expect to build 100% reliable services—ones that never fail. However, increasing reliability is worse for a service rather than better! Extreme reliability comes at a cost! Embrace the Risk! www.ingenieriadelcaos.com .
deployments is using the staging slots available in Azure App Service to stage a deployment before moving it to production. www.ingenieriadelcaos.com .
to generate alert rules that captures the target and criteria for alerting. Target Resource Defines the scope and signals available for alerting. • Virtual machines. • Storage accounts. • Log Analytics workspace. • Application Insights. Signals can be of the following types: metric, activity log, Application Insights, and log. www.ingenieriadelcaos.com .
constant iteration Analysis & Learning • Understand what happened and build new action plans and training • Siloed apps for reporting and analysis make it hard to summarize incidents, delaying post mortems Detection & Alerting • Monitoring systems send alerts for issues which need to be reviewed and escalated • Alerting systems generate noise which is lost in email, unclear where to escalate Remediation • Resolve the issues • Multiple tools, people and processes need to be coordinated for the implementation of long-term solutions for incidents. Process Overview & Challenges Containment • Data must be reviewed and the current situation assessed • Damage has to be contained quickly and efficiently to minimize the impact. Incident Management
responsibility for Microsoft and represents an investment that any customer using Microsoft Online Services can count on. Azure implements the five-stage process. Detect Assess Diagnose Stabilize Close www.ingenieriadelcaos.com .
network performance: • Instantly add scale to your applications • Load balance Internet and private network traffic • Improve application reliability via health checks • Flexible NAT rules for better security • Directly integrated into virtual machines and cloud services www.ingenieriadelcaos.com .