problems! 4/1/2010 4/23/2012 5/3/2010 TFS 2010 RTM 4/23/2011 Service Deployment 8/5/2011 Service Update 9/26/2011 //BUILD 2011 12/7/2011 Service Update 1/30/2012 Service Update 2/20/2012 Service Update 3/12/2012 Service Update 4/2/2012 Service Update
actionable and represent a real issue with the system. Alerts should create a sense of urgency – false alerts dilutes that Redundant alerts for same the issue Needed to set right thresholds and tune often Stateless alerts contributed to further noise
performance • All 3 related to same code defect • APM component mapped to feature team • Auto-dialer engaged Global DRI Eliminated alert noise ~928 alerts per week to ~22 and reduced DRI escalations by ~56%
DRAFT Microsoft Confidential 52 Service Availability & Health Metrics DRAFT DRAFT DRAFT Incident Count Incident Count DRAFT DRAFT DRAFT % of Incidents User Minutes DRAFT DRAFT DRAFT Error By Source Incidents by Severity User Impact Minutes During Incidents [TFS Only] 3 2 1 4 1. TFS Availability is on an improving trend. No Sev0/Sev1 LSIs for July. 2. App Insights switched from synthetic availability to real-user experience in Ibiza portal. A high volume of SEV-2 LSIs (72) contributed to customer impact in addition to intermittent UX errors. (UX fixes applied on 8/11 that improves availability) 3. App Insights was impacted by 3 long running LSIs related to ES maintenance, Ibiza updates and an Azure Storage outage. 4. TFS Service attainment (SLO) improved significantly MoM with focus on minimizing failed/slow commands and reviewing in weekly LiveSite reviews