Every major Azure incident, analyzed: a timestamped timeline, the root cause in plain language, the business impact, and whether it earned SLA credits.
A bug in an automated cryptographic key rotation process left an old signing key in a state where Azure AD could no longer validate the tokens it issued. Sign-in to Microsoft 365, the Azure portal, and any app that authenticates through Azure AD failed worldwide for several hours until Microsoft rolled the key back.
Three overlapping bugs in the Azure multi-factor authentication service combined so that MFA requests could not complete. Because so many tenants require MFA at sign-in, users could not finish authenticating to Microsoft 365 and Azure AD worldwide. It took Microsoft the better part of a day to fully mitigate the layered failure.
A severe thunderstorm near the South Central US datacenter caused voltage swells that damaged cooling and electrical infrastructure. Hardware overheated and shut down to protect itself, taking a large slice of the region offline. Recovery took more than a day, and dependencies on the affected region briefly rippled into some services worldwide.
A configuration change meant to improve Azure Storage performance contained a bug and was rolled out far more broadly than intended, bypassing the normal staged deployment. Storage front ends entered a loop and stopped serving requests, and because so many Azure services depend on Storage, the failure cascaded worldwide for hours.
An allowlist configuration error on Azure Storage scale units made storage unavailable across an availability zone in Central US. Because Virtual Machines cannot run without their disks, the failure cascaded into VMs, dependent Azure services, and Microsoft 365 for many hours - landing in the same week as the unrelated CrowdStrike incident.
Layer-7 DDoS attacks - attributed by Microsoft to the actor it tracks as Storm-1359 - flooded the web front ends of the Azure Portal, Outlook on the web, and OneDrive across several days in early June 2023. Authentication and portal access failed intermittently, showing how an attack on web tiers reads as an identity outage to the business.
During a planned router addition to Microsoft’s global WAN, a command with unintended effects caused routers to forward packets incorrectly worldwide. Azure, Teams, and Outlook were disrupted globally - worst for roughly 90 minutes - with full recovery within hours after the change was rolled back.
A name-server delegation change made during a planned Azure DNS migration disrupted resolution for Microsoft services globally. Because DNS sits in front of everything, customers could not reach Storage, SQL Database, and Microsoft 365 even though those services were themselves healthy - for roughly three hours.
7 min read
Recent incident log
Smaller incidents from the live feed (last 90 days). Major events graduate into full post-mortems above.
Active -Multiple Service impacted affecting West US Regions