Microsoft Azure outage: Cloud services disrupted after West US network failure, now resolved

July 24, 2026 at 6:13 PM GMT+8

Microsoft services experienced widespread disruption on July 23, 2026, after a network failure in the company’s West US cloud region affected access to Azure and multiple Microsoft 365 services. A faulty maintenance automation process removed critical network routes, causing connectivity problems across Microsoft cloud services for several hours.

Microsoft confirmed the disruption after 10:30 A.M ET, stating that some customers in North America were experiencing difficulties accessing Microsoft 365 services with thousands of users reporting problems with Outlook, Teams, SharePoint, OneDrive, Copilot, Azure, Xbox Live, and the Microsoft Store. Some users also reported issues downloading Windows updates and installing Microsoft Office applications.

The company’s Service Health Status dashboard showed service degradation affecting multiple services through specific network paths.

Microsoft posted on X “We’ve confirmed some users in North America are experiencing issues accessing or using various Microsoft 365 services. We’re analyzing service telemetry and diagnostic data to isolate the source of impact.”

Microsoft later confirmed that the incident lasted from 14:44 UTC to 19:41 UTC, affecting connectivity to Azure and other cloud services hosted in the West US region.

The Preliminary post incident review (PIR) on Azure status history tracked the blackout timeline:

14:44 UTC: Device maintenance began, and customers started experiencing service disruption.

14:45 UTC: Microsoft networking teams and incident responders began investigating traffic anomalies, routing issues, and service degradation signals.

15:00–16:00 UTC: Engineers analyzed telemetry and customer reports to determine the scope of the outage.

16:00–17:45 UTC: Investigation identified abnormal routing behavior linked to recent network maintenance activity.

17:45 UTC: Microsoft began rolling back the maintenance changes.

18:26 UTC: Rollback completed, restoring network stability.

19:41 UTC: All affected services fully recovered.

The issue affected traffic moving into and out of the region although workloads communicating entirely within the West US were not impacted.

“The issue initially presented as large-scale route churn in our Wide-Area Network (WAN). It was later found that route removal was from a datacenter in the West US region. Once confirmed, engineers began roll back at 17:45 UTC. By 18:26 UTC, the network was restored and healthy, and all impacted services had fully recovered by 19:41 UTC,” said Microsoft on the status dashboard.

The outage affected a broad range of Microsoft cloud services, including Azure App Service, Application Gateway, Azure Kubernetes Service (AKS), Azure Cosmos DB, Azure Database for PostgreSQL, Azure Monitor, Microsoft Graph, Microsoft Sentinel, Virtual WAN, VPN Gateway, and other related services.

Microsoft’s preliminary investigation determined that the incident was caused by an error in its automated maintenance change process.

The company performed routine device maintenance that required isolating selected network paths. Before maintenance begins, Microsoft’s systems convert maintenance requests into automated instructions and perform validation checks to confirm that redundant network connectivity remains available.

A software defect in that conversion process caused the system to incorrectly identify additional devices as part of the maintenance activity. This resulted in more IP routes being removed than intended.

The removed routes disrupted connectivity between Microsoft’s data center infrastructure and its wide-area network (WAN), affecting traffic entering and leaving the West US region. The issue initially appeared as large-scale routing instability across Microsoft’s WAN. Engineers later determined that the unexpected route removals originated from a West US data center.

Microsoft advised customers running critical applications to consider multi-region deployment strategies to improve resilience during regional outages. Microsoft recommends using Azure Service Health alerts and reviewing application reliability through the Azure Well-Architected Framework.

The incident highlights the operational risks of large-scale cloud infrastructure changes and the importance of automated safeguards, redundant architectures, and rigorous validation processes in hyperscale cloud environments.