2 min read

Microsoft maintenance bug cuts off Azure West US

A Microsoft maintenance bug removed Azure routes in West US, disrupting 27 services for almost five hours on July 23.

Image: The Register

A Microsoft fiber-maintenance error cut off access to Azure resources in the West US region for almost five hours on July 23, disrupting 27 services and leaving customers unable to reach workloads hosted in the company’s Californian cloud facilities.

Microsoft began what it described as “routine device maintenance” at 14:44 UTC (07:44 Pacific Time). The work immediately triggered service degradation, prompting networking, services, and incident-response teams to investigate traffic anomalies, routing behavior, packet loss, and recent configuration changes.

How the maintenance caused the outage

Microsoft’s maintenance process is designed to isolate specific network paths while ensuring that at least one of two redundant paths remains healthy. The company also performs safety checks intended to confirm the work will not affect customers.

That process failed because of a bug in the system that converts maintenance requests into instructions for network devices. The bug incorrectly included additional devices in the maintenance event, causing IP routes to be removed from more devices than intended.

Recommended reading

Nvidia’s Rubin servers turn to hot-liquid cooling

“The routes were removed between our datacenter and wide-area network, impacting traffic entering or exiting the region.”

Microsoft

The incident initially appeared as large-scale route churn across Microsoft’s Wide-Area Network. Further investigation showed that the route removal originated in a West US datacenter.

Microsoft identified a connection to recent fiber-maintenance activity sometime between 16:00 UTC and 17:45 UTC. At 17:45 UTC, it began rolling back the changes. The company’s WAN had recovered by 18:26 UTC, while all affected services were fully restored by 19:41 UTC.

The outage adds Microsoft to a recent series of cloud incidents involving AWS, Google, and Azure, highlighting how a maintenance operation intended to reduce risk can itself sever connectivity to an entire region.

Marcus Vance

Enterprise Editor

Marcus follows the money. He covers enterprise software, cloud architecture, and the tectonic shifts in Big Tech strategy. He translates dense earnings calls and complex M&A activity into actionable insights about where the industry is actually heading. If a tech giant makes a silent pivot, Marcus is usually the first to notice.

via The Register

/ Keep reading