Microsoft Azure West US Region Goes Dark for Five Hours After Routine Fiber Check Goes Wrong

Microsoft Azure West US Region Goes Dark for Five Hours After Routine Fiber Check Goes Wrong

Microsoft's Azure West US region took a five-hour hit on July 23 after a routine fiber maintenance job went sideways. The outage, which began at 14:44 UTC and stretched until 19:41 UTC, knocked out 27 Azure services and left countless customers across California and the broader West Coast scrambling.

Microsoft traced the problem to a bug in the request-conversion system used during planned fiber maintenance. The system is supposed to isolate specific network paths and verify that at least one of two redundant links stays live during the work. Instead, the bug flagged extra devices as part of the maintenance event, causing the network to yank IP routes from more hardware than anyone intended.

The routes were cut between Microsoft's data centers and its wide-area network, essentially blocking traffic entering or leaving the West US region. That kicked off a cascade of failures across Azure's service portfolio. Application Gateway, Azure Firewall, Azure Kubernetes Service, Azure VMware Solution, Virtual WAN, VPN Gateway, and Microsoft Sentinel all flagged degradation within minutes.

First responders saw what looked like large-scale route churn inside Microsoft's WAN — a classic sign of a routing table spinning out of control. Teams eventually narrowed the source to a single data center in the West US region. Restoring the missing routes took coordinated effort across multiple networking teams.

The Services That Took the Hit

The list of impacted services reads like an Azure product catalog. Application Gateway lost the ability to route traffic into the region, taking down any web app that relied on it. Azure Firewall went dark, leaving east-west traffic inspection in limbo. Azure Kubernetes Service clusters lost connectivity to their control planes. Virtual WAN and VPN Gateway both reported connectivity failures, which meant site-to-site VPNs and branch-office links to Azure dropped.

Microsoft Sentinel, the company's security information and event management platform, also suffered data ingestion gaps during the window. For organizations that rely on Sentinel for real-time threat detection, that blind spot could have let malicious activity slip through unnoticed.

Several major customers reported cascading failures in their own applications. One large enterprise told The Register that its multi-region deployment architecture kicked in after 90 seconds of health-check failures, but the fail-over to East US wasn't entirely smooth — cached session data was lost, forcing some users to re-authenticate.

Lessons for Cloud Customers

Microsoft's post-incident analysis pointed at the automated maintenance request change process. The company said it would conduct a comprehensive review focusing on safety checks in the request-conversion pipeline. A fix is in the works, though Microsoft hasn't shared a timeline.

In its official post-mortem, Microsoft recommended that customers with mission-critical workloads adopt a multi-region deployment strategy. That's standard advice after any regional cloud outage, but it's worth repeating: if your entire infrastructure sits in one Azure region, a five-hour maintenance screw-up can stop your business cold.

Server racks at the Wikimedia Foundation data center

The company also pushed organizations to configure Azure Service Health alerts so the right people get notified the moment something goes wrong. A five-hour outage is bad enough; a five-hour outage that nobody notices until the help desk phones start ringing is worse.

There's also a broader question here about the fragility of cloud infrastructure. A single bug in a maintenance automation script brought down 27 services for nearly five hours. For all the talk of cloud reliability and five-nines uptime, this kind of incident shows that the gap between theory and practice is still wide. Microsoft runs a sprawling network of data centers linked by fiber, and when that fiber work goes wrong, the blast radius is enormous.

Google Faces Its Own Data Center Heat

Microsoft isn't the only cloud giant dealing with data center headaches this week. Google held a community meeting at the Broxbourne Council offices Wednesday evening for residents living near its newly built Waltham Cross data center north of London — and got an earful.

CERN data centre showing server infrastructure

Residents told Google's site operations team the facility, which sits on a 33-acre site, was generating a loud humming noise, particularly in the evening. Others complained about light pollution from the facility shining into their bedrooms all night. One resident claimed the data center was already dragging down house prices and asked if Google would buy his home.

Google's server operations site manager, Dan Dale, said the company was working on lighting projects to ensure lights "won't be on all the time" and would only activate when site access is needed. On the cooling front, site manager Tihomir Lazic explained that Waltham Cross uses a closed-loop chiller system that doesn't consume much water — roughly 60 cubic meters per month for the office space, equivalent to about five households.

The meeting highlighted a growing tension between cloud providers and the communities where they build. Data centers need to be close to population centers for low-latency connections, but nobody wants a humming, brightly lit server barn in their backyard. As hyperscale build-out accelerates — data centers are projected to consume nearly a fifth of U.S. electricity by 2035 — these community clashes will only become more common.

The Bigger Picture

For the cloud and edge computing industry, this week's events underscore two realities. First, even the most sophisticated cloud platforms can break in unexpected ways. Microsoft's automation bug is a reminder that the software that manages cloud infrastructure needs the same rigorous testing as the software that runs on it.

Second, the physical side of cloud computing — the data centers themselves — is becoming a community-relations challenge that no amount of engineering can solve alone. Google's Waltham Cross meeting shows that clear communication, noise mitigation, and good-faith engagement with neighbors are just as important as server uptime.

Cloud providers are spending billions on new capacity. Azure's revenue alone hit record levels in recent quarters as AI workloads drive demand through the roof. But the industry's growth is creating friction points — reliability gaps in automation tooling and community pushback on data center construction — that won't fix themselves.

The Azure West US outage affected real businesses. One Bay Area SaaS company told SDxCentral that the five-hour downtime cost it an estimated $120,000 in lost transactions and support overtime. For a company running on a single-region deployment, there's no substitute for the multi-region architecture Microsoft now recommends. But even that isn't cheap — replicating infrastructure across regions can double cloud bills.

Still, the takeaway isn't that cloud computing is broken. It's that the industry is still maturing. Five-hour outages, community NIMBY pushback, and automation bugs are growing pains. The hyperscalers have the engineering talent and the balance sheets to fix these problems. Whether they do it fast enough — before customers get burned or communities push back harder — is the open question. Check our Cloud & Edge Computing section for ongoing coverage.

In the meantime, if you're running critical workloads on Azure West US, now might be a good time to check those Azure Service Health alerts. And if you live near Waltham Cross, maybe invest in some blackout curtains.

← Back to Home