Microsoft Azure networking disruption hits UK regions after infrastructure changes

ExpressRoute, VPN Gateway, Azure Firewall and VMware services were among those affected as Microsoft investigated connectivity problems across multiple Azure regions.

By The Register

Microsoft has restored Azure services following a networking incident that disrupted connectivity across multiple cloud regions, including UK South and UK West.

The problem began at 20:30 UTC on 30 September and affected customers using several Azure gateway and networking services.

Microsoft said some users experienced degraded or interrupted connectivity, while gateways could fail to load in the Azure Portal and network management operations were subject to failures or delays.

Services affected included Azure ExpressRoute Gateway, Azure Firewall, Application Gateway and Web Application Firewall, VPN Gateway and Azure VMware Solution.

The disruption initially affected 18 Azure regions, according to updates issued during the incident. These included UK South and UK West, alongside regions in Europe, the US, Asia, Africa and the Middle East.

ExpressRoute is used to create private connections between organisations’ infrastructure and Microsoft’s Azure cloud, while VPN Gateway provides encrypted connections between Azure networks and other locations.

The services are widely used by organisations operating hybrid cloud infrastructure, meaning connectivity problems can affect links between on-premises systems and workloads hosted within Azure.

During the incident, Microsoft said some VPN gateways retained connectivity but were operating with reduced redundancy rather than suffering a complete outage.

Some network management components also failed to recover automatically, resulting in management operation failures in affected regions. Engineers began restoring those components using alternative healthy instances where available.

Microsoft initially identified a correlation between the incident and infrastructure operating-system servicing activity and paused that maintenance while engineers investigated.

A later update provided further detail on the sequence of events.

The company said a recent change to a regional gateway management service caused a higher-than-expected workload when separate operating-system servicing maintenance progressed through multiple regions.

The gateway management service would normally scale automatically to handle additional demand, but Microsoft said increased demand on dependent services prevented the regional systems from scaling as expected.

Engineers reverted the contributing gateway management change, reducing the load and allowing affected services to recover.

Microsoft said customer impact lasted until 02:15 UTC on 1 October.

The company is continuing to investigate the scaling behaviour and what additional safeguards may be required to prevent a similar incident.

Microsoft’s published timeline shows engineers first began investigating ExpressRoute connectivity problems in UK South at 21:29 UTC.

By 22:27 UTC, the investigation had established that multiple Azure regions were affected. At 23:05 UTC, engineers identified the link with operating-system servicing activity and paused further servicing while recovery work continued.

The incident was subsequently marked as mitigated after affected services returned to expected operation.

Open article on Cheshire Today