"Global Standard" is a billing tier, not a failover plan
Azure OpenAI lost Sweden Central for six hours last week, then the gateways went in 18 regions. Microsoft's own docs say your global deployment dies with its primary region.
cat contents.txt
Last Tuesday, between 10:03 and 15:58 UTC, Azure OpenAI, Foundry Agent Service, Foundry Models and Cognitive Services all fell over in Sweden Central. Microsoft's incident write-up (tracking ID 7QL5-Z50) is short and worth reading twice. A backend service that looks up resource metadata started timing out against its database and cache. "An internal system process generated an increase in the overall amount of requests for the backend service which reached thresholds." The instances restarted, kept restarting, and your chat completions came back as intermittent 5XXs and latency spikes for the best part of six hours.
The line I keep rereading is from the recovery: availability came back "following the scale-out and the disabling of an automated health check." The health check was part of the load.
Then, the next evening, a different Microsoft team had its own week. From 20:30 UTC on September 30 to 02:15 UTC on October 1 (tracking ID 7Q30-010), ExpressRoute Gateway, VPN Gateway, Azure Firewall, Application Gateway and WAF degraded across multiple regions, because "a recent change to a regional gateway management service triggered a higher-than-expected load when an unrelated operating system servicing maintenance proceeded gradually through multiple regions." Translation: a config change and a patch rollout met in the middle, and the thing that was supposed to autoscale did not.
Two incidents, two root causes, one lesson for anyone shipping an LLM feature on Azure.
the word "global" did a lot of work in your head
I asked three teams this week what deployment type they were on. All three said Global Standard, and two of them said it with the confidence of people who believed that meant multi-region failover.
It does not. Microsoft's deployment types page says it in a note most people scroll past: "With Global Standard and Data Zone Standard deployment types, if the primary region experiences an interruption in service, all traffic initially routed to this region is affected." Global means Microsoft may process your prompt anywhere. It is a data-residency statement and a quota statement. It is not a promise that a dead region gets routed around for you.
So if your Foundry resource lived in Sweden Central on Tuesday, the "global" in the SKU name did nothing for you between 10:03 and 15:58. Same page, same note, and I had never read it either.
what the chain actually looks like
Draw the path from your user to the model. App, then Application Gateway or Firewall, then private link into the Foundry resource, then whatever region the resource was created in, then the model. Last week each of the two incidents took out a different hop. The model hop went Tuesday. The network hop went Wednesday. A feature that was "up" by your own dashboard could have been unreachable either day, for reasons that look identical from a client: timeouts and 5XX.
Which brings me to retries. The SDK default retry with backoff is exactly wrong against an incident like 7QL5-Z50, where the backend was dying under request volume. Every client retrying twice is more load on a service whose root cause was load. Retries belong at the gateway, where they can be counted and capped per backend, not in every pod.
the three things I actually changed
- Two resources, two regions, one gateway. Azure API
Management's
AI gateway
supports backend pools with priority, weights and a circuit
breaker whose trip duration comes from the backend's
Retry-Afterheader. Primary Foundry resource in one region, secondary in another, same model version pinned in both, the gateway fails over. This is boring, documented and has existed for over a year. I just had not done it for the "small" features. - Log the region as well as the model. My cost and latency logs recorded the deployment name. They did not record which resource and region served the call, so on Tuesday I could not tell a Sweden Central failure from a bad prompt until I read the status page. If you followed the cost observability piece, add one column. Region. Future you will thank present you at 10:14 UTC on a Tuesday.
- A canary that calls the model, not the health endpoint. The platform's own health check made its outage worse, so I am done trusting anyone's health endpoint, including mine. The canary sends a 20-token completion through the real gateway every minute per backend and trips the pool when p95 crosses a line. It costs about a dollar a day. Tuesday cost one team I know six hours of a support bot answering nothing.
the bill nobody mentions
Multi-region doubles the number of deployments you pay quota for, and if you are on provisioned throughput it doubles the reserved capacity you pay whether or not it serves a byte. That is the real reason the three teams I asked were single-region: the second resource looked like waste on the invoice. Six hours of a down feature during European business hours is also on an invoice. It just arrives as churn instead of a line item.
Read the Microsoft note again. "All traffic initially routed to this region is affected." They told you. It was in the docs the whole time, right under the part that says Global Standard is the one to start with.
tags: #azure #resilience #architecture