Region goes dark
One availability zone drops off the map and your retry logic makes it worse.
How it unfolds
- t+0Every pod in one availability zone goes dark at once. Your dashboards are still green they usually are.
- t+2Error rate climbs to 4%. Your retry logic, written for “a blip”, starts amplifying: every failed call is now three calls.
- t+6Latency budget blown. Connection pools saturate as the healthy zones absorb traffic they were never sized for.
- t+9The report writes itself: no backoff between retries, no cross-zone circuit breaker. The failure wasn't losing the zone it was the stampede your system created responding to it.
- t+12Kill switch. Steady state restored. Findings filed before lunch.
You'll find out
Whether your retry logic heals failure or causes it and whether the zones you kept were ever sized for the whole load.
Runs on: Network partition + cloud API