Scenarios

Pre-built disasters, ready to schedule.

Each scenario is a named failure story with a hypothesis, a script, an abort plan and a report. You pick the time Tuesday at 2 PM, not Saturday at 3 AM.

⚠ Scheduled anarchy verified infrastructure only

Five-alarm45 minOne AZ at a time

Region goes dark

One availability zone drops off the map and your retry logic makes it worse.

How it unfolds

  1. t+0Every pod in one availability zone goes dark at once. Your dashboards are still green they usually are.
  2. t+2Error rate climbs to 4%. Your retry logic, written for “a blip”, starts amplifying: every failed call is now three calls.
  3. t+6Latency budget blown. Connection pools saturate as the healthy zones absorb traffic they were never sized for.
  4. t+9The report writes itself: no backoff between retries, no cross-zone circuit breaker. The failure wasn't losing the zone it was the stampede your system created responding to it.
  5. t+12Kill switch. Steady state restored. Findings filed before lunch.

You'll find out

Whether your retry logic heals failure or causes it and whether the zones you kept were ever sized for the whole load.

Runs on: Network partition + cloud API

Spicy30 minDatabase tier only

The database failover you never tested

Managed Postgres failover sounds like someone else's problem. It isn't.

How it unfolds

  1. t+0We trigger failover on your replica via your cloud provider's own API, scoped to the database tier. The primary steps down.
  2. t+1Some requests fail fast. Most just… hang. Your connect timeout is 30 seconds; your users give up in four.
  3. t+4The app recovers mostly. One service never reconnects. Its pool is full of dead sockets it still considers alive.
  4. t+7The report shows which services spent the next hour reading from a stale replica. Nobody noticed. That's the part that should worry you.

You'll find out

Whether “failover” means 30 seconds of pain or an hour of quiet corruption and which services never reconnect at all.

Runs on: Managed DB failover (your cloud's own API)

Warm-up20 minOne dependency

Slow dependency

Nothing is down. Nothing is broken. Everything is two seconds slower.

How it unfolds

  1. t+0A third-party API you call on every checkout starts responding two seconds slower. No errors. No alerts. Just slow.
  2. t+3Thread pools fill. Queue depths grow. Autoscaling adds pods against a bottleneck none of them can fix.
  3. t+6The circuit breaker finally opens six minutes after the first user noticed. Your timeout budget said 800 ms; reality said 6,000.

You'll find out

Where your timeouts actually live, and how far one slow vendor ripples through everything that touches it.

Runs on: Latency injection (Toxiproxy)

Seven more in the library.

Spicy30 min

Fibre cut

The backhoe special: the link between two sites or zones goes silent mid-morning. Find out if your traffic reroutes or just queues until users give up.

Runs on: Network partition (Pumba)One link
Warm-up20 min

Noisy neighbour

CPU and memory starvation on one service, and the neighbours notice or worse, don't.

Runs on: Resource stress (stress-ng)One service
Spicy25 min

DNS is always the problem

We block name resolution and find out which services cache forever, which cache for nothing, and which one still hardcodes an IP from 2021.

Runs on: Network fault (Pumba)Name resolution only
Spicy60 min

Black Friday

k6 ramps traffic to ten times normal against a verified environment. Marketing calls it a forecast. We call it a rehearsal.

Runs on: Load generation (k6)Front door only
Warm-up15 min

The certificate expired at 3 AM

A simulated expired TLS cert on one endpoint. If your monitoring catches it before your customers do, this one's boring. Usually it isn't.

Runs on: TLS intercept (Toxiproxy)One endpoint
Spicy40 min

“Wait, wasn't that dev?”

A facilitated drill: a config change lands somewhere it shouldn't. The team has to notice, trace and roll back before the timer runs out.

Runs on: Facilitated process drillPeople, not systems
Surprise???

Oh My! moment

The facilitator draws a surprise from a shuffled deck. You find out what your runbooks actually say when nobody has read them.

Runs on: Random drawDepends on the draw

The full menu

Every kind of failure we can simulate.

Scenarios are recipes. These are the ingredients the fault types underneath them all, ready to be combined into whatever keeps you up at night.

⚠ Fibre cuts & network partitions

Sever the link between services, zones or sites. Find out what your system does when half of it simply stops answering.

⚠ Cloud region & AZ outages

Take a whole availability zone or region dark and watch your failover, retries and capacity math get tested for real.

⚠ External provider failover

Your ISP, CDN or DNS provider fails over or just fails. Simulate the handover and the outage before they do it to you.

⚠ Third-party API failure

The payment gateway, the maps API, the auth provider you depend on: slow it down, break it, and see how far the damage travels.

⚠ Database failover & corruption

Trigger managed failover, kill connections, and learn which services reconnect cleanly and which read stale data for an hour.

⚠ Resource starvation

CPU, memory and disk pressure on real workloads the noisy-neighbour problems that never show up in staging.

⚠ Traffic spikes

Load generation at ten times normal, against an environment you own. Rehearse Black Friday before it rehearses you.

⚠ Certificate & config rot

Expired certs, bad deploys, config drift the boring failures that cause most real outages and get the least practice.

⚠ People drills

Facilitated exercises where the team has to notice, trace and roll back a change before the timer runs out.

Request a GameDay

Tell us what keeps your systems up at night.

Pick the failure types most relevant to what you run. We'll draft a GameDay around them scoped, verified, aborted on your terms and walk your team through it.

Failure types to focus on (pick any)

Runs only against infrastructure you've verified you own.

Every scenario ships with the same anatomy.

Nothing here is “just run the command and see what happens”. Each one is a structured experiment you can defend to your manager.

✓

A hypothesis worth testing

✓

Steady-state checks before anything breaks

✓

Timed fault steps, one at a time

✓

Abort conditions that fire automatically

✓

A facilitator guide for running it as a team

✓

A post-run report you can hand to a reviewer

⚠ Under construction

Your failure isn't in the library?

The scenario builder lets you compose your own GameDay from the same verified tools pick a fault, set the timing, cap the blast radius, define the abort conditions. Coming to paid plans; early-access teams get first say on what we build next.

Request your GameDay →

Every scenario runs against infrastructure you've verified you own.