System status: mission critical

GameDays, without the boring parts.

Oh My! GameDay as a Service

You could build your own chaos engineering platform. We already did Chaos Monkey, Toxiproxy, k6, and more behind one dashboard. We deploy it faster than you can scaffold it, so your engineers spend their time on work that actually matters.

Every run targets infrastructure you've verified you own. Always.

Why we exist

We learned this at Amazon.

Chaos engineering is necessary, unglamorous infrastructure work. Every serious team needs to know their systems survive failure and almost none of them should spend a quarter building a fault-injection platform to find out. That's the trap: a year of engineering time spent on plumbing that never shows up in your product.

We genuinely love GameDays. There's a specific joy in watching a system survive something you broke on purpose, on your own schedule, instead of at 2am on a Saturday. So we built the thing we wished existed proven tools behind one dashboard and a single deploy command. We've already paid the setup cost, so your engineers don't have to.

“Chaos engineering infrastructure is undifferentiated heavy lifting. Your product isn't.”

The high cost of DIY

Your best engineers shouldn't spend months writing fault-injection boilerplate. Every hour on chaos plumbing is an hour not spent on what your company exists to build.

Already built, already proven

The same battle-tested tools the largest cloud teams rely on, wrapped in one dashboard and one deploy command. Amazon-grade GameDays in minutes, not quarters.

How it works

Four steps from zero to a real failure test.

01

Deploy the agent

One Docker command, into an environment you own.

02

Pick a scenario

From a catalog built on real, proven OSS tools.

03

Verify ownership

We check before anything runs. No exceptions.

04

Run it live

Watch it, hit the kill switch anytime, get a report after.

$ docker run --rm gdaas/agent init --env staging-us-east

Tool catalog

Real tools. Not gimmicks.

Every scenario is built on a named, proven open-source project or documented cloud API nothing invented, nothing untested at scale. You're not trusting a black box; you're getting the industry's own tools without the integration work.

State

Instance & Pod Killer

Kill instances and containers on a schedule.

Built on Chaos Monkey / Pumba
Network

Latency & Jitter Injector

Add delay and variance to network calls.

Built on Toxiproxy
Resource

CPU & Memory Stress

Pin resource usage to a target percentage.

Built on stress-ng
Load Generation

Load Swarm

Generate real concurrent traffic against a verified target.

Built on k6
Network

Network Partition Simulator

Split nodes to test split-brain handling.

Built on Blockade
Cloud API direct

Managed DB Failover

Trigger a controlled failover via your cloud provider's own API.

Built on AWS FIS / Azure Chaos Studio patterns
Dependency

Dependency Blackhole

Drop traffic to a downstream service and watch the fallback path.

Built on Toxiproxy
Resource

Disk & IO Saturation

Fill volumes and throttle IO to rehearse storage pressure.

Built on stress-ng
State

Clock Skew

Drift container clocks to surface time-sensitive assumptions.

Built on Pumba

Safety

We are not a booter.

Every target must be verified as yours before anything can run. Faults execute from inside through an agent you deploy into your own environment, or a scoped cloud role you grant us never from the outside in. Every run requires a metric-based stop condition that aborts it automatically when your system crosses a line you set, plus a manual kill switch that is always one click away. Every run produces an audit log you can hand to a reviewer.

Ownership verified

Kill switch always on

Blast-radius capped

Full audit log

Who it's for

Platform & SRE teams

Rehearse failure before it happens for real. Prove the failover path, the retry budget, and the on-call runbook work on a Tuesday afternoon, with the whole team watching, instead of discovering the gaps during an incident.

Compliance-driven teams

SOC 2, ISO 27001, DORA resilience testing is increasingly something you have to evidence, not just claim. Every GameDay ends with a timestamped report and audit log, so you walk away with an artifact, not just a memory.

The math

What is a year of plumbing worth?

Nobody knows exactly what their own resilience tooling costs which is exactly the problem. Set the sliders to your team and find out what the plumbing is quietly eating.

10 people
6 h/mo

Not on the tooling yet? Set it to zero then add the one-time build you would have to write first: a minimal viable fault-injection platform takes roughly 2,000 engineer-hours before anyone runs a single test.

Hours your team gets back, per year

720hrs

Ongoing maintenance you skip
720 h/yr
That is, in real terms
≈ 4.5 engineer-months
Plus the DIY build you never write
+ 2,000 h once
First-year total
≈ 17 engineer-months

At the default settings, that's more than a full engineer spending an entire year building and babysitting plumbing instead of your product.

Assumes 160 h per engineer-month. Move the sliders the conclusion doesn't change.

Pricing

Pay for what you break, not for a platform team.

Low-cost, usage-based access. No upfront infrastructure investment, no seat minimums, no year of platform engineering before your first run. Pricing is still being finalized get on the list and we'll bring you in early.

Join teams already running their first GameDay.