Chaos Engineering definition
Chaos engineering is the practice of deliberately injecting failures into a system, such as killing servers, adding network latency or breaking dependencies, to verify that it keeps working as expected. Run as controlled experiments with a hypothesis and a limited blast radius, it uncovers weaknesses before they cause real outages. Netflix popularized it with its Chaos Monkey tool.
Why break things on purpose?
Modern systems are built from many services, cloud components and third-party APIs, and their failure behavior is hard to predict from design documents. Timeouts may be missing, retries may amplify load, failover may depend on a credential nobody has used in a year. Chaos engineering tests these assumptions directly, in a controlled way and at a time the team chooses, instead of waiting for an incident at 3 a.m. to reveal them.
Netflix built Chaos Monkey to randomly terminate production instances, forcing engineers to design services that survive the loss of any single server. The discipline has since spread well beyond streaming to banks, retailers and cloud providers, often as part of site reliability engineering practice.
How to run a chaos experiment
A chaos experiment is a scientific test rather than random destruction. Each one follows the same basic steps, and it is stopped immediately if user impact goes beyond the limits agreed with stakeholders beforehand. Experiments are written down in advance so anyone can review them.
Good candidates for first experiments are failures that already happened in past incidents or near misses, because the team knows they are realistic and can confirm whether the fixes made afterward actually work. The steps are:
- Define steady state: measurable signals of normal behavior, such as orders per minute or error rate
- Form a hypothesis: for example, if one availability zone fails, checkout errors stay below 1 percent
- Limit the blast radius: start in staging or with a small share of production traffic
- Inject the failure: terminate instances, add latency, block a dependency or exhaust CPU
- Observe and compare against steady state using monitoring and traces
- Fix what you learn, then automate the experiment so it runs regularly
Example failure scenarios
Useful early experiments are simple. Terminate a random pod or instance and confirm traffic shifts without errors. Add 500 ms of latency to a payment provider call and check that timeouts, retries and user messaging behave sensibly. Make the cache unavailable and see whether the database survives the extra load. Fail over the primary database and measure how long writes are interrupted.
Each experiment usually finds something: an alert that never fired, a retry storm, a dashboard that hid the problem. Those findings feed directly into runbooks and observability improvements. Teams also run game days, scheduled sessions where people practice responding to injected failures together.
Tools and where to start
Tools include Chaos Monkey from Netflix, Gremlin, LitmusChaos and Chaos Mesh for Kubernetes, AWS Fault Injection Service and Azure Chaos Studio. Most can target specific instances, containers or network paths and stop automatically when a safety alarm triggers, which keeps experiments within the agreed blast radius.
Start only after basic monitoring, alerting and high availability are in place; chaos experiments on a system with no redundancy simply cause outages. Nexzem introduces chaos experiments gradually for clients running critical platforms, beginning in staging with clear hypotheses and rollback plans.