Testland
Browse all skills & agents

chaos-mesh

Configures Chaos Mesh for Kubernetes-native chaos engineering - picks fault types (PodChaos, NetworkChaos, StressChaos, IOChaos, TimeChaos, DNSChaos, KernelChaos, HTTPChaos), targets via label selectors, controls blast radius via namespace whitelists + selector filters, schedules via CronJobs, observes via dashboard. Distinct from Litmus by architecture (Chaos Mesh has its own dashboard + workflow orchestration; Litmus uses ChaosCenter UI). Use when the target system runs on Kubernetes and fault experiments should be declared as CRDs in the cluster alongside the workloads they target.

Install with skills.sh (any agent)

npx skills add testland/qa --skill chaos-mesh
View source

chaos-mesh

Overview

Per chaos-mesh-home (opens in new window):

"Chaos Mesh is a platform that 'brings various types of fault simulation to Kubernetes and has an enormous capability to orchestrate fault scenarios.'"

Per chaos-mesh-home (opens in new window), Chaos Mesh leverages "Kubernetes CustomResourceDefinitions (CRDs) for seamless integration with the Kubernetes ecosystem."

When to use

  • The platform is Kubernetes.
  • The team wants CRD-native chaos with a built-in dashboard.
  • Workflow orchestration matters (sequence + parallel experiments).
  • Physical machine support needed (Chaosd extension).

If LitmusChaos is already deployed, evaluate stack-fit before adding Chaos Mesh - both serve similar use cases with different ergonomics.

Step 1 - Install

curl -sSL https://mirrors.chaos-mesh.org/v2.6.3/install.sh | bash
# Or via Helm:
helm repo add chaos-mesh https://charts.chaos-mesh.org
helm install chaos-mesh chaos-mesh/chaos-mesh -n chaos-mesh --create-namespace

Per chaos-mesh-home (opens in new window), "no special dependencies required - Chaos Mesh deploys directly on Kubernetes clusters, including minikube and kind."

Step 2 - Fault types

Per chaos-mesh-home (opens in new window):

CRDEffect
PodChaosPod kill, container kill, pod failure
NetworkChaosLatency, packet loss, partition, bandwidth, corruption
StressChaosCPU stress, memory stress
IOChaosDisk read/write delay, errors
TimeChaosClock skew
DNSChaosDNS lookup failures
KernelChaosKernel-level fault injection
HTTPChaosHTTP request fault injection
JVMChaosJVM-level (exception, GC pause, method delay)

Plus Schedule for cron-style + Workflow for orchestration.

Step 3 - Author a NetworkChaos

apiVersion: chaos-mesh.org/v1alpha1
kind: NetworkChaos
metadata:
  name: checkout-network-latency
  namespace: app
spec:
  action: delay
  mode: one                    # or 'all', 'fixed', 'fixed-percent', 'random-max-percent'
  selector:
    namespaces:
      - app
    labelSelectors:
      app: checkout
  delay:
    latency: '500ms'
    correlation: '50'
    jitter: '50ms'
  duration: '5m'

Per chaos-mesh-home (opens in new window), Chaos Mesh provides "selector-based filtering using labels, annotations, and namespace whitelists to control 'blast radius' and target specific resources."

Step 4 - Author a PodChaos

apiVersion: chaos-mesh.org/v1alpha1
kind: PodChaos
metadata:
  name: checkout-pod-kill
  namespace: app
spec:
  action: pod-kill
  mode: fixed-percent
  value: '50'
  selector:
    namespaces:
      - app
    labelSelectors:
      app: checkout
  duration: '60s'

mode: fixed-percent + value: '50' kills 50% of matching pods.

Step 5 - Workflow orchestration

Per chaos-mesh-home (opens in new window): "Workflow Orchestration: Users can combine serial and parallel experiments to simulate complex, realistic failure scenarios matching actual system architecture."

apiVersion: chaos-mesh.org/v1alpha1
kind: Workflow
metadata:
  name: checkout-resilience-test
  namespace: app
spec:
  entry: combined-chaos
  templates:
    - name: combined-chaos
      templateType: Serial
      deadline: 30m
      children:
        - network-latency-step
        - then-pod-kill-step
        - then-stress-step
    - name: network-latency-step
      templateType: NetworkChaos
      networkChaos:
        action: delay
        mode: all
        selector: { ... }
        delay: { latency: 200ms }
        duration: 5m
    - name: then-pod-kill-step
      templateType: PodChaos
      podChaos:
        action: pod-kill
        mode: one
        selector: { ... }
    - name: then-stress-step
      templateType: StressChaos
      stressChaos:
        mode: all
        selector: { ... }
        stressors:
          cpu: { workers: 4, load: 80 }
        duration: 3m

Step 6 - Dashboard

Per chaos-mesh-home (opens in new window), Chaos Mesh ships a dashboard with RBAC.

kubectl port-forward -n chaos-mesh svc/chaos-dashboard 2333:2333
# Open http://localhost:2333

The dashboard provides authoring (visual experiment construction), running, observability, and replay.

Step 7 - Run + verdict

kubectl apply -f checkout-network-latency.yaml

# Watch state
kubectl get networkchaos checkout-network-latency -w

# Status / events
kubectl describe networkchaos checkout-network-latency

The chaos resource lifecycle is Created → Running → Stopped. Pair with external monitoring (Datadog / Prometheus / Grafana) to verify the steady-state hypothesis held.

Step 8 - Physical machine support

Per chaos-mesh-home (opens in new window): "Physical Machine Support: Chaosd (experimental) extends chaos testing to non-Kubernetes environments through PhysicalMachineChaos resources."

For VM / bare-metal targets:

apiVersion: chaos-mesh.org/v1alpha1
kind: PhysicalMachineChaos
metadata:
  name: vm-cpu-stress
spec:
  action: stress-cpu
  address:
    - 'http://10.0.0.5:31767'
  duration: 5m
  stress-cpu:
    load: 80
    workers: 4

The Chaosd agent runs on the target VM; the K8s CRD remotely triggers it.

Step 9 - CI integration

- name: Trigger chaos experiment
  run: |
    kubectl apply -f experiments/checkout-network-latency.yaml
    sleep 320  # 5min duration + buffer
    kubectl delete -f experiments/checkout-network-latency.yaml
- name: Check steady-state from Datadog
  run: ./scripts/datadog-verdict.sh

Anti-patterns

Anti-patternWhy it failsFix
mode: all without scopeAll matching pods affected; blast radius too wide.Start with mode: one or fixed-percent: 25.
No durationChaos persists until manual cleanup; risky.Always set duration (Step 3 example).
Targeting chaos-mesh namespaceCrashes the chaos infrastructure itself.Whitelist app namespace; deny-list chaos-mesh.
Disable RBAC on dashboardAnyone with cluster access can trigger chaos.Per chaos-mesh-home (opens in new window): RBAC is on by default - keep it on.
Skipping observability integrationChaos runs but verdict invisible.Wire dashboard + external monitoring.

Limitations

  • Kubernetes only (mostly). Chaosd is experimental; non-K8s is second-class.
  • Per-tool incompatibility. Chaos Mesh CRDs aren't Litmus ChaosEngines.
  • JVM / language-specific chaos. Available but requires agent installation in the target.
  • Resource overhead. Chaos controller + dashboard pods cost cluster resources.

References

  • cm (opens in new window) - Chaos Mesh overview: K8s-native, fault types, selector-based blast-radius control, workflow orchestration, dashboard with RBAC, Chaosd for physical machines.
  • litmus-chaos - sibling K8s alternative.
  • gremlin-chaos - multi-platform commercial alternative.
  • chaos-experiment-author - methodology this tool implements.

Related skills

chaos-drill-protocol

Run protocol for a chaos experiment that has already been designed: the four pre-flight gates (non-production target, measured healthy baseline, live observability, a rollback that has actually been exercised), how to pick a conservative blast-radius bound, the sampling cadence and abort criteria fixed in writing before injection, and the recovery-validation step with its tolerance and timeout. Owns execution safety only, not experiment design: the steady-state hypothesis, the fault to inject, and the experiment file come from elsewhere. Use when an experiment definition exists and a fault is about to be injected into a running system, and the go/no-go gates, abort thresholds, and recovery check still need to be agreed and written down before the fault starts.

chaos-experiment-author

Build-an-X workflow for a chaos experiment per the Principles of Chaos Engineering - defines steady-state hypothesis, picks the variables (real-world events: network latency, node failure, region outage), sets the blast radius (which percentage / namespace / user cohort), automates execution, and emits the verdict (steady-state held / didn't hold). Use to scope a chaos experiment before running it via Litmus / Chaos Mesh / Gremlin / Toxiproxy.

chaos-results-reporter

Aggregates chaos drill verdicts over time into a resilience trend report - per-experiment hypothesis-held / blast-radius / time-to-detect / time-to-recover, degradation trends across runs, action items, and a stakeholder summary. Use when a team has completed one or more chaos drills and needs a structured trend report showing whether resilience is improving, degrading, or stable across iterations.

failure-injection-test-author

Orchestrates WireMock fault stubs (HTTP-level fault: 500s, malformed JSON, slow responses) with Toxiproxy (TCP-level: latency, packet loss, reset) into a single resilience test scenario - the test starts both, applies fault per scenario, runs the SUT against the impaired endpoints, verifies the SUT's resilience patterns. Use when one test must reproduce a combined network + HTTP failure - a cross-layer failure mode from an incident postmortem that neither pure HTTP fault stubs nor pure TCP chaos can cover alone, because most real failures span both layers.

gremlin-chaos

Configures Gremlin (commercial) for cross-platform chaos engineering (fault injection, resilience testing) - installs the Gremlin agent on Linux / Windows / Kubernetes, picks attack types (resource, network, state, request), chains attacks into Scenarios (chaos experiments), integrates with the Reliability Score for forward-looking metrics. Use when the platform spans multiple environments (bare metal + cloud + serverless) and the team needs a commercial-supported solution per Gremlin's multi-platform support.

litmus-chaos

Configures LitmusChaos for Kubernetes-native chaos engineering - installs via Helm, picks ChaosExperiments from the ChaosHub (`pod-delete`, `network-latency`, `node-cpu-hog`, etc.), authors a ChaosEngine CR scoping the experiment + steady-state probes, runs as part of the cluster, exports Prometheus metrics for the verdict. Use when the platform is Kubernetes (CNCF-hosted; cloud-native). Prefer over chaos-mesh when the team wants a ChaosCenter web UI for workflow scheduling and ChaosHub catalog browsing; use chaos-mesh for fine-grained network-fault policies via its own CRD family.

steady-state-hypothesis-validator

Validates a chaos experiment's steady-state hypothesis before execution: checks that each probe metric is measurable and observable, that a recent baseline exists, that tolerances are numerically meaningful and SLI-backed, that the measurement window is defined, and that the chosen metrics would actually move under the target failure mode. Use when a chaos experiment has been authored (via chaos-experiment-author) and the team needs a pre-flight verdict before running the drill in any environment.

toxiproxy-chaos

Configures Toxiproxy for TCP-level fault injection - runs as a sidecar / proxy between client and upstream, applies toxics (latency, bandwidth, slow_close, timeout, slicer, limit_data, reset_peer) via control API. Focused on the proxy itself rather than an API-level chaos runner, including non-test usage (chaos in dev environments, integration tests, pre-prod simulation). Use when the team needs TCP-precise fault injection in development / integration environments without K8s or commercial tooling.