chaos-results-reporter
Aggregates chaos drill verdicts over time into a resilience trend report - per-experiment hypothesis-held / blast-radius / time-to-detect / time-to-recover, degradation trends across runs, action items, and a stakeholder summary. Use when a team has completed one or more chaos drills and needs a structured trend report showing whether resilience is improving, degrading, or stable across iterations.
Install with skills.sh (any agent)
npx skills add testland/qa --skill chaos-results-reporterchaos-results-reporter
Overview
Per principlesofchaos.org (opens in new window):
"Chaos Engineering is the discipline of experimenting on a system in order to build confidence in the system's capability to withstand turbulent conditions in production."
A single drill verdict tells you whether the system held on a specific day. A trend report answers the harder question: is the system getting more resilient over time, or are the same blast-radius categories failing on every run?
This skill walks the workflow for aggregating drill results - from a completed chaos drill, or from chaos-experiment-author Step 7 - into a structured trend report with per-experiment metrics, cross-run trend lines, action items, and a stakeholder summary.
Differentiation axis vs. chaos-experiment-author Step 7: that step produces a single-drill verdict. This skill aggregates multiple verdicts over time, computes trend direction, and produces a stakeholder-facing document. Differentiation axis vs. running a drill live: a real-time drill run produces reports as it executes. This skill runs post-hoc, after one or more drills have already completed and their reports exist on disk.
Hard-reject rule
If the input set contains zero completed drill reports (no hypothesis, verdict, timestamps, or observed metrics), halt immediately:
HALT: No drill results found. Provide at least one completed drill report
before running chaos-results-reporter.Do not fabricate metrics or assume a prior run exists.
How to use
Step 1 - Collect drill reports
Locate completed drill reports. Supported input forms:
For each report, extract the fields in the field-extraction map in references/drill-fields-and-signals.md. If any required field is missing from a report, flag that report as INCOMPLETE in the aggregate table and skip it from trend calculations. Do not guess or interpolate missing values.
Step 2 - Build the per-experiment summary table
Emit one row per drill run, sorted by date ascending:
| Date | Experiment | Verdict | Hypothesis held | Blast radius (peak err) | TTD | TTR |
|------------|-------------------------------|----------|-----------------|------------------------|--------|--------|
| 2026-01-10 | checkout-network-latency | PASSED | Yes | 1.2% (bound 5%) | n/a | 42 s |
| 2026-02-07 | checkout-network-latency | ABORTED | No | 6.1% (bound 5%) | 78 s | 3 m 2 s |
| 2026-03-01 | checkout-network-latency | PASSED | Yes | 0.9% (bound 5%) | n/a | 31 s |TTD (time-to-detect): how long from injection start until the blast-radius monitor triggered an abort or the team observed a signal. Per chaos-principles (opens in new window) principle "Minimize Blast Radius": reducing TTD is a leading indicator of maturing blast-radius containment.
TTR (time-to-recover): how long from experiment end until steady state returned. Maps to the ISTQB concept of recoverability under ISO/IEC 25010:2023 (opens in new window) Quality Characteristic: Reliability > Recoverability (the ability of software to recover data directly affected in the case of an interruption or failure and re-establish the desired state of the system).
Step 3 - Compute per-experiment trend
For each unique experiment_type with two or more runs, compute:
Hypothesis-held rate across runs:
held_rate = (count of PASSED runs) / (total runs for this experiment type)Flag the trend direction:
Blast-radius trend: compare peak observed error rate run over run. Flag WIDENING if the most recent peak is more than 20% above the earliest recorded peak for the same experiment type; NARROWING if it is more than 20% below; STABLE otherwise.
TTR trend: compare recovery time run over run. Flag IMPROVING if the median TTR in the second half of runs is shorter than in the first half; DEGRADING if longer; STABLE if within 20%.
Step 4 - Identify degradation signals
Scan all experiments for the degradation signals catalogued in references/drill-fields-and-signals.md. Emit each match as a named finding.
Per chaos-principles (opens in new window) principle "Automate Experiments to Run Continuously": a DEGRADING trend on an automated experiment is a signal that the system has drifted since the experiment was written. Do not treat a single passing run as permanent confidence.
Step 5 - Compile action items
For each HIGH or MEDIUM finding, emit a concrete action item with:
Step 6 - Emit the stakeholder summary
The stakeholder summary is a short (8-15 line) non-technical section for engineering leads and product owners. It must:
Example:
## Resilience trend summary - Q1 2026
**Period:** 2026-01-10 to 2026-03-01 | **Drills run:** 5 | **Passed:** 3 |
**Aborted:** 2 | **Failed:** 0
The checkout service maintained its target error rate in 3 of 5 runs. Two
network-latency drills in February were aborted when the error rate exceeded
the 5% budget, with the peak reaching 6.1%. Recovery time improved from
3 minutes in February to 31 seconds in March after the retry backoff was
tuned.
**Top action items:**
1. Confirm the February blast-radius breach root cause before the next
scheduled drill (checkout-network-latency).
2. Expand the pod-kill experiment blast radius from 1% to 5% of replicas
now that the network-latency experiment is stable.Output format (full report)
The full report consists of four sections in order:
Write the report to results/chaos/trend-report-<YYYY-MM-DD>.md (where the date is today's date) unless the user specifies a different output path.
Worked example
Three completed drill reports for checkout-network-latency sit in results/chaos/:
Steps 1-2: All three parse cleanly (no INCOMPLETE reports), yielding the per-experiment summary table sorted by date.
Step 3: Held rate is 2 of 3. Peak error moved from 1.2% (Jan) to 0.9% (Mar), more than 20% below the earliest peak, so the blast-radius trend is NARROWING. TTR fell from the 3 m 2 s February peak to 31 s in March, so the TTR trend is IMPROVING. Because the series is not yet 3 consecutive PASSED runs, the March pass is treated as regression toward the mean, not confirmed recovery (see Anti-patterns).
Step 4: No HIGH or MEDIUM signal fires for this experiment type: the abort was isolated to one run (no repeated blast-radius breach), TTR improved after the run (no "No TTR improvement after fix"), and the blast-radius trend is narrowing rather than widening.
Steps 5-6: With no HIGH or MEDIUM finding, Steps 4-5 add no action items for this experiment type. The stakeholder summary reads: "The checkout service maintained its target error rate in 2 of 3 network-latency runs; recovery time improved from 3 minutes in February to 31 seconds in March after the retry backoff was tuned."
The four-section report is written to results/chaos/trend-report-<YYYY-MM-DD>.md.
Anti-patterns
| Anti-pattern | Why it fails | Fix |
|---|---|---|
| Reporting on a single drill run | One data point can't show a trend | Collect at least 2 runs per experiment type before computing trend direction |
| Interpolating missing fields | Fabricated metrics corrupt the trend | Flag the report as INCOMPLETE and exclude it from calculations |
| Marking a DEGRADING trend as acceptable because the most recent run passed | One passing run after a degraded series is regression toward the mean, not confirmed recovery | Require 3 consecutive PASSED runs before reclassifying to STABLE or IMPROVING |
| Mixing experiment types in a single trend line | Network-latency and pod-kill have different blast-radius profiles | Keep trends per experiment type (Step 3) |
| Copying raw metric tables into the stakeholder summary | Non-technical readers lose the signal in the noise | Keep the summary prose-only; tables stay in the per-experiment section |
Limitations
References
Drill report field extraction and degradation signals
View source (opens in new window)Drill report field extraction and degradation signals
Reference tables for chaos-results-reporter. Step 1 uses the field-extraction map to parse each drill report into a normalized row; Step 4 uses the degradation-signals catalog to flag findings across runs.
Field extraction map
For each completed drill report, extract these fields from the listed source locations. If any required field is missing from a report, flag that report as INCOMPLETE in the aggregate table and skip it from trend calculations - do not guess or interpolate missing values.
| Field | Source location in drill report |
|---|---|
experiment_id | Header: Chaos drill report - <id> |
date | Start: timestamp, truncated to date |
experiment_type | Experiment: field |
hypothesis | Steady-state hypothesis from experiment YAML |
verdict | Verdict: field: PASSED, ABORTED, FAILED |
blast_radius_observed | Peak error rate, Peak affected replicas, Peak latency p99 |
blast_radius_bound | Blast-radius bound: field |
time_to_detect | Time from injection start to first abort criterion breach (or "n/a" if not aborted) |
time_to_recover | Recovery time: field |
abort_reason | Abort reason: field (empty if verdict is PASSED) |
Degradation signals catalog
Scan all experiments for these signals. Emit each match as a named finding; each HIGH or MEDIUM finding becomes an action item in Step 5.
| Signal | Condition | Severity |
|---|---|---|
| Repeated blast-radius breach | Same experiment type aborted for the same abort reason in 2+ consecutive runs | HIGH |
| No TTR improvement after fix | ABORTED run was followed by a code change (per git log or team note), but TTR did not improve in the next run | HIGH |
| Widening blast radius | Blast-radius trend is WIDENING for any experiment type | MEDIUM |
| Declining held rate | Held rate dropped more than 20 percentage points between first and second half of run history | MEDIUM |
| Single data point | Any experiment type has only one run | LOW (informational) |
Related skills
chaos-drill-protocol
Run protocol for a chaos experiment that has already been designed: the four pre-flight gates (non-production target, measured healthy baseline, live observability, a rollback that has actually been exercised), how to pick a conservative blast-radius bound, the sampling cadence and abort criteria fixed in writing before injection, and the recovery-validation step with its tolerance and timeout. Owns execution safety only, not experiment design: the steady-state hypothesis, the fault to inject, and the experiment file come from elsewhere. Use when an experiment definition exists and a fault is about to be injected into a running system, and the go/no-go gates, abort thresholds, and recovery check still need to be agreed and written down before the fault starts.
chaos-experiment-author
Build-an-X workflow for a chaos experiment per the Principles of Chaos Engineering - defines steady-state hypothesis, picks the variables (real-world events: network latency, node failure, region outage), sets the blast radius (which percentage / namespace / user cohort), automates execution, and emits the verdict (steady-state held / didn't hold). Use to scope a chaos experiment before running it via Litmus / Chaos Mesh / Gremlin / Toxiproxy.
chaos-mesh
Configures Chaos Mesh for Kubernetes-native chaos engineering - picks fault types (PodChaos, NetworkChaos, StressChaos, IOChaos, TimeChaos, DNSChaos, KernelChaos, HTTPChaos), targets via label selectors, controls blast radius via namespace whitelists + selector filters, schedules via CronJobs, observes via dashboard. Distinct from Litmus by architecture (Chaos Mesh has its own dashboard + workflow orchestration; Litmus uses ChaosCenter UI). Use when the target system runs on Kubernetes and fault experiments should be declared as CRDs in the cluster alongside the workloads they target.
failure-injection-test-author
Orchestrates WireMock fault stubs (HTTP-level fault: 500s, malformed JSON, slow responses) with Toxiproxy (TCP-level: latency, packet loss, reset) into a single resilience test scenario - the test starts both, applies fault per scenario, runs the SUT against the impaired endpoints, verifies the SUT's resilience patterns. Use when one test must reproduce a combined network + HTTP failure - a cross-layer failure mode from an incident postmortem that neither pure HTTP fault stubs nor pure TCP chaos can cover alone, because most real failures span both layers.
gremlin-chaos
Configures Gremlin (commercial) for cross-platform chaos engineering (fault injection, resilience testing) - installs the Gremlin agent on Linux / Windows / Kubernetes, picks attack types (resource, network, state, request), chains attacks into Scenarios (chaos experiments), integrates with the Reliability Score for forward-looking metrics. Use when the platform spans multiple environments (bare metal + cloud + serverless) and the team needs a commercial-supported solution per Gremlin's multi-platform support.
litmus-chaos
Configures LitmusChaos for Kubernetes-native chaos engineering - installs via Helm, picks ChaosExperiments from the ChaosHub (`pod-delete`, `network-latency`, `node-cpu-hog`, etc.), authors a ChaosEngine CR scoping the experiment + steady-state probes, runs as part of the cluster, exports Prometheus metrics for the verdict. Use when the platform is Kubernetes (CNCF-hosted; cloud-native). Prefer over chaos-mesh when the team wants a ChaosCenter web UI for workflow scheduling and ChaosHub catalog browsing; use chaos-mesh for fine-grained network-fault policies via its own CRD family.
steady-state-hypothesis-validator
Validates a chaos experiment's steady-state hypothesis before execution: checks that each probe metric is measurable and observable, that a recent baseline exists, that tolerances are numerically meaningful and SLI-backed, that the measurement window is defined, and that the chosen metrics would actually move under the target failure mode. Use when a chaos experiment has been authored (via chaos-experiment-author) and the team needs a pre-flight verdict before running the drill in any environment.
toxiproxy-chaos
Configures Toxiproxy for TCP-level fault injection - runs as a sidecar / proxy between client and upstream, applies toxics (latency, bandwidth, slow_close, timeout, slicer, limit_data, reset_peer) via control API. Focused on the proxy itself rather than an API-level chaos runner, including non-test usage (chaos in dev environments, integration tests, pre-prod simulation). Use when the team needs TCP-precise fault injection in development / integration environments without K8s or commercial tooling.