failure-injection-test-author
Orchestrates WireMock fault stubs (HTTP-level fault: 500s, malformed JSON, slow responses) with Toxiproxy (TCP-level: latency, packet loss, reset) into a single resilience test scenario - the test starts both, applies fault per scenario, runs the SUT against the impaired endpoints, verifies the SUT's resilience patterns. Use when one test must reproduce a combined network + HTTP failure - a cross-layer failure mode from an incident postmortem that neither pure HTTP fault stubs nor pure TCP chaos can cover alone, because most real failures span both layers.
Install with skills.sh (any agent)
npx skills add testland/qa --skill failure-injection-test-authorfailure-injection-test-author
Overview
Real production failures span layers:
A test using only WireMock (HTTP fault stubs) misses TCP-level chaos. A test using only Toxiproxy misses payload-level faults. Production failures combine both.
This skill builds a workflow that chains WireMock + Toxiproxy into one orchestrated test scenario - closer to production reality.
When to use
For pure HTTP fault stubs, see wiremock-stubs (in the qa-test-data plugin). For pure TCP chaos, see toxiproxy-chaos.
Step 1 - Topology
[ SUT (App) ] → [ Toxiproxy ] → [ WireMock ] → (returns canned response or 500)
↓ ↓ ↓
resilience network chaos HTTP fault stub
patterns (latency, etc) (500, malformed JSON, etc)The SUT connects to Toxiproxy; Toxiproxy forwards to WireMock; WireMock returns the configured response. The combined chain exercises both layers.
Step 2 - docker-compose setup
# docker-compose.test.yml
services:
wiremock:
image: wiremock/wiremock:3
ports: ["8081:8080"]
volumes:
- ./wiremock-mappings:/home/wiremock/mappings
toxiproxy:
image: ghcr.io/shopify/toxiproxy:latest
ports:
- "8474:8474"
- "8080:8080" # what the SUT connects to
app:
build: .
environment:
EXTERNAL_API_URL: http://toxiproxy:8080The SUT's EXTERNAL_API_URL points at Toxiproxy:8080; Toxiproxy forwards to wiremock:8080.
Step 3 - Configure the proxy
Once both containers are up:
# Tell Toxiproxy where to forward
curl -d '{"name":"external-api","listen":"0.0.0.0:8080","upstream":"wiremock:8080"}' \
http://toxiproxy:8474/proxiesStep 4 - Per-scenario test setup
// tests/resilience.spec.ts
import { Toxiproxy } from 'toxiproxy-node-client';
import axios from 'axios';
const toxiproxy = new Toxiproxy('http://toxiproxy:8474');
const wiremockBase = 'http://wiremock:8080';
beforeEach(async () => {
// Reset both
await axios.delete(`${wiremockBase}/__admin/mappings`);
const proxy = await toxiproxy.get('external-api');
for (const toxic of await proxy.toxics()) {
await proxy.removeToxic(toxic.name);
}
});
test('SUT retries on TCP reset followed by 500 then succeeds', async () => {
// 1. Stub WireMock: first call returns 500, second returns 200
await axios.post(`${wiremockBase}/__admin/mappings`, {
request: { method: 'GET', url: '/api/orders/1' },
response: { status: 500 },
priority: 1,
scenarioName: 'retry-test',
requiredScenarioState: 'Started',
newScenarioState: 'after-first',
});
await axios.post(`${wiremockBase}/__admin/mappings`, {
request: { method: 'GET', url: '/api/orders/1' },
response: { status: 200, jsonBody: { id: 1, status: 'fulfilled' } },
priority: 2,
scenarioName: 'retry-test',
requiredScenarioState: 'after-first',
});
// 2. Configure Toxiproxy: reset_peer toxic
const proxy = await toxiproxy.get('external-api');
await proxy.addToxic({
name: 'reset-on-first-byte',
type: 'reset_peer',
attributes: { timeout: 0 },
});
// 3. Trigger SUT
const result = await sut.fetchOrder(1);
// 4. Assert: SUT recovered after retry
expect(result).toEqual({ id: 1, status: 'fulfilled' });
// 5. Verify the WireMock log shows 2 attempts
const requests = await axios.get(`${wiremockBase}/__admin/requests`);
expect(requests.data.requests).toHaveLength(2);
});The test verifies: SUT made 2 calls (per WireMock log) and the second succeeded - the retry pattern works under TCP-reset + HTTP-500 combined fault.
Step 5 - Scenario catalog
Common scenarios:
| Scenario name | TCP toxic | HTTP fault | Verifies |
|---|---|---|---|
| Slow + 500 | latency 2000ms | 500 status | Retry honors timeout + retry-on-5xx |
| Reset + retry success | reset_peer (1 hit) | 200 (next call) | Retry handles connection reset |
| Slow body | bandwidth 1KB/s | 200 with large payload | Read timeout fires |
| Malformed JSON | (none) | 200 + invalid JSON | Parser handles gracefully |
| Cascade: timeout + 503 | timeout 5000ms | 503 | Circuit breaker opens after N timeouts |
| Network partition | timeout (forever) | (n/a) | Fallback to cached / null |
Step 6 - Verdict
Each scenario produces a per-resilience-pattern verdict:
## Failure injection results - `<sha>`
| Scenario | SUT behavior | Verdict |
|-----------------------|-----------------------------------------|---------|
| Slow + 500 | Retried 3 times; succeeded on 3rd | ✅ |
| Reset + retry success | Retried; succeeded | ✅ |
| Slow body | Read timeout at 5s; aborted | ✅ |
| Malformed JSON | ParseError thrown; defaulted to empty | ✅ |
| Cascade: timeout + 503 | Circuit breaker opened after 3 timeouts | ✅ |
| Network partition | Fell back to cached value | ⚠ partial - fallback returned stale > 1h |Step 7 - CI integration
- run: docker compose -f docker-compose.test.yml up --wait --wait-timeout 120
- run: npx jest tests/resilience.spec.ts
- run: docker compose -f docker-compose.test.yml down --volumesAnti-patterns
| Anti-pattern | Why it fails | Fix |
|---|---|---|
| Mocking the HTTP client instead of using WireMock | Mock can't simulate TCP-level faults. | Real Toxiproxy + WireMock chain (Step 1). |
| Forgetting to reset toxics + stubs between tests | Cross-test contamination. | beforeEach reset (Step 4). |
| Single-scenario tests (just 500, no TCP) | Real failures combine layers; single-layer tests miss them. | Author scenarios spanning both layers (Step 5). |
| Per-test docker-compose up / down | Slow; per-test setup overhead. | Per-suite docker-compose up (Step 7). |
| Not verifying the WireMock request log | Test passes even if SUT didn't actually retry (just got lucky). | Assert on __admin/requests count (Step 4 example). |
Limitations
References
Related skills
chaos-drill-protocol
Run protocol for a chaos experiment that has already been designed: the four pre-flight gates (non-production target, measured healthy baseline, live observability, a rollback that has actually been exercised), how to pick a conservative blast-radius bound, the sampling cadence and abort criteria fixed in writing before injection, and the recovery-validation step with its tolerance and timeout. Owns execution safety only, not experiment design: the steady-state hypothesis, the fault to inject, and the experiment file come from elsewhere. Use when an experiment definition exists and a fault is about to be injected into a running system, and the go/no-go gates, abort thresholds, and recovery check still need to be agreed and written down before the fault starts.
chaos-experiment-author
Build-an-X workflow for a chaos experiment per the Principles of Chaos Engineering - defines steady-state hypothesis, picks the variables (real-world events: network latency, node failure, region outage), sets the blast radius (which percentage / namespace / user cohort), automates execution, and emits the verdict (steady-state held / didn't hold). Use to scope a chaos experiment before running it via Litmus / Chaos Mesh / Gremlin / Toxiproxy.
chaos-mesh
Configures Chaos Mesh for Kubernetes-native chaos engineering - picks fault types (PodChaos, NetworkChaos, StressChaos, IOChaos, TimeChaos, DNSChaos, KernelChaos, HTTPChaos), targets via label selectors, controls blast radius via namespace whitelists + selector filters, schedules via CronJobs, observes via dashboard. Distinct from Litmus by architecture (Chaos Mesh has its own dashboard + workflow orchestration; Litmus uses ChaosCenter UI). Use when the target system runs on Kubernetes and fault experiments should be declared as CRDs in the cluster alongside the workloads they target.
chaos-results-reporter
Aggregates chaos drill verdicts over time into a resilience trend report - per-experiment hypothesis-held / blast-radius / time-to-detect / time-to-recover, degradation trends across runs, action items, and a stakeholder summary. Use when a team has completed one or more chaos drills and needs a structured trend report showing whether resilience is improving, degrading, or stable across iterations.
gremlin-chaos
Configures Gremlin (commercial) for cross-platform chaos engineering (fault injection, resilience testing) - installs the Gremlin agent on Linux / Windows / Kubernetes, picks attack types (resource, network, state, request), chains attacks into Scenarios (chaos experiments), integrates with the Reliability Score for forward-looking metrics. Use when the platform spans multiple environments (bare metal + cloud + serverless) and the team needs a commercial-supported solution per Gremlin's multi-platform support.
litmus-chaos
Configures LitmusChaos for Kubernetes-native chaos engineering - installs via Helm, picks ChaosExperiments from the ChaosHub (`pod-delete`, `network-latency`, `node-cpu-hog`, etc.), authors a ChaosEngine CR scoping the experiment + steady-state probes, runs as part of the cluster, exports Prometheus metrics for the verdict. Use when the platform is Kubernetes (CNCF-hosted; cloud-native). Prefer over chaos-mesh when the team wants a ChaosCenter web UI for workflow scheduling and ChaosHub catalog browsing; use chaos-mesh for fine-grained network-fault policies via its own CRD family.
steady-state-hypothesis-validator
Validates a chaos experiment's steady-state hypothesis before execution: checks that each probe metric is measurable and observable, that a recent baseline exists, that tolerances are numerically meaningful and SLI-backed, that the measurement window is defined, and that the chosen metrics would actually move under the target failure mode. Use when a chaos experiment has been authored (via chaos-experiment-author) and the team needs a pre-flight verdict before running the drill in any environment.
toxiproxy-chaos
Configures Toxiproxy for TCP-level fault injection - runs as a sidecar / proxy between client and upstream, applies toxics (latency, bandwidth, slow_close, timeout, slicer, limit_data, reset_peer) via control API. Focused on the proxy itself rather than an API-level chaos runner, including non-test usage (chaos in dev environments, integration tests, pre-prod simulation). Use when the team needs TCP-precise fault injection in development / integration environments without K8s or commercial tooling.