toxiproxy-chaos
Configures Toxiproxy for TCP-level fault injection - runs as a sidecar / proxy between client and upstream, applies toxics (latency, bandwidth, slow_close, timeout, slicer, limit_data, reset_peer) via control API. Focused on the proxy itself rather than an API-level chaos runner, including non-test usage (chaos in dev environments, integration tests, pre-prod simulation). Use when the team needs TCP-precise fault injection in development / integration environments without K8s or commercial tooling.
Install with skills.sh (any agent)
npx skills add testland/qa --skill toxiproxy-chaostoxiproxy-chaos
Overview
Toxiproxy is Shopify's "TCP proxy to simulate network and system conditions for chaos and resiliency testing" (toxiproxy-readme (opens in new window)). It sits between client and upstream; you configure toxics via its HTTP control API.
This skill is the infrastructure / dev-environment angle. The test-suite-driven angle is in api-chaos-runner (in the qa-api-testing plugin); both rely on the same Toxiproxy primitive.
When to use
For test-suite integration, see api-chaos-runner.
Step 1 - Install + run
# Pull the official image
docker pull ghcr.io/shopify/toxiproxy:latest
# Run as a daemon
docker run --rm -p 8474:8474 -p 5432:5432 ghcr.io/shopify/toxiproxy:latest
# Or natively (Linux):
brew install toxiproxy # macOS
# Then: toxiproxy-serverPort 8474 is the control API; other ports are listeners for proxied traffic.
Step 2 - Define a proxy
Via the control API:
curl -d '{"name":"orders-db","listen":"0.0.0.0:5432","upstream":"orders-db-real:5432"}' \
http://localhost:8474/proxiesOr via the CLI:
toxiproxy-cli create -l 0.0.0.0:5432 -u orders-db-real:5432 orders-db
toxiproxy-cli listThe application connects to localhost:5432 (Toxiproxy listener); Toxiproxy forwards to orders-db-real:5432.
Step 3 - Toxic catalog
Per toxiproxy-readme (opens in new window), the canonical toxics:
| Toxic | Effect |
|---|---|
latency | Add latency to all data passing through |
down | Force the proxy down (no connections accepted) |
bandwidth | Cap bandwidth in kbps |
slow_close | Delay TCP socket close |
timeout | Stop forwarding traffic after a delay; let connection time out |
slicer | Slice TCP data into smaller bits |
limit_data | Cap total bytes through the proxy |
reset_peer | Reset connection on the next byte |
Step 4 - Add toxics
# 500ms latency on every request through the orders-db proxy
toxiproxy-cli toxic add -t latency -a latency=500 orders-db
# Bandwidth cap at 50 KB/s
toxiproxy-cli toxic add -t bandwidth -a rate=50 orders-db
# Force the proxy down (kill the connection)
toxiproxy-cli toxic add -t timeout -a timeout=5000 orders-db
# Remove all toxics on this proxy
toxiproxy-cli toxic remove orders-db -n latencyToxics can apply on upstream (data going from client → server) or downstream (server → client) directions. Default: both.
Step 5 - Direction-specific toxics
toxiproxy-cli toxic add -t latency -a latency=500 -n upstream-latency --downstream=false orders-db
toxiproxy-cli toxic add -t latency -a latency=200 -n downstream-latency --upstream=false orders-dbUseful when the client / server have asymmetric tolerances.
Step 6 - Language SDKs
Per toxiproxy-readme (opens in new window), SDKs exist for Python, Node, Go, Ruby:
# Python
from toxiproxy import Toxiproxy
client = Toxiproxy()
proxy = client.create('orders-db', '0.0.0.0:5432', 'orders-db-real:5432')
proxy.add_toxic(name='latency', type='latency', attributes={'latency': 500})
# ... run app ...
proxy.destroy()// Node
const Toxiproxy = require('toxiproxy-node-client');
const client = new Toxiproxy('http://localhost:8474');
const proxy = await client.createProxy({ name: 'orders-db', listen: '0.0.0.0:5432', upstream: 'orders-db-real:5432' });
await proxy.addToxic({ type: 'latency', attributes: { latency: 500 } });The SDKs make integration into test fixtures (per playwright-fixture-builder) clean.
Step 7 - docker-compose integration
# docker-compose.test.yml
services:
toxiproxy:
image: ghcr.io/shopify/toxiproxy:latest
ports:
- 8474:8474 # control
- 5432:5432 # proxied DB
- 8080:8080 # proxied API
app:
environment:
DB_HOST: toxiproxy
DB_PORT: 5432
EXTERNAL_API_URL: http://toxiproxy:8080The app points at Toxiproxy; tests configure toxics via the control API.
Step 8 - Use cases
| Use case | How |
|---|---|
| Test resilience patterns | Inject latency / failure; verify retry / timeout / circuit-breaker |
| Reproduce a production incident | Replicate the network conditions; debug locally |
| Pre-prod simulation | Chaos in staging; verify the team's runbook |
| Dev-time exploration | "What if the DB is slow?" - engineer toggles a toxic |
| Integration test fixtures | Per-test toxic add / remove via SDK (per Step 6) |
Anti-patterns
| Anti-pattern | Why it fails | Fix |
|---|---|---|
| Toxiproxy in production | Adds proxy latency + failure surface to real traffic. | Test / staging only. |
| Forgetting to remove toxics after test | Subsequent tests inherit chaos; flaky. | Cleanup in afterEach / context manager. |
| One global Toxiproxy for parallel tests | Tests fight over shared toxic state. | One Toxiproxy per parallel worker. |
| Skipping direction (default = both) | Asymmetric scenarios miss. | Explicit --upstream=false / --downstream=false (Step 5). |
| Manual control-API curl in tests | Verbose; error-prone. | Use language SDK (Step 6). |
Limitations
References
Related skills
chaos-drill-protocol
Run protocol for a chaos experiment that has already been designed: the four pre-flight gates (non-production target, measured healthy baseline, live observability, a rollback that has actually been exercised), how to pick a conservative blast-radius bound, the sampling cadence and abort criteria fixed in writing before injection, and the recovery-validation step with its tolerance and timeout. Owns execution safety only, not experiment design: the steady-state hypothesis, the fault to inject, and the experiment file come from elsewhere. Use when an experiment definition exists and a fault is about to be injected into a running system, and the go/no-go gates, abort thresholds, and recovery check still need to be agreed and written down before the fault starts.
chaos-experiment-author
Build-an-X workflow for a chaos experiment per the Principles of Chaos Engineering - defines steady-state hypothesis, picks the variables (real-world events: network latency, node failure, region outage), sets the blast radius (which percentage / namespace / user cohort), automates execution, and emits the verdict (steady-state held / didn't hold). Use to scope a chaos experiment before running it via Litmus / Chaos Mesh / Gremlin / Toxiproxy.
chaos-mesh
Configures Chaos Mesh for Kubernetes-native chaos engineering - picks fault types (PodChaos, NetworkChaos, StressChaos, IOChaos, TimeChaos, DNSChaos, KernelChaos, HTTPChaos), targets via label selectors, controls blast radius via namespace whitelists + selector filters, schedules via CronJobs, observes via dashboard. Distinct from Litmus by architecture (Chaos Mesh has its own dashboard + workflow orchestration; Litmus uses ChaosCenter UI). Use when the target system runs on Kubernetes and fault experiments should be declared as CRDs in the cluster alongside the workloads they target.
chaos-results-reporter
Aggregates chaos drill verdicts over time into a resilience trend report - per-experiment hypothesis-held / blast-radius / time-to-detect / time-to-recover, degradation trends across runs, action items, and a stakeholder summary. Use when a team has completed one or more chaos drills and needs a structured trend report showing whether resilience is improving, degrading, or stable across iterations.
failure-injection-test-author
Orchestrates WireMock fault stubs (HTTP-level fault: 500s, malformed JSON, slow responses) with Toxiproxy (TCP-level: latency, packet loss, reset) into a single resilience test scenario - the test starts both, applies fault per scenario, runs the SUT against the impaired endpoints, verifies the SUT's resilience patterns. Use when one test must reproduce a combined network + HTTP failure - a cross-layer failure mode from an incident postmortem that neither pure HTTP fault stubs nor pure TCP chaos can cover alone, because most real failures span both layers.
gremlin-chaos
Configures Gremlin (commercial) for cross-platform chaos engineering (fault injection, resilience testing) - installs the Gremlin agent on Linux / Windows / Kubernetes, picks attack types (resource, network, state, request), chains attacks into Scenarios (chaos experiments), integrates with the Reliability Score for forward-looking metrics. Use when the platform spans multiple environments (bare metal + cloud + serverless) and the team needs a commercial-supported solution per Gremlin's multi-platform support.
litmus-chaos
Configures LitmusChaos for Kubernetes-native chaos engineering - installs via Helm, picks ChaosExperiments from the ChaosHub (`pod-delete`, `network-latency`, `node-cpu-hog`, etc.), authors a ChaosEngine CR scoping the experiment + steady-state probes, runs as part of the cluster, exports Prometheus metrics for the verdict. Use when the platform is Kubernetes (CNCF-hosted; cloud-native). Prefer over chaos-mesh when the team wants a ChaosCenter web UI for workflow scheduling and ChaosHub catalog browsing; use chaos-mesh for fine-grained network-fault policies via its own CRD family.
steady-state-hypothesis-validator
Validates a chaos experiment's steady-state hypothesis before execution: checks that each probe metric is measurable and observable, that a recent baseline exists, that tolerances are numerically meaningful and SLI-backed, that the measurement window is defined, and that the chosen metrics would actually move under the target failure mode. Use when a chaos experiment has been authored (via chaos-experiment-author) and the team needs a pre-flight verdict before running the drill in any environment.