ghz-load
Wraps ghz, the gRPC load testing tool, for throughput and latency benchmarking. Covers test invocation (--proto + --call + host:port; or --protoset for compiled descriptors), load parameters (-n total requests, -c concurrency, -r RPS rate limit, -z duration), output formats (json/csv/html/influx-summary for CI consumption), the metrics reported (RPS achieved, latency p50/p95/p99, status-code distribution, errors), and CI integration patterns for regression gating. Use when benchmarking a gRPC service's throughput or detecting latency regressions in CI.
Install with skills.sh (any agent)
npx skills add testland/qa --skill ghz-loadghz-load
Overview
Per ghz.sh/docs/usage (opens in new window), ghz accepts a .proto (or compiled protoset), a method, a host:port, and load parameters, and emits per-request metrics + a summary. This skill wraps ghz for two use cases: ad-hoc throughput measurement and CI regression gating.
When to use
Authoring
Install
Per ghz.sh/docs/install (opens in new window):
# Homebrew
brew install ghz
# Go install
go install github.com/bojand/ghz/cmd/ghz@latestVerify:
ghz --versionConfigure a config.json
Per ghz docs, --config=path reads a JSON / TOML config. For reproducibility, commit it to the repo:
{
"proto": "./proto/user.proto",
"import-paths": ["./proto", "./vendor"],
"call": "user.v1.UserService/GetUser",
"host": "localhost:8080",
"insecure": true,
"total": 10000,
"concurrency": 50,
"rps": 0,
"data": {
"id": "user-1"
},
"format": "json",
"output": "ghz-report.json",
"skipFirst": 100
}skipFirst: 100 discards the first 100 requests (cold-cache / warmup). Set per-service.
Running
Basic load test
ghz --proto=./proto/user.proto \
--call=user.v1.UserService/GetUser \
--insecure \
-n 10000 -c 50 \
-d '{"id": "user-1"}' \
localhost:8080Per ghz.sh/docs/usage (opens in new window):
| Flag | Meaning |
|---|---|
--proto | Path to the .proto file |
--protoset | Path to a compiled descriptor set (alternative) |
--import-paths | Comma-separated proto import paths |
--call | package.Service/Method |
--insecure | "Use plaintext and insecure connection" |
-n, --total=N | "Number of requests to run. Default is 200" |
-c, --concurrency=N | "Number of request workers to run concurrently" |
-r, --rps=N | "Requests per second (RPS) rate limit"; 0 = unlimited |
-z, --duration=N | Total duration (30s, 5m) - alternative to -n |
-t, --timeout=N | Per-request timeout (default 20s) |
-d | JSON message payload |
-D | Path to a JSON file containing the payload |
Rate-limited (RPS pinning)
ghz --proto=./proto/user.proto \
--call=user.v1.UserService/GetUser \
--insecure \
-c 50 -z 60s -r 200 \
-d '{"id":"user-1"}' \
localhost:8080Run for 60s, 50 concurrent workers, capped at 200 RPS. Useful for confirming the service can sustain the target rate.
Duration mode
ghz -z 5m -c 100 ...When testing for stability / soak, -z beats -n - the test ends after N minutes regardless of throughput.
Streaming methods
Streaming RPCs aren't natively load-tested by ghz; it sends one unary call per worker per request. For streaming load see grpc-streaming-test-author.
Parsing results
Summary (default summary format)
Summary:
Count: 10000
Total: 20.45 s
Slowest: 120.34 ms
Fastest: 2.15 ms
Average: 10.23 ms
Requests/sec: 488.94Status code distribution:
Status code distribution:
[OK] 9983 responses
[DeadlineExceeded] 17 responsesPer grpc-status-code-mapping-reference, any non-OK is a flag for investigation.
JSON output for CI consumption
ghz --config=ghz.config.json
# Writes ghz-report.jsonSchema highlights:
{
"count": 10000,
"total": 20450000000,
"average": 10230000,
"fastest": 2150000,
"slowest": 120340000,
"rps": 488.94,
"latencyDistribution": [
{"percentage": 50, "latency": 8000000},
{"percentage": 95, "latency": 25000000},
{"percentage": 99, "latency": 80000000}
],
"statusCodeDistribution": {"OK": 9983, "DeadlineExceeded": 17},
"errorDistribution": {}
}Latencies are nanoseconds.
Other formats
Per ghz.sh/docs/usage (opens in new window), --format options: summary (default), csv, json, pretty, html, influx-summary, influx-details. Use html for shareable single-file reports; influx-* to ship metrics to InfluxDB.
CI integration
Baseline regression gate
# .github/workflows/grpc-perf.yml
name: grpc-perf
on:
pull_request:
paths:
- "service/**"
- "proto/**"
jobs:
ghz-baseline:
runs-on: ubuntu-latest
services:
service-under-test:
image: my-grpc-service:pr-${{ github.event.pull_request.number }}
ports: [8080]
steps:
- uses: actions/checkout@v5
- name: Install ghz
run: |
curl -L https://github.com/bojand/ghz/releases/download/v0.120.0/ghz-linux-x86_64.tar.gz | tar xz
sudo mv ghz /usr/local/bin/
- name: Warm + load
run: |
ghz --config=tests/perf/ghz.config.json
- name: Restore baseline
uses: actions/cache@v4
with:
path: baseline-ghz-report.json
key: ghz-baseline-${{ github.base_ref }}
- name: Compare
run: python tests/perf/compare-ghz.py baseline-ghz-report.json ghz-report.jsoncompare-ghz.py checks p99 latency is within +10% of baseline; fails otherwise:
import json, sys
baseline = json.load(open(sys.argv[1]))
current = json.load(open(sys.argv[2]))
def p99(report):
for entry in report["latencyDistribution"]:
if entry["percentage"] == 99:
return entry["latency"]
return None
p99_baseline = p99(baseline) / 1_000_000 # ms
p99_current = p99(current) / 1_000_000
delta = (p99_current - p99_baseline) / p99_baseline
if delta > 0.10:
print(f"❌ p99 regressed: {p99_baseline:.1f}ms → {p99_current:.1f}ms ({delta*100:.1f}%)")
sys.exit(1)
print(f"✅ p99: {p99_baseline:.1f}ms → {p99_current:.1f}ms ({delta*100:+.1f}%)")Standalone benchmark report
ghz --config=ghz.config.json --format=html --output=ghz-report.htmlGenerate a single HTML report attached to the PR for human review of distribution shape (long tail, bimodal, etc.).
Anti-patterns
| Anti-pattern | Why it fails | Fix |
|---|---|---|
-n 100 for a "load test" | Sample size too small; metrics noisy | At least -n 5000 or -z 30s |
No skipFirst | Cold cache / JIT warmup inflates latencies | skipFirst: ~5-10% of total |
Unbounded --concurrency | Tests the load generator, not the service | Match -c to expected production concurrency |
| Single-payload load test | Misses cache / branch-prediction noise | Vary -d payloads via -D <file> |
| Compare summary across runs without statistical context | Single-run noise → false regressions | Run N=3 times; compare distributions, not single numbers |
--insecure against TLS-required services | Connection failure dominates results | Match prod TLS config |
| Treating non-OK as transport failure | Status codes have meaning per grpc-status-code-mapping-reference | Inspect distribution; classify per AIP-194 |
| Load-testing on shared CI runner | Other jobs perturb CPU; noisy | Dedicated runner or isolate via Docker resource limits |
Limitations
References
Related skills
buf-cli-lint-breaking-build
Wraps the buf CLI for protobuf PR gating: `buf build` (compile .proto), `buf lint` (STANDARD rules: snake_case fields, Service suffix), `buf breaking --against {ref}` (detect wire/codegen breakage vs a git/BSR baseline), and `buf format`. Use as the CI proto-lint + breaking-change gate, or to debug a breaking failure by rule ID (e.g. FIELD_NO_DELETE_UNLESS_NUMBER_RESERVED) and pick the FILE/PACKAGE/WIRE_JSON/WIRE ruleset per consumer. This is the detection TOOL that enforces the rules; for the catalog of what is breaking and why use protobuf-versioning-strategy-reference, and for cross-service schema contract testing use protobuf-compat-checking - not this.
grpc-interceptor-test-author
Authors unit tests for gRPC interceptor logic: Go grpc.UnaryServerInterceptor/UnaryClientInterceptor, Java ServerInterceptor/ClientInterceptor, and grpc-js client interceptors. Covers auth (Unauthenticated on bad token), retry (backoff on Unavailable), logging/tracing (metadata extraction + propagation), error-mapping (status translation), and chained interceptor ordering - by calling the interceptor directly with a spy handler, no live backend. Use when a gRPC interceptor is written or modified. Different test surface from grpc-streaming-test-author (multi-message stream sequences) and grpc-mock (service handler logic) - use those, not this, for streams or handlers.
grpc-mock
Wraps gRPC server-mocking patterns for client-side tests: Go bufconn (in-memory net.Listener via google.golang.org/grpc/test/bufconn) + mockgen-generated interface mocks, Python pytest-grpc fixtures + unittest.mock patching of stubs, JVM grpc-mock library / in-process gRPC server (InProcessServerBuilder), Node @grpc/grpc-js fake server with NewServer-on-port-0. Use when writing client-side tests that need a controllable gRPC server response (success cases, error cases per grpc-status-code-mapping-reference, timeouts, and single-response error injection) without spinning up a real backend. For multi-message streaming-sequence tests (server-streaming, bidi), use grpc-streaming-test-author instead. Distinct from grpcurl-cli (ad-hoc CLI invocation against a real server) and ghz-load (perf against a real server).
grpc-status-code-mapping-reference
Pure-reference catalog of gRPC standard status codes - the 17 canonical codes (OK..UNAUTHENTICATED), their numeric values, semantics, retry behaviour per AIP-194 (only UNAVAILABLE is auto-retry-safe), and the gRPC-to-HTTP status mapping used by grpc-gateway (NOT_FOUND→404, INVALID_ARGUMENT→400, PERMISSION_DENIED→403, UNAUTHENTICATED→401, RESOURCE_EXHAUSTED→429, FAILED_PRECONDITION→400 not 412, ABORTED→409, UNAVAILABLE→503, DEADLINE_EXCEEDED→504, etc.). Use when designing a gRPC service's error vocabulary, writing assertions in gRPC client tests, configuring retry policies, or mapping gRPC errors to HTTP via a gateway. Consumed by buf-cli-lint-breaking-build, ghz-load, grpcurl-cli, grpc-mock, grpc-streaming-test-author.
grpc-streaming-test-author
Workflow-driven skill that builds gRPC streaming-RPC test suites from a proto definition. Classifies each RPC by pattern (unary, server-streaming, client-streaming, bidi), then emits the required categories per pattern - ordering preservation, completion semantics (server close after stream end, client half-close), cancellation, deadline handling, partial-stream failure. Produces skeletons for Go (bufconn + Send/Recv), Python (iterators), JVM (StreamObserver), Node (call.write/end). Use when adding tests for a new streaming RPC or auditing a suite for uncovered categories. Different test surface from grpc-interceptor-test-author (interceptor layer) and grpc-mock (harness); for wire-level streaming semantics use grpc-streaming-tests, not this.
grpcurl-cli
Wraps grpcurl, the curl-equivalent CLI for gRPC. Covers descriptor sources (server reflection default, --import-path + --proto for proto files, --protoset for compiled descriptor sets), service discovery (`list`, `describe`), invoking unary RPCs (`-d '{...}'`, `-d @file.json`, `-d @` for stdin), streaming RPCs (newline-delimited JSON via stdin), TLS configuration (--cacert, --cert, --key, --insecure, --plaintext), header injection (-H 'Authorization: Bearer ...'), and exit codes. Use for ad-hoc gRPC debugging, smoke testing, scriptable PR-time gates, and CLI-based interaction with reflective gRPC services.
protobuf-versioning-strategy-reference
Pure-reference catalog of protobuf3 versioning and breaking-change rules: field-number reservation (reserve on delete; 1..536870911; 19000-19999 reserved), wire-safe vs wire-incompatible changes (add/remove safe with reservation; changing a field number always breaks), compatible type conversions (int32/uint32/int64/uint64/bool; sint32/sint64; string/bytes for UTF-8; enum/int), oneof + map constraints, and buf's four breaking categories (FILE/PACKAGE/WIRE_JSON/WIRE) with rule IDs. Use when designing a schema change or picking a buf breaking ruleset. This is the catalog of what is breaking and why, not a scanner; to detect changes in CI use buf-cli-lint-breaking-build, for the gRPC status-code vocabulary use grpc-status-code-mapping-reference, and for cross-service contract testing use protobuf-compat-checking.