go-race-detector-workflow
Runs the Go race detector and goroutine-leak checker end-to-end: instrument with `go test -race`, read race reports, configure GORACE options, stress with `-count`/`-cpu`, detect goroutine leaks with go.uber.org/goleak, and gate both checks in CI. Use when a Go service has shared state accessed by concurrent goroutines, when a race-related incident needs a regression harness, or when adding `-race` to a CI matrix for a Go module. Does not cover barrier-based deterministic interleaving or forced goroutine scheduling; use race-condition-test-author for that.
Install with skills.sh (any agent)
npx skills add testland/qa --skill go-race-detector-workflowgo-race-detector-workflow
The race detector (compiled into the binary via ThreadSanitizer) and goroutine leaks are two separate failure classes: the detector finds concurrent unsynchronized access, not leaked goroutines. This skill walks both checks, from first run to CI gate.
race-condition-test-author covers multi-language deterministic interleaving (barriers, jcstress, TSan for C/C++). This skill focuses exclusively on Go: the -race flag, GORACE tuning, stress amplification, and goleak.
Step 1 - Enable the race detector
Per go.dev/doc/articles/race_detector (opens in new window), add -race to any go command:
go test -race ./...
go run -race main.go
go build -race ./cmd/serverThe flag compiles ThreadSanitizer instrumentation into the binary. It requires cgo and a C compiler (on Linux/FreeBSD/Windows; Darwin ships its own). Supported platforms as of Go 1.22: linux/amd64, linux/arm64, linux/ppc64le, linux/s390x, linux/loong64, freebsd/amd64, netbsd/amd64, darwin/amd64, darwin/arm64, windows/amd64 (mingw-w64 runtime v8+ required).
Per go.dev/doc/articles/race_detector (opens in new window), expected overhead:
Step 2 - Read a race report
A detected race prints two goroutine stacks to stderr:
WARNING: DATA RACE
Write at 0x00c0000b4010 by goroutine 7:
main.(*Cache).set+0x6c
/home/user/app/cache.go:38
Previous read at 0x00c0000b4010 by goroutine 6:
main.(*Cache).get+0x44
/home/user/app/cache.go:22
Goroutine 7 (running) created at:
main.runWorker+0x34
/home/user/app/main.go:71The report names the conflicting accesses (read vs. write), the memory address, and the goroutine creation sites. Fix by protecting all accesses to the address with the same synchronization primitive (mutex, atomic, or channel hand-off).
Step 3 - Tune GORACE options
Per go.dev/doc/articles/race_detector (opens in new window), set GORACE before the command:
GORACE="log_path=/tmp/race/report halt_on_error=1 history_size=2" \
go test -race ./...Useful options:
| Option | Default | When to change |
|---|---|---|
log_path | stderr | Set to a file path so CI can archive race reports as artifacts |
halt_on_error | 0 | Set to 1 to stop immediately on first race; useful for local debugging |
history_size | 1 | Increase to 2-7 when report stacks look truncated (trades memory for depth) |
strip_path_prefix | "" | Strip module root from paths so report lines are repo-relative |
exitcode | 66 | Override if your CI treats specific exit codes differently |
Step 4 - Stress with -count and -cpu
The race detector only fires on races that actually execute. A single go test -race run on a lightly-contended path may produce zero output and still miss a real race. Amplify coverage:
# Run each test 10 times per package
go test -race -count=10 ./...
# Exercise multiple GOMAXPROCS values
go test -race -cpu=1,2,4,8 ./...
# Combine: 5 runs at each GOMAXPROCS
go test -race -count=5 -cpu=1,2,4 ./...-cpu sets GOMAXPROCS for each comma-separated value, then re-runs. Running at GOMAXPROCS=1 surfaces sequencing bugs; higher values surface true parallel races. Combining both increases scheduler interleaving diversity without extra code.
Step 5 - Check loop-variable capture with go vet
Per the Go vet documentation at go.dev/cmd/vet (opens in new window) (loopclosure: "check references to loop variables from within nested functions"), go vet flags the classic loop-variable-capture anti-pattern that frequently causes races when goroutines close over a range variable:
// Before Go 1.22 - race: all goroutines capture the same &v
for _, v := range items {
go func() { process(v) }() // vet warns here
}
// Fix: copy the variable
for _, v := range items {
v := v
go func() { process(v) }()
}Run before -race to filter out this class early:
go vet ./...
go test -race ./...In Go 1.22+, range variables are per-iteration by default; the capture pattern is still worth auditing in code that may be compiled with older toolchains.
Step 6 - Detect goroutine leaks with goleak
A goroutine that starts but never stops is a leak: the race detector ignores it (no concurrent access violation), but the goroutine holds resources and inflates memory over time.
Install per github.com/uber-go/goleak (opens in new window):
go get -u go.uber.org/goleakPer-test: VerifyNone
import "go.uber.org/goleak"
func TestWorkerPool(t *testing.T) {
defer goleak.VerifyNone(t)
pool := NewWorkerPool(4)
pool.Submit(func() { /* work */ })
pool.Shutdown()
// VerifyNone fires after Shutdown() returns;
// any still-running worker goroutine fails the test.
}Per github.com/uber-go/goleak (opens in new window), VerifyNone is incompatible with t.Parallel(): goleak cannot associate a specific goroutine with a specific parallel sub-test.
Package-level: VerifyTestMain
For packages that use t.Parallel(), wrap the test runner instead:
func TestMain(m *testing.M) {
goleak.VerifyTestMain(m)
}VerifyTestMain runs the full test binary, then checks for leaked goroutines once all tests have completed.
Filtering expected goroutines
Third-party libraries sometimes leave intentional background goroutines. Silence a known one by its top-of-stack function:
goleak.VerifyNone(t,
goleak.IgnoreTopFunction("database/sql.(*DB).connectionOpener"),
)Full filter-option catalog (IgnoreAnyFunction, IgnoreCurrent, Cleanup, and when to prefer each) is in references/goleak-filter-options.md.
Step 7 - CI matrix
Gate both checks in CI. Run -race in at least one matrix dimension (per go.dev/doc/articles/race_detector (opens in new window): "It is recommended to always run race-enabled tests"):
jobs:
test:
strategy:
matrix:
go-version: ["1.22", "1.23"]
race: ["", "-race"]
steps:
- uses: actions/setup-go@v5
with:
go-version: ${{ matrix.go-version }}
- name: Run tests
env:
GORACE: "log_path=/tmp/race/report halt_on_error=0"
run: |
go vet ./...
go test ${{ matrix.race }} -count=3 -cpu=1,4 -timeout=10m ./...
- name: Upload race reports
if: failure()
uses: actions/upload-artifact@v4
with:
name: race-reports-${{ matrix.go-version }}-${{ matrix.race }}
path: /tmp/race/report*Because -race adds the Step 1 execution overhead, set -timeout to at least 5-10x your non-race run time. Upload log_path files on failure so the report survives the run.
Anti-patterns
| Anti-pattern | Why it fails | Fix |
|---|---|---|
Run -race once, see no output, ship | Race detector only finds races that execute in that run | Use -count/-cpu matrix (Step 4) |
Skip -race in CI for "release" builds | Race that appears in production, not in CI | Gate at least one matrix dimension with -race (Step 7) |
Use defer goleak.VerifyNone(t) with t.Parallel() | goleak cannot associate goroutines to parallel sub-tests | Use VerifyTestMain instead (Step 6) |
IgnoreCurrent() at test-file scope | Snapshot is taken once at import time; masks leaks added before each test | Call IgnoreCurrent() inside each test function, not at package init |
Trust -race to catch goroutine leaks | -race detects concurrent unsynchronized access, not leaked goroutines | Add goleak (Step 6); both gates are complementary |
Set history_size to max (7) always | 128K access history per goroutine multiplies memory cost; can OOM CI runners | Start at 1; raise only when reports show truncated stacks |
Limitations
References
goleak filter options
View source (opens in new window)goleak filter options
Filter options for goleak.VerifyNone / goleak.VerifyTestMain, per pkg.go.dev/go.uber.org/goleak (opens in new window). Pass one or more as trailing arguments to suppress goroutines that are expected rather than leaked.
IgnoreTopFunction
Ignores any goroutine whose top-of-stack frame is the named function. Prefer this when the library goroutine is identifiable by name.
goleak.VerifyNone(t,
goleak.IgnoreTopFunction("database/sql.(*DB).connectionOpener"),
)IgnoreAnyFunction (v1.3.0+)
Ignores any goroutine whose stack contains the named function at any depth, not just the top frame. Use when the identifying frame is not at the top.
goleak.VerifyNone(t,
goleak.IgnoreAnyFunction("google.golang.org/grpc.(*ccBalancerWrapper).watcher"),
)IgnoreCurrent
Snapshots the goroutines already running at call time and ignores exactly those at verification.
opt := goleak.IgnoreCurrent()
// ... test logic ...
goleak.VerifyNone(t, opt)Prefer IgnoreTopFunction over IgnoreCurrent when the library goroutine is identifiable by name: IgnoreCurrent silences goroutines that were already running at snapshot time, which can mask leaks introduced before the snapshot.
Cleanup
goleak.Cleanup(func(int)) registers a function goleak calls with the exit code after the leak check, e.g. to log instead of failing the process. Used with VerifyTestMain when the default exit behavior needs to change.
References
Related skills
async-ordering-tests
Test async ordering - event-loop / queue / channel ordering assertions, JS Promise microtask vs macrotask ordering, Python `asyncio.gather` vs `asyncio.wait_for` semantics, Go goroutine + channel happens-before relationships, async/await re-entrancy. Use deterministic schedulers (sinon fake timers, asyncio test mode) to remove run-to-run variance. Use when a callback fires twice, a later response overwrites an earlier one, or a cancelled parent task leaves a child still running - bugs where completion order, not shared memory, is the defect.
deadlock-detection-harness
Build deadlock-detection harnesses - extract lock-acquire-order graph via instrumentation, run cycle detection (DFS) to spot inconsistent ordering, use lock-acquire timeouts to surface rather than hang, JVM `jstack` / `gdb thread apply all bt` for postmortem analysis. Pair with ThreadSanitizer's `detect_deadlocks=1` for runtime detection. Use when a service that holds two or more locks hangs in production with no crash or error, or before release when lock acquisition order across code paths has never been proven consistent.
jepsen-patterns
Reference for Jepsen-style distributed-systems testing - consistency models hierarchy (linearizability vs sequential vs causal vs monotonic-reads vs eventual), nemesis primitives (network partitions, clock skew, kill nodes), workload generators, Knossos + Elle linearizability checkers. Reference-only because Jepsen tests are typically Clojure-bespoke per system; use this skill to evaluate vendor claims and structure your own test. Use when a datastore vendor advertises a consistency guarantee that has to be checked before adoption, when reading a published Jepsen report for gaps, or when a custom replicated store needs its own consistency test scoped.
mvcc-isolation-tests
Build per-database MVCC isolation-level tests - Read Uncommitted vs Read Committed vs Repeatable Read vs Serializable; verify which anomalies are prevented at each level (dirty read, non-repeatable read, phantom read, serialization anomaly, write skew). Per PostgreSQL transaction isolation docs; analogous patterns for MySQL InnoDB, SQL Server, and DynamoDB. Use when two concurrent transactions can touch the same rows (balance debit, seat booking, stock decrement), or before changing a service's default isolation level.
race-condition-test-author
Build deterministic race-condition tests - identify shared mutable state, drive interleavings via barriers / latches / manual scheduling; use ThreadSanitizer (clang `-fsanitize=thread`) for C/C++/Go data race detection; use jcstress (`@JCStressTest` + `@Actor` + `@Outcome`) for JVM stress; use Loom virtual-thread interleavings for parallel testing. Use when a defect only reproduces under load on shared in-process state (cache, counter, connection pool, lazy-init singleton), or when writing the regression test for a race-condition incident before the fix lands.