test-isolation-patterns
Pure reference catalog of test-isolation and fixture-lifecycle patterns - the four-phase test pattern (Meszaros), fixture scope (per-test / per-describe / shared / global), the Fresh-Fixture vs Shared-Fixture trade-off (Fowler), parallel-safety patterns, and cleanup discipline (afterEach / afterAll / tagged-cleanup), plus a pattern-selection guide and a worked leaking-state diagnosis. The database-isolation strategies (transaction-rollback / database-per-worker / template-database) and network / external-service stubbing live in references/. This is the architecture-tier reference, not a file-level fixture-coupling style rule. Use when designing fixture scope and isolation strategy, auditing fixture coupling or retry/wait policy, or moving a suite to parallel execution.
Install with skills.sh (any agent)
npx skills add testland/qa --skill test-isolation-patternstest-isolation-patterns
Overview
A test that fails sometimes for non-obvious reasons is non-deterministic. Per Martin Fowler - Eradicating Non-Determinism in Tests (opens in new window): "A test is non-deterministic when it passes sometimes and fails sometimes, without any noticeable change in the code, tests, or environment… Once you start ignoring a regression test failure, then that test is useless and you might as well throw it away." The dominant cause is broken isolation - one test affecting another, the environment leaking, fixtures sharing state. This catalog is the canonical reference for the isolation patterns that prevent it.
This skill is a pure reference - no execution steps. It is the catalog cited when auditing fixture coupling, retry/wait policy consistency, and CI integration health. It complements test-code-conventions §6 (which is the file-level rule against global-fixture hubs) with the cross-cutting architecture patterns. It also complements flake-pattern-reference, which catalogs flake symptoms; this skill catalogs the prevention patterns.
When to use
How to use
Pattern 1 - The four-phase test pattern
Canonical source: Gerard Meszaros - xUnit Test Patterns: Refactoring Test Code (2007) (opens in new window). Referenced in the Wikipedia entry on test fixture (opens in new window).
Every test has four phases:
| Phase | What |
|---|---|
| 1. Setup | Establish the pre-conditions / fixture |
| 2. Exercise | Interact with the System Under Test |
| 3. Verify | Determine whether the expected outcome was obtained |
| 4. Teardown | Return to a clean state |
Phases 1 and 4 together are fixture management. Patterns 2 through 5 below cover how to do them safely; isolating external stores and services is covered in the deep references.
Pattern 2 - Fixture scope
The framework's test runner offers three or four scopes; the team picks the tightest scope that meets the constraint.
| Scope | Lifecycle | Use when |
|---|---|---|
| Per-test (function-scoped) | Setup before each test; teardown after each | Default. Maximally isolated. Slowest. Always parallel-safe. |
| Per-describe (class / module-scoped) | Setup before the first test in the group; teardown after the last | Setup is expensive and the group of tests genuinely shares it (read-only) |
| Shared (session / worker-scoped) | Setup once for the whole run; teardown at end | Setup is unaffordable per-describe (e.g., spinning up a Docker stack) and the tests don't mutate it |
| Global (module-loading) | Setup at module-import time; no teardown | Anti-pattern in nearly all cases. Use only for truly immutable language-level fixtures (constants, configuration). |
The single rule that prevents most flake: never share mutable fixtures across tests. If a fixture is mutated by any test, it must be per-test scoped.
Framework-specific scope syntax (illustrative; cite the per-framework skill for tool-specific details)
Anti-patterns
| Anti-pattern | Why it fails |
|---|---|
| Per-describe fixture that any test in the describe mutates | One test fails; the next "starts" from the mutated state |
| Shared fixture mutated through a leaky abstraction (e.g., factory returns a shared object) | Cross-test mutation without an obvious culprit; flake follows |
| Per-test scope for genuinely expensive setup (a 30s Docker spin-up per test) | Suite time explodes; team skips tests |
| Global fixture for anything that has state | Cannot reset between test runs; CI run pollutes the next run |
Inheritance hierarchy of fixtures (BaseTest → AppTest → DomainTest → SpecificTest) | Depth-3+ chains break unpredictably (§A2) |
Pattern 3 - Fresh Fixture vs Shared Fixture trade-off
Canonical source: Martin Fowler - Eradicating Non-Determinism in Tests (opens in new window).
Fowler's framing: "I prefer the former [Fresh Fixture], as it's often easier - and in particular easier to find the source of a problem… [but] rebuilding the database each time can add a lot of time to test runs, so that argues for switching to a clean-up strategy."
| Approach | Setup cost | Isolation | When |
|---|---|---|---|
| Fresh Fixture (rebuild from scratch every test) | High | Maximum | Default; use unless measured slow |
| Cleanup strategy (preserve the fixture, undo changes at teardown) | Low | Strong if cleanup is comprehensive | When Fresh Fixture's cost is prohibitive |
| Persistent Fresh Fixture (fresh per test, persisted via transaction-rollback) | Low | Maximum | The pragmatic middle for DB-backed tests |
The transaction-rollback pattern (Persistent Fresh Fixture): Begin a transaction at test start; do all the test's DB work inside it; rollback at test end. The database is materially unchanged across tests. The pattern works for any DB that supports transactions; integration-test frameworks like DatabaseCleaner (Ruby), pytest-django's db fixture, Spring's @Transactional test annotation all implement it. The five DB-level strategies are catalogued in references/database-store-isolation.md.
Anti-patterns
| Anti-pattern | Why it fails |
|---|---|
| Fresh Fixture that takes 60+ seconds per test | Suite time becomes prohibitive; team skips tests |
| Cleanup strategy that misses one mutation surface (cache; queue; file system) | Cross-test coupling through the missed surface |
| Transaction-rollback that doesn't actually rollback (autocommit, DDL changes) | Silent state leakage |
| Shared Fixture documented as "immutable" but tests mutate it anyway | The documentation is unverified; flake follows |
Pattern 4 - Parallel safety
Canonical source: Fowler - Eradicating Non-Determinism in Tests (opens in new window) on isolation as the parallel-safety prerequisite, plus Luo et al. FSE 2014 (opens in new window), which attributes 20% of flakes to concurrency problems (race conditions and deadlocks).
Parallel execution magnifies every isolation bug. The patterns that make parallel safe:
| Pattern | What it does |
|---|---|
| Worker-scoped fixtures | Each parallel worker has its own state (DB, file system path, port range) |
| Unique identifiers per test | Test names, file paths, generated IDs include the worker ID (worker_${WORKER_ID}_user_${TEST_ID}) |
| Ephemeral output paths | Tests write to tmp/${WORKER_ID}/${TEST_ID}/ and clean up at teardown |
| Port range allocation | Each worker gets a port range (30000 + WORKER_ID * 100) to avoid binding conflicts |
| No global singletons | No process.env writes, no global config mutation, no static state |
| Idempotent setup | Re-running the setup produces the same state (so a flaky-and-retried test isn't tainted) |
Anti-patterns
| Anti-pattern | Why it fails |
|---|---|
process.env.X = "..." in a test (writes to a shared global) | Worker N's env-write affects worker M's reads |
| Hard-coded port 3000 in tests (port collisions) | First worker binds; others fail |
Tests writing to /tmp/test.log (path collision) | Workers stomp each other's files |
| Test-name-based DB seeding (collides across workers if names overlap) | Cross-worker state pollution |
Per-test setup that does setTimeout / sleep to "let things settle" | Flake source: async-wait is the largest flake category at 45% per Luo et al. 2014 (opens in new window); use proper event-based synchronisation |
Pattern 5 - Cleanup discipline
Canonical source: Meszaros's xUnit Test Patterns (2007) - the Garbage-Collected Teardown vs In-line Teardown vs Implicit Teardown vs Setup Decorator patterns.
The four canonical cleanup approaches:
| Pattern | Mechanism |
|---|---|
| In-line Teardown | Each test explicitly cleans up at end (last line of the test body) |
| Implicit Teardown | afterEach / afterAll hooks the runner calls automatically |
| Garbage-Collected Teardown | Cleanup happens when the language's GC reclaims the fixture (typed in C# / Java with IDisposable / AutoCloseable) |
| Tagged Cleanup | Fixture registers itself with a "cleanup queue" at setup; queue drains at suite end |
Rule: Implicit Teardown via the runner's afterEach hook is the default. In-line Teardown is acceptable when the cleanup is specific to one test. Tagged Cleanup is for fixtures whose lifetime is variable (held across multiple tests, then released).
Anti-patterns
| Anti-pattern | Why it fails |
|---|---|
| No teardown ("the next test will clean up") | Failing test orphans state; the next test fails too |
| Teardown that swallows errors silently | Real cleanup failures are invisible; flake follows |
Teardown that depends on test-pass state (if (test.passed) cleanup()) | Failing tests don't clean up; cascading flake |
| Teardown order-dependent on setup order | Refactoring setup breaks teardown |
Worked example - diagnosing a leaking-state test
Symptom. updates a member's role passes when its file runs alone but fails intermittently in the full suite - the classic order-dependent flake (Fowler (opens in new window)).
The suite (before):
describe("member admin", () => {
let user;
beforeAll(() => { user = createUser({ role: "member" }); }); // per-describe scope
it("promotes to admin", () => {
promote(user);
expect(user.role).toBe("admin"); // mutates the shared fixture
});
it("updates a member's role", () => {
expect(user.role).toBe("member"); // fails: user was mutated above
});
});Diagnosis. Walk Pattern 2: user is a per-describe (beforeAll) fixture that the first test mutates. The second test inherits the mutated state. This is the Pattern 2 anti-pattern row "Per-describe fixture that any test in the describe mutates", and it violates the single rule that prevents most flake: never share mutable fixtures across tests.
Fix. Move the fixture to per-test (beforeEach) scope so each test gets a Fresh Fixture (Pattern 3):
describe("member admin", () => {
let user;
beforeEach(() => { user = createUser({ role: "member" }); }); // per-test scope
it("promotes to admin", () => {
promote(user);
expect(user.role).toBe("admin");
});
it("updates a member's role", () => {
expect(user.role).toBe("member"); // always fresh; order-independent
});
});Each test now starts from a clean object - maximally isolated and parallel-safe. If createUser were genuinely expensive, the middle path is Pattern 3's Persistent Fresh Fixture (transaction-rollback), not a shared mutable object.
Deep references
Once fixture scope is chosen, the two external-dependency concerns have their own deep catalogs:
Cross-cutting anti-patterns
| Anti-pattern | Why it fails |
|---|---|
| Implicit ordering (test B depends on test A's side effects) | Per Fowler (opens in new window): "isolation… gives you more flexibility in running subsets of tests and parallelizing tests." Ordering breaks both. |
| Tests that "sleep until it works" | Timing-fragile; async-wait is 45% of all flakes per Luo et al. 2014 (opens in new window) |
| Tests that read system time without overrides | Tests fail at midnight / DST / leap year |
| Tests that read random data without seeding | Non-reproducible failures |
| Tests that depend on file-system layout | OS / CI-runner-specific failures |
| Tests that depend on locale / timezone of the runner | Internationalisation-dependent flake |
Pattern-selection guide
| Scenario | Recommended pattern |
|---|---|
| Default (unit / integration test) | Per-test fixture scope + Fresh Fixture |
| DB-backed integration test | Per-test fixture + transaction-rollback (Persistent Fresh Fixture) |
| Slow expensive E2E setup | Per-describe Shared Fixture documented as immutable + transactional teardown |
| Parallel execution | Worker-scoped DB + unique IDs per worker + ephemeral output paths |
| External service interaction | Stubs by default; contract tests at API surface; real-network only in smoke / canary |
| Multi-worker DB-heavy suite | Database-per-worker + template-database cloning |
| Mutation-heavy unit tests | Per-test fixture + in-memory mock |
Hand-off targets
References
Database and external-store isolation
View source (opens in new window)Database and external-store isolation
Deep reference for test-isolation-patterns SKILL.md. Consult after picking fixture scope, when the system under test reads or writes a database, cache, or queue. This is the dominant source of test flake at scale; five canonical strategies, each with trade-offs.
Transaction-rollback (the default)
Each test runs in a transaction; teardown rollbacks. Works for: relational DBs with full transaction support. Doesn't work for: DDL changes, multiple DB connections, queues, caches.
Database-per-test-worker
Each parallel worker gets its own database (named app_test_worker_1, app_test_worker_2, etc.). Created once at startup; reused across tests within the worker; dropped at suite end. Works for: parallel execution with mutation-heavy tests. Cost: pre-suite setup time + N× DB storage.
Template database / pristine clone
Pre-create a template database with seed data; clone it per test (or per worker). PostgreSQL's CREATE DATABASE … TEMPLATE template_db is the canonical mechanism. Works for: tests needing complex seed state. Cost: template maintenance.
Containerised DB-per-test
Each test gets a fresh Docker container (Testcontainers (opens in new window) is the canonical library). Maximum isolation; highest cost. Works for: integration tests where the DB version / extensions / config matter. Don't use for: unit tests.
In-memory substitution
Use SQLite in-memory instead of the production DB engine. Fast; works for simple SQL. Doesn't work for: production-specific features (PostgreSQL JSON, Postgres extensions, MySQL spatial types). Cited as an anti-pattern by Fowler on integration tests (opens in new window) when the production engine has features the in-memory substitute lacks.
Anti-patterns
| Anti-pattern | Why it fails |
|---|---|
| Tests that mutate a shared DB without isolation | Cross-test coupling; the dominant source of flake at scale |
| In-memory substitution masking production-engine differences | Tests pass locally; fail in production |
| Transaction-rollback for tests that do DDL (CREATE TABLE in test) | DDL is auto-commit in most engines; rollback doesn't undo it |
| Database-per-worker without a maximum-worker limit | Storage explodes; CI cost surges |
| Containerised DB-per-test for unit tests | 5-second container startup × 1000 unit tests = unworkable |
Network and external-service isolation
View source (opens in new window)Network and external-service isolation
Deep reference for test-isolation-patterns SKILL.md. Consult when a test touches a service the team does not control (third-party HTTP APIs, external endpoints). Tests should not depend on external services they don't control; three patterns cover the choice.
Stub (canned response)
Use when the test doesn't care about the network itself. Reach for a stub library - nock (opens in new window), WireMock (opens in new window), Mountebank (opens in new window) - or the msw-handlers / wiremock-stubs / mountebank-imposters skills in qa-test-data.
Contract test
Use when the test cares whether the service contract holds. Pact (opens in new window) or schemathesis (opens in new window) verify the contract rather than a canned body.
Real network call in a controlled environment
Use for a smoke / canary test in a staging tier with a dedicated test partition, where exercising the live service is the point of the test.
Anti-patterns
| Anti-pattern | Why it fails |
|---|---|
| Unit tests calling the real external API | Tests fail when the API is down; tests pass when the API silently changes |
| Stubs that drift from production response shape | Tests pass with stubs that don't match reality |
| One global stub for the whole suite | Tests cross-couple through the stub configuration |
| Contract test with no contract refresh | Stub goes stale; tests pass while production breaks |
See also the msw-handlers, wiremock-stubs, and mountebank-imposters skills in qa-test-data for stub implementation, and the SKILL's pattern-selection guide for when each pattern applies.
Related skills
object-model-patterns
Pure reference catalog of the canonical object-model architecture patterns for test automation frameworks - Page Object Model (Fowler), Screenplay (Marcano/Palmer/Hill), Component Object, App Actions (Cypress idiom), Service Object, Repository, and Screen Object (the desktop/mobile sibling of Page Object covering Windows UIA, macOS XCTest, Linux AT-SPI, Appium / Espresso) - each with its canonical citation, when-to-use rules, refuse-to-mix anti-patterns, and a worked example. This is the architecture-tier reference - what each pattern *is* - not file-level style rules and not tool-specific configuration. Use when designing, reviewing, or migrating a test framework's object-model architecture.
test-code-conventions
Pure-reference catalog of test-code conventions: AAA structure (Arrange / Act / Assert), per-test single-responsibility, descriptive naming (`{sut}_{scenario}_{expected}`), assertion specificity, mocking rationale (state vs behavior, fake vs mock), fixture-coupling rules, and the magic-number / hard-coded-string anti-patterns; the E2E selector-priority and web-first-assertion conventions live in references/. Use as the shared rule book a test-code review cites back to, or as onboarding for what makes a test code-reviewable; to score a test's quality on weighted axes use test-design-scorecard, and for setup/teardown isolation specifically use test-isolation-patterns.
test-design-scorecard
Scores test files 1 to 5 on six design axes (AAA phase separation, single-responsibility, naming, fixture coupling, magic literals, setup time) using explicit per-level anchors that settle what separates a 2 from a 4, then turns the scores into growth-framed feedback and a per-author trend report; the per-PR and per-author rollup examples and trend-reporting conventions live in references/. Owns the scoring and the write-up only: the conventions being scored live in a separate conventions catalog such as `test-code-conventions`, and block-or-approve gating belongs to an adversarial review. Use when a test diff needs a graded coaching read rather than a merge verdict: onboarding a new engineer, a team deliberately ramping up test discipline, or a quarterly per-author trend where the output is a conversation, not a gate.
test-framework-architecture-audit
Audits an existing test automation framework across eight architecture-tier axes and bands each one PASS, WARN, or FAIL: page-object coverage and purity, base-class inheritance depth, fixture scope and coupling, helper sprawl, naming-convention drift, retry and wait consistency, documented-versus-actual convention drift, and CI integration health. Carries the numeric cut behind every band and labels which cuts are practitioner conventions rather than published standards. Measures the framework's own structure (page objects, base classes, fixtures, helpers, conventions), not the suite's tier mix or flake rate, and not the design of a framework that does not exist yet. Use when a test framework has grown for a release or more without structural review, before a major refactor, or when a team suspects its written test conventions no longer match what the code actually does.
test-framework-blueprint
Build-an-X workflow that takes an SDET from no test suite to a complete framework design in seven steps - inventory the SUT, choose runner + language, directory layout + fixture architecture, object-model decision, test data + mocking wiring, reporting + CI integration, conventions doc + review gates - producing a written framework blueprint (directory tree, fixture list, chosen patterns, CI matrix) plus an implementation order. This is the whole-framework design workflow - not the Step 2 runner-choice decision on its own, not the Step 4 object-model pattern catalog it defers to, and not the scaffolder that generates the harness skeleton once the blueprint exists. Use when designing a test automation framework from scratch or re-architecting one that grew organically.
test-step-design-patterns
Pure reference catalog of test-step design patterns at the architecture tier - step granularity (one logical action per step), abstraction layers (mechanical → page → business), step extraction rules (when to inline / when to extract to a helper / when to extract to a Page Object method), the declarative-vs-imperative phrasing rule, FIRST principles (Fast / Independent / Repeatable / Self-validating / Timely), and the AAA / Given-When-Then mapping. This is the cross-framework architecture-tier reference for what a step IS, when it should exist, and where it should live - not file-level AAA style rules and not Gherkin-specific translation. Use when designing or reviewing the step layer of a test framework - for example when writing or reviewing E2E or integration tests, when the step count per test is high, or when refactoring recorded or codegen test output into readable steps.
test-suite-health-audit
Measures an existing test suite's current state on four axes: per-file tier classification (unit / integration / E2E, first match wins), pyramid ratio against whatever target the team already committed to, per-layer flake rate, and defects-caught-per-run-minute ROI per tier, then reduces them to one categorical verdict (Healthy, Needs pruning, Needs refactor, Cannot assess). Reports severity against a target ratio but never prescribes one: choosing the target unit:integration:E2E mix and the rebalancing plan belongs to a pyramid-balancing capability such as `test-pyramid-balancer`. Use when a suite has grown for a year or more without review and someone needs a defensible read on whether it is healthy, over-grown, or structurally inverted before deciding what to delete or rewrite.