test-step-design-patterns
Pure reference catalog of test-step design patterns at the architecture tier - step granularity (one logical action per step), abstraction layers (mechanical → page → business), step extraction rules (when to inline / when to extract to a helper / when to extract to a Page Object method), the declarative-vs-imperative phrasing rule, FIRST principles (Fast / Independent / Repeatable / Self-validating / Timely), and the AAA / Given-When-Then mapping. This is the cross-framework architecture-tier reference for what a step IS, when it should exist, and where it should live - not file-level AAA style rules and not Gherkin-specific translation. Use when designing or reviewing the step layer of a test framework - for example when writing or reviewing E2E or integration tests, when the step count per test is high, or when refactoring recorded or codegen test output into readable steps.
Install with skills.sh (any agent)
npx skills add testland/qa --skill test-step-design-patternstest-step-design-patterns
Overview
This skill is a pure reference - no execution steps. It is the catalog cited when auditing step granularity at the architecture tier - within an "Act" phase, what is one step? It complements test-code-conventions §1 (AAA structure at the file level).
When to use
How to use
Pattern 1 - FIRST principles
Canonical source: Robert C. "Uncle Bob" Martin - Clean Code (2008), chapter 9 "Unit Tests" (opens in new window), reaffirmed in Clean Coder and across his blog. The FIRST mnemonic is the foundational quality bar for every test step.
| Letter | Principle |
|---|---|
| F - Fast | Tests must be fast. Slow tests don't get run; tests that don't get run don't catch bugs. |
| I - Independent | Tests must not depend on each other. Each test sets up its own world. Per Fowler on test isolation (opens in new window), this is the prerequisite to parallel execution and selective re-runs. |
| R - Repeatable | Tests run consistently on any environment (laptop, CI, prod-like). No "runs only on Tuesdays" / "passes on Linux only." |
| S - Self-validating | Tests pass or fail on their own. No human reads logs to determine the outcome. |
| T - Timely | Tests are written close to (ideally before) the production code. Stale tests rot. |
FIRST is the underlying rationale for most of the patterns below. A step that violates FIRST is the smell; the patterns prescribe the fix.
Pattern 2 - Step granularity
A "step" is the minimal unit a test reader can name in business terms. One logical action per step - not one click, not one assertion, but one meaningful operation.
The single-purpose-step rule
Each step should do exactly one of: Arrange (setup), Act (the operation under test), Assert (verify outcome), or Annotate (logging / labelling). Steps that mix two are split.
| Smell | Refactor |
|---|---|
await page.click('#submit'); expect(toast).toBeVisible(); | Two steps: the Act (submit) and the Assert (toast visible). Split. |
const user = await createUser({...}); await login(user); | Two Arrange steps. Acceptable as one logical "Arrange a logged-in user" if the helper is named that way (Pattern 4). |
await page.click('#row-1'); await page.click('#row-1-edit'); await page.fill('#name', 'New'); | Three mechanical clicks comprising one business action ("edit row 1's name"). Extract to a single business step (Pattern 4). |
Step count per test
Aim for 3-8 steps per test body (counting Arrange / Act / Assert phases as steps). >15 is the threshold flagged as a smell - the test is doing too much or operating at the wrong abstraction.
Anti-patterns
| Anti-pattern | Why it fails |
|---|---|
One step that does everything: await fullCheckoutFlow(); | Single line of action; the test reads as "did the helper work?" not as "did the SUT work?" |
| 30 mechanical clicks comprising one logical flow | Brittle to UI changes; the test reads as a script, not a specification |
Mixed Act + Assert in one line (expect(await page.click(...)).toBeTruthy()) | Cannot distinguish "the action failed" from "the assertion failed" in the diagnostic |
One step with two unrelated assertions (expect(cart.total).toBe(10); expect(user.role).toBe('admin');) | Test fails on the first assert; the second is never evaluated. Split into single-responsibility tests |
Pattern 3 - Abstraction layers
The dominant test-code smell at scale is mechanical steps in the test body. The fix is layered abstraction. Three layers, named consistently:
| Layer | Vocabulary | Example |
|---|---|---|
| Business layer (the test body) | Domain verbs the PM / business stakeholder would recognise | customer.signsIn(), cart.addsItem(sku), checkout.placesOrder() |
| Page / Component layer (Page Objects, Tasks, Service Objects) | Page-specific or component-specific operations | LoginPage.submit({ email, password }), CartPage.applyCoupon(code) |
| Mechanical layer (the framework primitives) | Click, type, navigate, request | page.click(), page.fill(), request.post() |
The test body lives at the business layer. The test never reaches into the mechanical layer directly. If it does, the team has either no abstraction or the abstraction leaks.
Bad vs good (cross-framework)
Bad (test body at the mechanical layer):
test('places an order', async ({ page }) => {
await page.goto('/login');
await page.fill('#email', '[email protected]');
await page.fill('#password', 'pass');
await page.click('#submit');
await page.goto('/product/sku-001');
await page.click('button.add-to-cart');
await page.goto('/checkout');
await page.fill('#shipping-address', '123 Main St');
await page.click('#place-order');
await expect(page.locator('.confirmation')).toBeVisible();
});Good (test body at the business layer):
test('places an order', async ({ customer, cart, checkout }) => {
await customer.signsIn();
await cart.addsItem('sku-001');
await checkout.placesOrderWithDefaultShipping();
await expect(checkout.confirmation).toBeVisible();
});The mechanics live in the Page Object / Task / fixture. The test body reads as the specification, not the implementation.
When to keep the test mechanical
Some tests legitimately operate at the mechanical layer - testing the mechanical surface itself:
For these, the test body at the mechanical layer is correct. Tag them (@a11y, @visual) so reviewers don't refactor them by mistake.
Pattern 4 - Step extraction rules
When should a step move out of the test body?
| Heuristic | Action |
|---|---|
| The step appears in 3+ tests | Extract to a Page Object method / Task / fixture |
| The step is mechanical (click, fill, navigate) but the test isn't about that mechanic | Extract |
| The step is business-meaningful and only used here | Keep inline; name it well |
| The step is 5+ lines of mechanical operations | Always extract |
| The step requires explanatory comments | The comment is a smell; the abstraction is missing - extract |
| The step is the first thing in 80% of tests | Extract to a fixture (per-test or per-describe setup) |
The "rule of three" for extraction
Per Refactoring (Fowler 1999, 2nd ed. 2018) (opens in new window), duplicate code is acceptable at first occurrence (write it inline). At the second occurrence, note the duplication. At the third occurrence, extract.
This applies to test steps: don't pre-extract on the first test ("we might need this later" is YAGNI). Extract when the third test needs the same step.
Anti-patterns
| Anti-pattern | Why it fails |
|---|---|
| Extracting every step on the first test | YAGNI; the abstraction doesn't match real usage |
| Never extracting (every test is 30 lines of mechanics) | Mechanical leakage; brittle |
Extracting to a helper that doesn't belong to a layer (helpers/random-stuff.ts) | Sprawl; the team can't find the helper |
| Extracted helper that takes 8 parameters | The helper is doing too much; split it |
Helper that hides Act vs Arrange (named setupAndDoThing()) | Test reads as if it's doing one thing when it does two |
Pattern 5 - AAA / Given-When-Then mapping
Two equivalent step-grouping vocabularies (Arrange/Act/Assert = Given/When/Then). The team picks one, uses it consistently, and keeps the phases visually separable (blank-line or comment separation). Full mapping table, when-AAA-vs-when-Given-When-Then guidance, and the phase-separation rule with code: references/step-grouping-and-phrasing.md.
Pattern 6 - Declarative vs imperative step phrasing
Even at the business layer, prefer declarative phrasing ("the customer signs in") over imperative ("enters email, enters password, clicks submit"). The test: would the wording need to change if the implementation changed? If yes, rewrite declaratively. Full table, when-imperative-is-correct cases, and anti-patterns: references/step-grouping-and-phrasing.md.
Pattern 7 - Step naming
Per Roy Osherove's The Art of Unit Testing (2013) (opens in new window), test names follow the pattern <system_under_test>_<scenario>_<expected_outcome>. Step naming (within the test) follows similar discipline:
| Smell | Refactor |
|---|---|
await doThing() | Name what the thing is: await customer.signsIn() |
await test1() | Helpers don't get numeric names; describe what they do |
await x = await getX() | Single-letter variables hide what's being created |
await page.click('#submit') | Wrap in a named Page Object method: await loginPage.submit() |
The reader test
A reviewer should be able to read the test body aloud and have it sound like a specification:
"A customer signs in. They add SKU-001 to their cart. They place an order with default shipping. The order is confirmed."
If reading aloud doesn't produce a specification - if it produces "click, type, click, click, navigate, click, expect-truthy" - the steps are at the wrong abstraction.
Worked example
A reviewer receives an E2E spec places an order written entirely at the mechanical layer (the "Bad" block under Pattern 3): goto login, fill email, fill password, click submit, goto product, click add-to-cart, goto checkout, fill address, click place-order, then assert the confirmation.
Applying the patterns:
Result - the test body now reads as a specification: it collapses to exactly the "Good" block under Pattern 3 (four business-layer steps, no mechanical leakage).
The reader test (Pattern 7) now passes: "A customer signs in, adds SKU-001 to the cart, places an order with default shipping, and the order is confirmed."
Pattern selection guide
| Scenario | Pattern |
|---|---|
| Test reads as 20 clicks-and-types | Extract to Page Objects / Tasks (Pattern 4) and rewrite at business layer (Pattern 3) |
| Test does Arrange and Act in the same line | Split (Pattern 2 single-purpose rule) |
| Three tests share the same first 4 lines | Extract to fixture / helper (Pattern 4 rule of three) |
| Test phases not visually separable | Add blank lines or AAA comments (Pattern 5) |
| Step name reads as implementation detail | Rewrite declaratively (Pattern 6) |
| Step requires explanatory comment to understand | The abstraction is missing - extract (Pattern 4) |
| Test mixes business and mechanical vocabulary | Pick one layer per test (Pattern 3) |
Cross-cutting anti-patterns
| Anti-pattern | Why it fails |
|---|---|
| Tests with 30+ steps (one logical thing per "step" but the whole test does 10 logical things) | Single-responsibility violation; split into multiple tests |
Helpers that wrap one framework call (async function click(sel) { await page.click(sel); }) | Wrapping for the sake of wrapping; adds no abstraction |
Step extracted into a helper but the helper takes a boolean flag that branches behavior | Two helpers masquerading as one |
| Step that does retry / wait / fallback inside | Hides flakiness; the test passes when it should fail |
| Tests that read top-down look fine but the helpers contain hidden assertions | The test seems to assert one thing; actually asserts more (or different things) |
| Tests where the assertion is inside the Page Object method | Violates object-model-patterns no-assertion rule |
Hand-off targets
References
Step grouping and phrasing
View source (opens in new window)Step grouping and phrasing
Reference detail for test-step-design-patterns Patterns 5 and 6: how test steps are grouped into phases and how they are phrased. The SKILL.md spine summarizes and links here; this file holds the full tables, guidance, and code.
AAA / Given-When-Then mapping (Pattern 5)
Two equivalent step-grouping vocabularies. Same content, different name conventions.
| AAA (xUnit tradition) | Given-When-Then (BDD tradition) |
|---|---|
| Arrange | Given |
| Act | When |
| Assert | Then |
| (Annotate) | (no equivalent; goes in step comments) |
Both express the same three phases. The team picks one and uses it consistently. Don't mix: if some tests use AAA comments and others use Given-When-Then helpers, the cognitive cost grows.
When AAA is the right choice
When Given-When-Then is the right choice
The phase-separation rule
Whatever vocabulary the team uses, the phases must be visually separable. The pattern from test-code-conventions §1:
test('places an order', async () => {
// Arrange
const customer = await aCustomer().withCartItem('sku-001').build();
// Act
const order = await customer.placesOrder();
// Assert
expect(order.status).toBe('confirmed');
});Or, with blank-line separation (no comments needed):
test('places an order', async () => {
const customer = await aCustomer().withCartItem('sku-001').build();
const order = await customer.placesOrder();
expect(order.status).toBe('confirmed');
});Either works; lacking either fails the §1 AAA review.
Declarative vs imperative step phrasing (Pattern 6)
Even at the business layer, two phrasings exist:
| Imperative | Declarative |
|---|---|
| "The customer enters the email and password and clicks submit" | "The customer signs in" |
| "The customer adds 5 items to the cart and proceeds to checkout" | "The customer initiates checkout with a 5-item cart" |
| "The customer is created with role=admin and org_id=42" | "An admin customer in org 42" |
The declarative form is preferred per Cucumber's better-Gherkin guidance (opens in new window) (which applies broadly, not just to Gherkin): "scenarios should describe the intended behaviour of the system, not the implementation."
The test: would the wording need to change if the implementation changed? If yes, the step is imperative; rewrite declaratively.
When imperative is correct
Anti-patterns
| Anti-pattern | Why it fails |
|---|---|
Declarative step that hides a critical mechanic (customer.signsIn() when the test is about SSO redirect) | The test no longer tests what it claims to test |
| Imperative step that re-describes the system internals | Brittle to refactors; the test breaks when the implementation changes for unrelated reasons |
| Mixing imperative and declarative within one test | Reader can't tell what abstraction level they're operating at |
Related skills
object-model-patterns
Pure reference catalog of the canonical object-model architecture patterns for test automation frameworks - Page Object Model (Fowler), Screenplay (Marcano/Palmer/Hill), Component Object, App Actions (Cypress idiom), Service Object, Repository, and Screen Object (the desktop/mobile sibling of Page Object covering Windows UIA, macOS XCTest, Linux AT-SPI, Appium / Espresso) - each with its canonical citation, when-to-use rules, refuse-to-mix anti-patterns, and a worked example. This is the architecture-tier reference - what each pattern *is* - not file-level style rules and not tool-specific configuration. Use when designing, reviewing, or migrating a test framework's object-model architecture.
test-code-conventions
Pure-reference catalog of test-code conventions: AAA structure (Arrange / Act / Assert), per-test single-responsibility, descriptive naming (`{sut}_{scenario}_{expected}`), assertion specificity, mocking rationale (state vs behavior, fake vs mock), fixture-coupling rules, and the magic-number / hard-coded-string anti-patterns; the E2E selector-priority and web-first-assertion conventions live in references/. Use as the shared rule book a test-code review cites back to, or as onboarding for what makes a test code-reviewable; to score a test's quality on weighted axes use test-design-scorecard, and for setup/teardown isolation specifically use test-isolation-patterns.
test-design-scorecard
Scores test files 1 to 5 on six design axes (AAA phase separation, single-responsibility, naming, fixture coupling, magic literals, setup time) using explicit per-level anchors that settle what separates a 2 from a 4, then turns the scores into growth-framed feedback and a per-author trend report; the per-PR and per-author rollup examples and trend-reporting conventions live in references/. Owns the scoring and the write-up only: the conventions being scored live in a separate conventions catalog such as `test-code-conventions`, and block-or-approve gating belongs to an adversarial review. Use when a test diff needs a graded coaching read rather than a merge verdict: onboarding a new engineer, a team deliberately ramping up test discipline, or a quarterly per-author trend where the output is a conversation, not a gate.
test-framework-architecture-audit
Audits an existing test automation framework across eight architecture-tier axes and bands each one PASS, WARN, or FAIL: page-object coverage and purity, base-class inheritance depth, fixture scope and coupling, helper sprawl, naming-convention drift, retry and wait consistency, documented-versus-actual convention drift, and CI integration health. Carries the numeric cut behind every band and labels which cuts are practitioner conventions rather than published standards. Measures the framework's own structure (page objects, base classes, fixtures, helpers, conventions), not the suite's tier mix or flake rate, and not the design of a framework that does not exist yet. Use when a test framework has grown for a release or more without structural review, before a major refactor, or when a team suspects its written test conventions no longer match what the code actually does.
test-framework-blueprint
Build-an-X workflow that takes an SDET from no test suite to a complete framework design in seven steps - inventory the SUT, choose runner + language, directory layout + fixture architecture, object-model decision, test data + mocking wiring, reporting + CI integration, conventions doc + review gates - producing a written framework blueprint (directory tree, fixture list, chosen patterns, CI matrix) plus an implementation order. This is the whole-framework design workflow - not the Step 2 runner-choice decision on its own, not the Step 4 object-model pattern catalog it defers to, and not the scaffolder that generates the harness skeleton once the blueprint exists. Use when designing a test automation framework from scratch or re-architecting one that grew organically.
test-isolation-patterns
Pure reference catalog of test-isolation and fixture-lifecycle patterns - the four-phase test pattern (Meszaros), fixture scope (per-test / per-describe / shared / global), the Fresh-Fixture vs Shared-Fixture trade-off (Fowler), parallel-safety patterns, and cleanup discipline (afterEach / afterAll / tagged-cleanup), plus a pattern-selection guide and a worked leaking-state diagnosis. The database-isolation strategies (transaction-rollback / database-per-worker / template-database) and network / external-service stubbing live in references/. This is the architecture-tier reference, not a file-level fixture-coupling style rule. Use when designing fixture scope and isolation strategy, auditing fixture coupling or retry/wait policy, or moving a suite to parallel execution.
test-suite-health-audit
Measures an existing test suite's current state on four axes: per-file tier classification (unit / integration / E2E, first match wins), pyramid ratio against whatever target the team already committed to, per-layer flake rate, and defects-caught-per-run-minute ROI per tier, then reduces them to one categorical verdict (Healthy, Needs pruning, Needs refactor, Cannot assess). Reports severity against a target ratio but never prescribes one: choosing the target unit:integration:E2E mix and the rebalancing plan belongs to a pyramid-balancing capability such as `test-pyramid-balancer`. Use when a suite has grown for a year or more without review and someone needs a defensible read on whether it is healthy, over-grown, or structurally inverted before deciding what to delete or rewrite.