test-framework-blueprint
Build-an-X workflow that takes an SDET from no test suite to a complete framework design in seven steps - inventory the SUT, choose runner + language, directory layout + fixture architecture, object-model decision, test data + mocking wiring, reporting + CI integration, conventions doc + review gates - producing a written framework blueprint (directory tree, fixture list, chosen patterns, CI matrix) plus an implementation order. This is the whole-framework design workflow - not the Step 2 runner-choice decision on its own, not the Step 4 object-model pattern catalog it defers to, and not the scaffolder that generates the harness skeleton once the blueprint exists. Use when designing a test automation framework from scratch or re-architecting one that grew organically.
Install with skills.sh (any agent)
npx skills add testland/qa --skill test-framework-blueprinttest-framework-blueprint
Overview
This skill is a build-an-X workflow: it walks an SDET through designing a test automation framework end to end and ends with two artifacts the team can act on:
It is the connective tissue between the pattern catalogs and the scaffolder. The catalogs (object-model-patterns, test-isolation-patterns, test-step-design-patterns, test-data-patterns in qa-test-data) say what each pattern IS; the scaffolder emits a skeleton once decisions are made. Neither walks the decisions in order. This skill does.
When to use
Do not use this skill to:
Step 1 - Inventory the system under test
Before any tool is named, record four facts about the SUT. Every later decision keys off them.
| Inventory item | Questions to answer |
|---|---|
| App stack | Languages, frameworks, persistence (e.g. Node + React + Postgres). Which external services does it call (payments, email, auth provider)? |
| Deployment shape | Monolith / services / serverless? Can a full stack run locally (compose file, dev server) or only in a shared environment? |
| Change shape | Where do PRs land - one monorepo, or per-service repos? Do most changes touch the API, the UI, or both? The layer that changes most needs the fastest feedback. |
| Team skills | What languages do the engineers writing and maintaining tests already know? Per framework-choice-advisor Step 1, the framework-language mismatch is the #1 maintenance cost. |
Decision output: a coverage-layers table stating which layers get automated coverage in this framework and which are explicitly out of scope (already covered elsewhere, or deferred). Example shape:
| Layer | In this framework? | Rationale |
|---|---|---|
| Unit | No | Lives in each package, owned by devs |
| API integration | Yes | Most PRs touch the API |
| Web E2E | Yes (thin) | Critical paths only |
| Contract | Deferred | Single team owns both sides today |
Step 2 - Choose runner + language
Two criteria dominate; everything else is tie-breaking:
For the full trade-off matrices (cross-browser, mobile, parallelization, ecosystem, hire-ability) use framework-choice-advisor; when a repo already has an E2E convention in package.json, prefer continuing with it unless there is a reason to switch.
Decision output: one runner + one language, with the rejected alternatives and the reason recorded in the blueprint (the rejection rationale is what stops the debate from reopening every quarter).
Step 3 - Directory layout + fixture architecture
Layout (worked stack: Playwright + TypeScript)
tests/
e2e/ # browser tests, grouped by user journey
invoicing/
auth/
api/ # request-fixture tests, grouped by resource
fixtures/
db.ts # worker-scoped database fixtures
auth.ts # test-scoped authenticated-session fixtures
index.ts # merged export the specs import
pages/ # object model (Step 4)
builders/ # test-data builders (Step 5)
playwright.config.tsRules of thumb: group specs by user-facing domain (not by page or by developer); keep fixtures in their own modules per concern; specs import one merged test object, never raw @playwright/test.
Fixture scoping decisions
Playwright fixtures establish each test's environment and are set up on-demand (only the ones a test needs), and Playwright offers exactly two scopes (Playwright test-fixtures docs (opens in new window)):
| Scope | Lifecycle | Blueprint use |
|---|---|---|
| Test (default) | Set up before and torn down after each test | Anything a test mutates: pages, sessions, seeded records |
| Worker | Set up once per worker process, reused across test files whose worker fixtures match | Expensive shared infrastructure tests only read, or per-worker isolated stores (database-per-worker) |
The blueprint records, per fixture: name, scope, what it provides, and whether tests mutate it. The single rule from test-isolation-patterns Pattern 2 applies verbatim: never share mutable fixtures across tests.
The Playwright fixture mechanics to standardize (test.extend, auto, mergeTests, option fixtures) and the pytest mapping for Python teams live in references/fixture-mechanics.md.
Step 4 - Object-model decision
Pick exactly one object-model pattern; mixing two in one codebase is the top cross-cutting anti-pattern in object-model-patterns. The short decision rule (full when-to-use rules, canonical citations, and per-pattern anti-patterns live in that catalog - defer to it, do not restate it):
| Choose | When |
|---|---|
| Page Object Model | Page-oriented SUT, 3+ engineers, classic runner; the default |
| + Component Objects | Component-architected frontend (React/Vue) with shared nav/modals; refinement of POM, not a competitor |
| Screenplay | Suite will exceed ~200 tests or has multiple actor types sharing interactions |
| App Actions | Cypress idiom; SUT exposes a programmatic state API and setup dominates runtime |
Decision output: the pattern name + the catalog link, plus the deferral rule: do not build the object-model layer until roughly 10 tests exist (see Anti-patterns); the blueprint names the pattern, the implementation order delays it.
Step 5 - Test data + mocking wiring
Three sub-decisions, each deferring to its own deeper tool:
Decision output: a one-line entry per external dependency (real / stubbed / contract-tested) and the seed + builder choices.
Step 6 - Reporting + CI integration
Decision output: the CI matrix table (trigger × suite × shards × retry).
Step 7 - Conventions doc + review gates
The blueprint ends as a living docs/test-conventions.md in the repo. It contains, at minimum: the coverage-layers table (Step 1), the runner decision with rejected alternatives (Step 2), the fixture table and scoping rules (Step 3), the object-model pattern name + catalog link (Step 4), the real/stubbed dependency list (Step 5), and the CI matrix (Step 6).
A conventions doc nobody enforces drifts. Wire the enforcement loop from this plugin:
Worked example
A full worked example - "Ledgerly", a B2B invoicing web app - walking all seven steps and ending with the implementation order lives in references/worked-example.md.
Anti-patterns
| Anti-pattern | Why it fails |
|---|---|
| Copying the framework from a previous job regardless of change shape | The old framework encoded the old SUT's inventory (Step 1); a UI-heavy framework on an API-heavy product tests the wrong layer slowly |
| Building abstraction layers before ~10 tests exist | Abstractions extracted from zero usage guess wrong; extract from observed duplication (the rule-of-three framing in test-step-design-patterns) |
| One mega base-class every test inherits | Depth-3+ hierarchies break unpredictably on root changes (§A2); compose fixtures instead |
| Choosing the runner before the team-skills inventory | Framework-language mismatch is the #1 maintenance cost per framework-choice-advisor |
| Designing the CI matrix for scale on day one (8 shards, 3 retries) | Retries hide flake in a young suite; shards add cost below the ci-test-job-conventions §1 runtime thresholds |
| Skipping the written blueprint ("the code is the doc") | Documented-vs-actual drift becomes undetectable; the drift check needs a documented side to compare against |
Limitations
References
Fixture mechanics - Playwright and pytest
View source (opens in new window)Fixture mechanics - Playwright and pytest
Deep mechanics for Step 3 of test-framework-blueprint. The blueprint decisions (which scope, what each fixture provides, whether tests mutate it) live in SKILL.md; this file carries the framework-specific wiring.
Playwright mechanics to standardize in the conventions doc
All per the test-fixtures docs (opens in new window):
The pytest equivalent
If Step 2 chose Python, the same architecture maps onto pytest fixtures. Per the pytest fixtures how-to (opens in new window), tests request fixtures by declaring them as arguments; available scopes are function (the default), class, module, package, and session, where scope controls destruction (a function-scoped fixture "is destroyed at the end of the test"; a session-scoped one at the end of the test session). The fixtures/index.ts merged-export convention becomes conftest.py (fixtures there are accessible to "tests from multiple test modules in the directory"), teardown code goes after yield, and { auto: true } becomes @pytest.fixture(autouse=True). Playwright's worker scope has no direct pytest twin; session scope plus per-worker IDs (e.g. pytest-xdist worker id) fills the same database-per-worker role.
Worked example - "Ledgerly", a B2B invoicing web app
View source (opens in new window)Worked example - "Ledgerly", a B2B invoicing web app
A full walk of the seven-step test-framework-blueprint on one product, ending with the implementation order. The steps themselves are in SKILL.md.
The product: Ledgerly lets accountants create, send, and reconcile invoices. Node/Express API + Postgres, React (Vite) frontend, Stripe for payments, one monorepo, full stack runs locally via Docker Compose. Team of four: three TypeScript-fluent product engineers, one SDET. No test suite beyond scattered React unit tests.
Step 1 - Inventory. Change shape: 70% of PRs touch the API or API + UI together; UI-only PRs are rare. External dependency: Stripe. Coverage decision: API integration layer (primary), thin web E2E for the five critical journeys (create invoice, send, pay via Stripe redirect, reconcile, export), unit stays with the packages, contract testing deferred (one team owns both sides).
Step 2 - Runner. Team language is TypeScript; one runner can cover both chosen layers, so Playwright Test takes the API tier (built-in request fixture, isolated per test per the test-fixtures docs (opens in new window)) and the E2E tier. Rejected: Cypress (would still need a second runner for the API tier), Jest + supertest (second runner for E2E).
Step 3 - Layout + fixtures.
tests/
api/
invoices/
payments/
e2e/
invoicing/
reconciliation/
fixtures/
db.ts
auth.ts
stripe-stub.ts
index.ts
pages/
builders/
playwright.config.ts
docs/test-conventions.mdFixture list (the blueprint's core table):
| Fixture | Scope | Provides | Mutated by tests? |
|---|---|---|---|
workerDb | worker | Database-per-worker (ledgerly_test_w${workerIndex}), migrated once per worker | yes, via test-scoped children |
seededAccount | test | One fresh accountant account + org in the worker DB | yes |
authedPage | test | page logged in as seededAccount | yes |
api | test | request context pre-authenticated against the API | yes |
stripeStub | worker, auto: true | Asserts the Stripe stub container is up; fails fast if a test would hit real Stripe | no |
db.ts, auth.ts, and stripe-stub.ts are separate test.extend() modules combined with mergeTests() into fixtures/index.ts, per the test-fixtures docs (opens in new window).
Step 4 - Object model. POM + Component Objects: the SUT is a React SPA with page-shaped flows and a shared nav/sidebar, suite projected well under 200 tests, so Screenplay overhead is not justified per the selection matrix in object-model-patterns. App Actions rejected (not Cypress; no exposed store API). POM construction is deferred in the implementation order until ~10 specs exist.
Step 5 - Data + mocking. Seed: empty DB + per-test creation through one invoiceBuilder and one accountBuilder (Test Data Builder per test-data-patterns in qa-test-data); no shared seed set yet. Isolation: database-per-worker (test-isolation-patterns Pattern 4b) because invoice tests are mutation-heavy. Dependencies: Postgres real (in compose), Stripe stubbed by a stub container in docker-compose.test.yml (stub tooling chosen per the detected runtime), email captured by a local SMTP sink.
Step 6 - CI matrix. Reporters per the test-reporters docs (opens in new window): junit + blob on CI, html locally.
| Trigger | Suite | Shards | Retry |
|---|---|---|---|
| Per-PR | tests/api + tests/e2e/invoicing (smoke) | none (est. < 5 min) | 0 |
| Merge to main | full tests/ | none until runtime > 10 min, then 2-4 per ci-test-job-conventions §1 | 1 on runner failure only |
| Nightly | full tests/ against staging | as merge | 1, failures auto-filed |
Step 7 - Conventions + gates. docs/test-conventions.md holds all six decision outputs above. A test-code critic wired as a PR check on tests/**; a framework-architecture audit scheduled quarterly.
Implementation order (each step waits on the previous):
Related skills
object-model-patterns
Pure reference catalog of the canonical object-model architecture patterns for test automation frameworks - Page Object Model (Fowler), Screenplay (Marcano/Palmer/Hill), Component Object, App Actions (Cypress idiom), Service Object, Repository, and Screen Object (the desktop/mobile sibling of Page Object covering Windows UIA, macOS XCTest, Linux AT-SPI, Appium / Espresso) - each with its canonical citation, when-to-use rules, refuse-to-mix anti-patterns, and a worked example. This is the architecture-tier reference - what each pattern *is* - not file-level style rules and not tool-specific configuration. Use when designing, reviewing, or migrating a test framework's object-model architecture.
test-code-conventions
Pure-reference catalog of test-code conventions: AAA structure (Arrange / Act / Assert), per-test single-responsibility, descriptive naming (`{sut}_{scenario}_{expected}`), assertion specificity, mocking rationale (state vs behavior, fake vs mock), fixture-coupling rules, and the magic-number / hard-coded-string anti-patterns; the E2E selector-priority and web-first-assertion conventions live in references/. Use as the shared rule book a test-code review cites back to, or as onboarding for what makes a test code-reviewable; to score a test's quality on weighted axes use test-design-scorecard, and for setup/teardown isolation specifically use test-isolation-patterns.
test-design-scorecard
Scores test files 1 to 5 on six design axes (AAA phase separation, single-responsibility, naming, fixture coupling, magic literals, setup time) using explicit per-level anchors that settle what separates a 2 from a 4, then turns the scores into growth-framed feedback and a per-author trend report; the per-PR and per-author rollup examples and trend-reporting conventions live in references/. Owns the scoring and the write-up only: the conventions being scored live in a separate conventions catalog such as `test-code-conventions`, and block-or-approve gating belongs to an adversarial review. Use when a test diff needs a graded coaching read rather than a merge verdict: onboarding a new engineer, a team deliberately ramping up test discipline, or a quarterly per-author trend where the output is a conversation, not a gate.
test-framework-architecture-audit
Audits an existing test automation framework across eight architecture-tier axes and bands each one PASS, WARN, or FAIL: page-object coverage and purity, base-class inheritance depth, fixture scope and coupling, helper sprawl, naming-convention drift, retry and wait consistency, documented-versus-actual convention drift, and CI integration health. Carries the numeric cut behind every band and labels which cuts are practitioner conventions rather than published standards. Measures the framework's own structure (page objects, base classes, fixtures, helpers, conventions), not the suite's tier mix or flake rate, and not the design of a framework that does not exist yet. Use when a test framework has grown for a release or more without structural review, before a major refactor, or when a team suspects its written test conventions no longer match what the code actually does.
test-isolation-patterns
Pure reference catalog of test-isolation and fixture-lifecycle patterns - the four-phase test pattern (Meszaros), fixture scope (per-test / per-describe / shared / global), the Fresh-Fixture vs Shared-Fixture trade-off (Fowler), parallel-safety patterns, and cleanup discipline (afterEach / afterAll / tagged-cleanup), plus a pattern-selection guide and a worked leaking-state diagnosis. The database-isolation strategies (transaction-rollback / database-per-worker / template-database) and network / external-service stubbing live in references/. This is the architecture-tier reference, not a file-level fixture-coupling style rule. Use when designing fixture scope and isolation strategy, auditing fixture coupling or retry/wait policy, or moving a suite to parallel execution.
test-step-design-patterns
Pure reference catalog of test-step design patterns at the architecture tier - step granularity (one logical action per step), abstraction layers (mechanical → page → business), step extraction rules (when to inline / when to extract to a helper / when to extract to a Page Object method), the declarative-vs-imperative phrasing rule, FIRST principles (Fast / Independent / Repeatable / Self-validating / Timely), and the AAA / Given-When-Then mapping. This is the cross-framework architecture-tier reference for what a step IS, when it should exist, and where it should live - not file-level AAA style rules and not Gherkin-specific translation. Use when designing or reviewing the step layer of a test framework - for example when writing or reviewing E2E or integration tests, when the step count per test is high, or when refactoring recorded or codegen test output into readable steps.
test-suite-health-audit
Measures an existing test suite's current state on four axes: per-file tier classification (unit / integration / E2E, first match wins), pyramid ratio against whatever target the team already committed to, per-layer flake rate, and defects-caught-per-run-minute ROI per tier, then reduces them to one categorical verdict (Healthy, Needs pruning, Needs refactor, Cannot assess). Reports severity against a target ratio but never prescribes one: choosing the target unit:integration:E2E mix and the rebalancing plan belongs to a pyramid-balancing capability such as `test-pyramid-balancer`. Use when a suite has grown for a year or more without review and someone needs a defensible read on whether it is healthy, over-grown, or structurally inverted before deciding what to delete or rewrite.