Testland
Browse all skills & agents

qa-vendor-evaluator

Build-an-X workflow that produces a side-by-side **commercial-vendor** evaluation matrix for QA tools - test-management platforms (TestRail / Qase / Xray / Zephyr / TestCollab), no-code platforms (mabl / Testim / Functionize / TestSigma / Reflect), visual regression services (Applitools / Percy / Chromatic), and commercial AI copilots - scoring each on capability fit, cost model, integration depth, vendor lock-in risk, exit cost, contractual posture, and customer-reference data. Scoped to commercial procurement - contract, lock-in, and exit-cost axes - not to choosing an open-source code-first framework on architectural fit. Use for commercial procurement decisions only - refuses to recommend a winner; the team owns the procurement choice.

Install with skills.sh (any agent)

npx skills add testland/qa --skill qa-vendor-evaluator
View source

qa-vendor-evaluator

Overview

QA managers procure commercial tools every 12 - 24 months: a new test-management platform, a no-code automation vendor, a visual-regression service, an AI copilot tier. The Capgemini World Quality Report 2025-26 (opens in new window) identifies integration friction (37% of teams) as the dominant blocker for AI-in-testing adoption - the proximate failure mode is teams that adopted a vendor without scoring integration cost in advance. This skill produces the structured side-by-side comparison the manager carries into the procurement decision, with every score citing its source.

This is decision-support, not recommendation. The skill refuses to pick a winner - the team owns the procurement choice. The output is the evidence pack: capability matrix, cost-model breakdown, integration / lock-in / exit-cost analysis, and contractual posture per vendor. The manager (or a procurement-committee) makes the call against the team's NFR priorities.

When to use

  • A 12 - 24 month procurement cycle is approaching for a commercial QA tool.
  • The team is evaluating a switch from one vendor to another (mid-contract renewal, post-acquisition consolidation).
  • A new tool category is being adopted (the team has no incumbent - e.g., first visual-regression service).
  • A vendor is being added to a multi-vendor stack and the team wants to ensure no overlap with incumbents.
  • An RFP / RFI / procurement-committee process needs a structured evaluation artifact.

Do not use this skill when:

  • The decision is between open-source code-first frameworks (Playwright vs Cypress vs Selenium) - use framework-choice-advisor. Different axis entirely (architecture, not procurement; no contract, lock-in, or exit-cost dimensions).
  • Only one vendor is being evaluated - comparison requires ≥2 candidates. (For a single-vendor go/no-go, use the team's standard procurement checklist; this skill needs comparison anchors.)
  • The team has no defined NFR priorities - the matrix cannot score "capability fit" without knowing what the team needs the tool to do.

Step 1 - Capture the inputs

Required:

InputNotes
≥2 vendor candidatesThe vendors being compared. Halts on 1; recommends ≥3 for a healthy comparison (avoids the two-choice false-binary).
Team profileTeam size, existing stack (CI, test framework, observability, tracker), seat / volume profile, geography, regulated-industry flag (if any).
NFR prioritiesOrdered list of what the team needs most: capability fit / cost / integration / lock-in / contract / support / data-residency. The order matters - the matrix's weighted score depends on it.
Time horizon12 / 24 / 36 month decision window. Drives the lock-in and exit-cost scoring (longer horizon = more weight on lock-in).
Per-vendor dataPricing page, feature page, integration docs, customer-reference reviews (Gartner Peer Insights, G2, Capterra), vendor-published case studies (tagged as vendor-data).

The skill halts with INSUFFICIENT_INPUT if any required input is missing.

Step 2 - Score on the seven procurement axes

Per the canonical commercial-procurement framework, seven axes drive the decision. Score each vendor on each axis, with every score citing its source:

  • A1 Capability fit - feature coverage vs the team's NFR list, plus documented limits.
  • A2 Cost model - pricing model, year-1/2/3 cost at the team's scale, hidden costs.
  • A3 Integration depth - CI, tracker, observability, test-framework, SSO, reverse data flow; often under-weighted despite integration friction blocking 37% of teams (Capgemini WQR 2025-26 (opens in new window)). Score each integration native (1.0) / API-buildable (0.7) / community-plugin (0.5) / not available (0.0).
  • A4 Lock-in risk - format portability, artifact export, data residency, migration path.
  • A5 Exit cost - test re-authoring, history portability, re-training, egress fees, early-termination.
  • A6 Contractual posture - SLA, support, security audits, on-prem, DPA/BAA, sub-processors.
  • A7 Customer-reference data - Gartner / G2 ratings, practitioner signal, community anecdote (tagged, not equal-weighted with surveyed data).

The full per-sub-axis rubric for all seven axes is in references/scoring-rubric.md.

Step 3 - Emit the comparison matrix

Emit a single markdown document: the vendors compared, the team profile, the ordered NFR priorities, a per-axis matrix (A1-A7), a weighted-score table using the NFR order as weights, an explicit "What this skill did NOT do" disclaimer (does not pick the winner, negotiate, validate vendor claims, or replace a reference call), and an evidence appendix tracing every cell to a source. The full worked comparison document is in references/example-matrix.md.

Step 4 - Hand off to procurement / decision committee

The matrix is the input to the decision, not the decision itself. Downstream:

  1. Vendor sales call to re-verify pricing, SLA, and roadmap claims.
  2. Customer reference call to 2 - 3 named customers per finalist.
  3. Security review of SOC 2 / ISO 27001 reports.
  4. Pilot / POC on the top-2 candidates (typically 30 - 60 day evaluation).
  5. Procurement-committee decision with the matrix as the structured evidence pack.

Anti-patterns

Anti-patternWhy it failsFix
Scoring "capability fit" without the team's NFR prioritiesThe matrix favours feature-richest vendor regardless of fitStep 1 NFR priorities are mandatory inputs
Equal-weight matrixTreats integration, cost, and capability as equally important - almost never trueStep 3 weighted score per team's NFR order
Picking the vendor by gut from the matrixPer the Capgemini WQR (opens in new window), integration cost is under-weighted; gut decisions favour capability over integrationThe weighted-score column is the discipline; the team must justify deviations
Skipping A4 / A5 (lock-in / exit cost) because "we'll figure it out later"The dominant cost surfaces at year-2+; skipping these axes optimises for year-1 happinessThese axes are mandatory in the matrix
Treating vendor-data and Gartner / G2 data identicallyVendor-data is marketing; reviewed data is signalA7 explicitly separates them
Using the matrix as the procurement decisionProcurement requires sales / reference / security calls beyond the matrixStep 4 hand-off lists the required downstream actions
Comparing only 2 vendorsTwo-choice procurement is a false binary; the team often missed a third optionStep 1 recommends ≥3 candidates
Skipping the weighted score because "it feels mechanical"Without weighting, the matrix is decorationStep 3 weighted score is required
Auto-recommending the highest-scored vendorThe team's context (existing relationship, hire-ability, contract leverage) is outside the matrix; auto-recommend strips that contextThe "What this skill did NOT do" block explicitly disclaims the recommendation

Limitations

  • Vendor-data dominates the matrix. Pricing pages, feature lists, and case studies are vendor-published. The skill flags these explicitly; the team verifies in sales calls.
  • Customer-reference scoring is shallow. Public-review averages (Gartner, G2) are noisy. Real customer calls are the high-signal version; the matrix supplements but doesn't replace.
  • Categories drift fast. A 2026 evaluation is stale by 2028; product features, pricing, and vendor stability all change. Re-run before each renewal cycle.
  • No competitive-dynamics axis. The matrix doesn't account for vendor M&A risk (TestRail's parent Idera, mabl's growth stage, etc.). Add as an A8 if the team's horizon is >24 months.
  • Local market variation. Pricing and contract terms differ by geography; the matrix uses US-list pricing unless the team specifies otherwise.
  • No regulatory-specific guidance. Regulated-industry teams (healthcare, finance, automotive) have additional contractual axes (BAA, HIPAA, FDA Class II evidence) the skill doesn't enumerate per industry. Layer those onto A6.
  • Honest about being a draft. The matrix is the manager's starting point; the procurement committee customises weights, adds axes, and replaces data after the sales / reference / security calls.

Hand-off targets

  • Open-source framework selection (different decision)framework-choice-advisor.
  • Per-vendor integration playbooks after the vendor is pickedtestrail-integration, xray-integration, zephyr-integration, currents-integration.
  • Vendor evaluation feeding a quarterly OKR (e.g., "adopt vendor X by Q3")qa-okr-author.
  • Compliance vendor evaluation (regulated industries) → augment the matrix with the qa-compliance plugin's per-framework reference skills.

References

  • Capgemini World Quality Report 2025-26 - 37% of teams cite integration friction as the dominant AI-in-testing blocker; load-bearing for axis A3: https://www.capgemini.com/insights/research-library/world-quality-report-2025-26/
  • Gartner Peer Insights - AI-augmented software testing category: https://www.gartner.com/reviews/market/ai-augmented-software-testing-tools
  • G2 / Capterra - methodology disclosure for review-density and recency scoring (general SaaS evaluation context; not QA-specific): https://www.g2.com/about
  • ISTQB glossary - test automation framework (the open-source / commercial boundary): https://glossary.istqb.org/en_US/term/test-automation-framework
  • ISO/IEC 25010 - quality model for non-functional requirements (used in A1 capability scoring): https://en.wikipedia.org/wiki/ISO/IEC_25010
  • framework-choice-advisor - sibling reference for open-source framework selection; this skill is its commercial-procurement complement.
  • testrail-integration, xray-integration, zephyr-integration, currents-integration - per-vendor integration baselines that feed A3.
  • qa-okr-author - when the procurement outcome ladders into a quarterly OKR.

Worked example - vendor comparison matrix

View source (opens in new window)

Worked example - vendor comparison matrix

Deep reference for the qa-vendor-evaluator SKILL.md, Step 3. The full comparison document the skill emits: per-axis matrix, weighted score, the explicit "What this skill did NOT do" disclaimer, and the evidence appendix. Values are illustrative.

# Vendor evaluation - `<category>` - `<team>` - 2026-07

## Vendors compared

| Code | Vendor | Pricing page | Cited integration doc |
|---|---|---|---|
| V1 | TestRail (Gurock / Idera) | https://www.testrail.com/pricing/ | https://support.testrail.com/hc/en-us/articles/7077873061908 |
| V2 | Qase | https://www.qase.io/pricing/ | https://docs.qase.io/en/articles/6417206-github |
| V3 | Xray (Xpand IT, for Jira) | https://marketplace.atlassian.com/apps/1211769/xray-test-management-for-jira | https://docs.getxray.app/display/XRAYCLOUD/REST+API |

## Team profile

- Size: 12 QA engineers
- Stack: Playwright + Jest, GitHub Actions, Linear (tracker), Datadog (observability)
- Geography: distributed US + EU; data-residency: EU required
- Regulated-industry: no
- Time horizon: 24 months

## NFR priorities (manager-supplied, ordered)

1. Integration with Linear + GitHub Actions
2. Cost at year-2 (team will grow to 18 engineers)
3. Data residency (EU)
4. Test-history portability (exit-cost matters; 24-month horizon)
5. SSO (SAML / OIDC)
6. Capability fit
7. Customer-reference depth

## Per-axis matrix

### A1 - Capability fit

| Vendor | Score (0-1.0) | Strengths | Gaps |
|---|---|---|---|
| TestRail | 0.85 | Mature test-case management, custom fields, bulk import / export | API rate limits documented at 180/min - may bind at scale |
| Qase | 0.80 | Modern UI, AI-assisted case authoring | Smaller plugin ecosystem |
| Xray | 0.95 | Deep Jira integration, BDD-native | Heavyweight Jira dependency the team doesn't have |

### A2 - Cost model

| Vendor | Year-1 (12 eng) | Year-2 (18 eng) | Hidden costs |
|---|---|---|---|
| TestRail | $5,328 (12 × $37/seat/mo Professional × 12) | $7,992 | SSO, automated backups, priority support are Enterprise-tier only |
| Qase | $4,320 (12 × $30/seat/mo Business × 12) | $6,480 | None at this tier; SSO included from Business plan |
| Xray | Quote from the Marketplace listing - Xray licenses by total Jira user tier, not by tester seat | Same tier rule at 18 engineers; re-quote if the Jira tier changes | Requires Jira Software seats if not already licensed |

### A3 - Integration depth

| Vendor | CI (GitHub Actions) | Tracker (Linear) | Observability (Datadog) | Test-framework (Playwright) | SSO |
|---|---|---|---|---|---|
| TestRail | Native (1.0) | API-buildable (0.7) | Community plugin (0.5) | API + `testrail-cli` (0.9) | SAML / OIDC (1.0) |
| Qase | Native action (1.0) | Native (1.0) | Webhook (0.7) | Native @qase/playwright (1.0) | SAML (Business+) (0.8) |
| Xray | API only (0.7) | API-buildable (0.7) | None native (0.0) | xray-junit-extensions (0.9) | SAML / OIDC (1.0) |

### A4 - Vendor lock-in risk

| Vendor | Format | Export | Lock-in score |
|---|---|---|---|
| TestRail | Proprietary case format; bulk CSV export | Documented CSV / XML export, JSON via API | Moderate (0.6) - export possible, but tests need re-authoring on migration |
| Qase | YAML / JSON case format; native import / export | First-class export to JSON / YAML | Low (0.85) - portable artifacts |
| Xray | BDD-native (Gherkin), JUnit / Cucumber export | Tied to Jira issue model; export possible but harder to disentangle | Moderate-high (0.5) - Jira coupling is the lock-in axis |

### A5 - Exit cost (24-month migration scenario)

| Vendor | Test re-authoring | History portability | Total exit cost (hand-wave) |
|---|---|---|---|
| TestRail | Cases portable as CSV; ~30% needs re-authoring for new tool | History exportable via API | ~3 person-months |
| Qase | YAML / Gherkin cases mostly portable; ~10% re-authoring | Native export | ~1 person-month |
| Xray | Gherkin scenarios portable; Jira-issue history harder to extract | API export; needs custom tooling | ~4 person-months |

### A6 - Contractual posture

| Vendor | SLA | Support | Security | EU residency |
|---|---|---|---|---|
| TestRail | 99.9% (Cloud), no SLA for self-hosted | Email; phone on Enterprise | SOC 2 Type II + ISO 27001 (cite vendor security page) | EU AWS region available on Enterprise |
| Qase | 99.9% on Business+ | Email + chat; CSM on Enterprise | SOC 2 Type II (cite) | EU region available on Business |
| Xray | Bound to Jira's SLA | Email + chat | SOC 2 Type II inherited from Xpand IT | Tied to Jira region |

### A7 - Customer-reference data

| Vendor | Gartner Peer Insights | G2 (recency / density) | Practitioner-signal |
|---|---|---|---|
| TestRail | 4.4/5 (382 reviews, mostly 2023-25) | 4.2/5, 250+ reviews | Cited in Lisa Crispin's *Agile Testing Condensed* |
| Qase | 4.6/5 (180 reviews, 2024-26) | 4.7/5, 120+ reviews | Featured in TestBash 2025 case studies |
| Xray | 4.4/5 (290 reviews) | 4.4/5, 200+ reviews | Heavy enterprise adoption signal; lighter mid-market |

## Weighted score per NFR priorities

| Axis | Weight (per team NFR order) | TestRail | Qase | Xray |
|---|---|---|---|---|
| A3 Integration | 0.25 | 0.78 | 0.92 | 0.62 |
| A2 Cost | 0.20 | 0.65 | 0.95 | 0.70 |
| A6 EU residency | 0.15 | 0.80 | 0.90 | 0.70 |
| A5 Exit cost | 0.15 | 0.60 | 0.90 | 0.50 |
| A6 SSO | 0.10 | 1.00 | 0.80 | 1.00 |
| A1 Capability fit | 0.10 | 0.85 | 0.80 | 0.95 |
| A7 Customer-reference | 0.05 | 0.90 | 0.85 | 0.85 |
| **Total** | **1.00** | **0.76** | **0.89** | **0.69** |

## What this skill did NOT do

- Pick the winner. The matrix and weighted scores are the input to the procurement decision; the team owns the choice. The team may legitimately pick the lower-scored vendor for reasons outside the matrix (existing relationship, hiring-pool, founder preference).
- Negotiate the contract. Once a vendor is picked, contract terms (discount, multi-year commit, SLA tier) are a separate procurement conversation.
- Validate vendor claims. Where the matrix cites vendor-data (pricing, feature lists, case studies), the data is vendor-published and should be re-verified in a sales call before commitment.
- Replace a reference call. Customer references should be called directly, not just scored from public-review averages.

## Evidence appendix

Every cell in the matrix above traces to a source URL or cited document. The full appendix lists every source (per axis × per vendor) so the team can spot-check.

Seven-axis procurement scoring rubric

View source (opens in new window)

Seven-axis procurement scoring rubric

Deep reference for the qa-vendor-evaluator SKILL.md, Step 2. The per-sub-axis scoring rubric for each of the seven procurement axes. Score each vendor on each axis, with every score citing its source.

Axis A1 - Capability fit

How well does the vendor's feature set match the team's documented NFR priorities?

Sub-axisScoring rubric (per-vendor)
Core feature coverage% of team's required features present (cite the team's requirement list)
Advanced / aspirational featuresFeatures the team doesn't need today but might in 2 years
Documented limitsPer-account / per-test / per-user caps that may bind

Axis A2 - Cost model

How is the vendor priced and what does it cost at the team's scale?

Sub-axisScoring rubric
Pricing modelPer-seat / per-test / per-execution / flat licence / hybrid
Cost at current team sizeYear-1 cost, cited to vendor pricing page (vendor-data)
Cost at projected team sizeYear-2 and year-3 projections; the team's growth plan drives this
Hidden costsAdd-ons (parallel execution, premium support, SSO, audit log, on-prem option)
Volume-discount commitmentsAnnual commits, multi-year discounts (cite to vendor sales channel)

Axis A3 - Integration depth with existing stack

Per the Capgemini WQR 2025-26 (opens in new window) finding (37% blocked by integration friction), this axis is often under-weighted in procurement. The skill weights it explicitly.

Sub-axisScoring rubric
CI integrationNative plugin? REST API? Webhook? CLI? Cite the integration doc URL per vendor
Tracker integrationJira / Linear / GitHub Issues / Azure DevOps
Observability integrationDatadog / Grafana / New Relic / Sentry
Test-framework bindingPlaywright / Cypress / Selenium / per-language
SSO / IAMSAML / OIDC / SCIM provisioning
Reverse data flowCan the team export raw test results? In what format?

For each integration, score: native (1.0) / API-buildable (0.7) / community-plugin (0.5) / not available (0.0). Use the team's existing integration skills (testrail-integration, xray-integration, zephyr-integration, currents-integration) as the per-vendor baseline.

Axis A4 - Vendor lock-in risk

The cost of being unable to leave.

Sub-axisScoring rubric
Proprietary data formatsAre tests in a portable format (Gherkin / standard JSON / open spec) or vendor-proprietary DSL?
Test artifact portabilityCan the team export tests, results, history? In what format? Vendor-published export tools count toward portability.
Data residencyWhere is data stored? Can the team request export and deletion?
Migration pathAre there documented or community-tested migration paths off this vendor (mabl -> Playwright, Testim -> Cypress)?

Axis A5 - Exit cost

The team has decided to leave in 24 months - what does it cost?

Sub-axisScoring rubric
Test re-authoring effortIf tests are vendor-DSL-bound, the migration is "rewrite from scratch." If tests are portable (Gherkin, standard fixtures), migration is mostly mechanical.
History portabilityCan the team take its test-result history? Defect-history correlation depends on this.
Re-trainingHow long to retrain the team on the new vendor?
Data egress feesSome vendors charge for bulk data export. Cite the contract clause if applicable.
Contract early-termination costMulti-year commits often have early-termination fees.

Axis A6 - Contractual posture

Procurement / legal / security review needs structured data.

Sub-axisScoring rubric
SLA tierUptime guarantee, support response time, escalation paths
Support tierEmail-only / chat / phone / dedicated CSM
Security audit availabilitySOC 2 Type II report? ISO 27001? Penetration-test results?
On-prem / private-cloud optionFor regulated industries; cite the vendor's deployment options page
Data-processing agreementGDPR-compliant DPA available? HIPAA BAA?
Sub-processor disclosureVendor's sub-processor list (the regulated-industry view)

Axis A7 - Customer-reference data

Independent (not vendor-published) signal.

Sub-axisScoring rubric
Gartner Peer InsightsRating, review density, recency. Cite the category report URL.
G2 / CapterraRating, review density. Flag if reviews are sparse or stale (<10 reviews in last 12 months).
Practitioner blog / conference signalHas the vendor been written about by recognised practitioners (Lisa Crispin, James Whittaker, etc.)? Cite the source.
Reddit / r/QualityAssurance / Hacker NewsAnecdotal community signal - tag as such. Don't weight equally with surveyed data.

Related skills

attack-surface-test-checklist

Maps a code change to the security tests worth running against it. Classifies changed paths and file contents into nine attack surfaces (authentication, session management, input handling, file upload, deserialization, access control, API and web service, cryptography, data protection), attaches the matching OWASP ASVS 4.0.3 verification requirements, OWASP Top 10 2021 category IDs, and OWASP WSTG section numbers to each active surface, then emits a per-surface manual and automated test checklist bounded by what actually changed. Surfaces with no changed lines are excluded rather than carried as filler. Use when a pull request, release branch, or feature is about to be security tested and the team needs a targeted test list instead of a generic application-wide checklist.

code-change-shape-classifier

Classifies a code change set into four shapes (pure-logic, service-layer, ui-heavy, data-heavy) from file-path and file-content signals, computes the shape distribution over a window of git history, and attaches a relative per-layer test cost model (unit 1x, service 3x, UI 10x) so downstream planning works from one shared input. Produces the classification only: it does not prescribe a target unit:service:UI ratio, does not estimate hours, and does not select which tests to run. Use when a pull request, release branch, or epic needs its change shape labelled before test effort, pyramid balance, or coverage depth is decided.

definition-of-done

Pure-reference + checklist-generator for the team's Definition of Done (DoD) - explains the Scrum Guide's DoD definition ("a formal description of the state of the Increment when it meets the quality measures required for the product"), proposes a starter DoD with the 7-10 lines most teams need (code reviewed, unit tests, docs, AC met, deployed to staging, smoke passed, no a11y regressions, telemetry wired, observability in place), and emits a per-PR checklist a reviewer enforces. Use when the team doesn't have a DoD or wants to revise theirs.

dod-adherence-review

Audits an existing Definition of Done checklist line by line against repository evidence (review records, diffs, CI runs, coverage reports, scan output) and tags every line met, not met, or unverifiable, refusing to pass a line on self-attestation or on a claim with no matching diff. Covers the line-pattern-to-evidence mapping for the common checklist shapes (code reviewed, coverage threshold, docs updated, acceptance criteria covered, staging deploy plus smoke, no new accessibility regressions, telemetry wired), the entry-stage versus exit-stage split many teams run, the roll-up verdict rules, and the audit table that gets emitted. Does not author, revise, or soften the checklist. Use when a story or pull request is about to be marked done and a committed Definition of Done exists that nobody has actually checked the work against.

e2e-suite-budget

Caps E2E suite size by computing per-test ROI - (regressions caught × value) ÷ (runtime × flake rate × maintenance) - then ranks every end-to-end test and recommends which bottom-decile ones to retire, move to a lower layer, or fix. Use when CI is slow or E2E-dominated, flaky failures are rising, or quarterly to keep suite size within maintenance capacity. For strategic unit:service:UI layer ratios use test-pyramid-balancer, for the minimal per-deploy critical-path gate use smoke-suite-gate, and for quarantining flaky tests use flaky-test-quarantine; this prunes low-signal tests by ROI.

framework-choice-advisor

Pure reference catalog for picking a test automation framework - covers Playwright / Cypress / Selenium / WebdriverIO / Appium / Espresso / XCUITest / RestAssured / Karate / k6 / Locust with side-by-side tradeoffs on speed, cross-browser, mobile, parallelisation, language support, ecosystem maturity, CI integration; a decision tree for matching project NFRs to framework choice; and reference directory / fixture / CI layouts for the chosen stack. This is the **upstream selection step**: it decides which tool to adopt, not how to configure a tool already chosen, and not how to rebalance the unit / integration / E2E mix of an existing suite. Use when starting a new test-automation suite from scratch, before installing any tool.

heuristic-test-design-reference

Reference catalog of the four canonical heuristic test-design models - Bach's Heuristic Test Strategy Model (HTSM) with SFDPOT product elements, Whittaker's 'How to Break Software' attack patterns, Bolton's FEW HICCUPPS consistency oracles, and the ISO/IEC 25010 quality characteristics - for use when the tester has no user story, no acceptance criteria, and no documentation. This is the zero-documentation case: it does not read from a written story, and it yields test-case ideas rather than session charters. Use as the reference layer when generating coverage for a feature with no documented input.

post-mortem-author

Build-an-X workflow that produces a blameless post-mortem from an incident - captures the timeline (chronological event sequence with sources), root cause analysis (what + why, not who), impact (users / revenue / SLO debt), action items (with owners + due dates + measurable success criteria), and "what went well" (intentional). Per Google SRE: "Blameless postmortems are a tenet of SRE culture." Use after every user-visible incident, not just severe ones.

product-risk-register-builder

Build-an-X workflow that produces a product-level risk register catalogue - per-feature / per-component product risks (functionality, performance, security, usability, compatibility, reliability) that persist across releases, distinct from per-release risk matrices. Walks the author through risk identification by ISO 25010 quality characteristic, scoring per impact × likelihood, and linking each register entry to mitigations + owners + review cadence. Output is a Markdown register the team reviews quarterly and that seeds release-level risk matrices. Use for long-lived product-quality risks; complements risk-matrix for per-release risks.

project-risk-register-builder

Build-an-X workflow producing a project-level risk register - risks to project execution (schedule slippage, environment instability, people / staffing, vendor / dependency, scope creep) rather than the product itself. Walks the author through ISO 31000-aligned identification, impact × likelihood scoring, mitigation strategy (avoid / mitigate / transfer / accept), and ownership; the project manager reviews it weekly. Use for release-execution risk. For product-quality risks use product-risk-register-builder, for the per-release product risk table use risk-matrix, and to sign off accepting one specific risk use risk-acceptance-decision-author.

qa-okr-author

Build-an-X workflow that drafts a QA team's quarterly OKR set - one to three Objectives, each with 3 - 5 measurable Key Results - from the team's current state (risk matrix, defect-trend narrative, test-run history, test-pyramid balance, compliance coverage). Every numeric target cites its source artifact (e.g., a defect-trend baseline's 2026-Q1 escape rate). QA-specific by design - generic OKR generators (Tability, Asana, ClickUp) don't know test metrics; the differentiation is the domain. Produces the OKR set itself - not the test-strategy document it sits inside, and not the risk-score calibration behind the baselines. Use at the start of each quarter to draft the OKR set the manager edits and the team commits to.

risk-acceptance-decision-author

Build-an-X workflow that produces a structured risk-acceptance decision document - for risks the team has decided to accept (rather than mitigate / transfer / avoid). Walks the author through the ISO 31000 risk-acceptance criteria (rationale, sign-off, scope, review trigger, exit conditions), captures stakeholder approval, and links to the originating risk register entry. Output is a Markdown decision artefact that lives alongside the risk register and provides audit-defensible justification for the team's acceptance choice. Use when a risk register entry's Strategy column is set to Accept, or an already-accepted risk comes up for its scheduled re-review, an audit, or a post-incident look-back.

risk-coverage-mapper

Build-an-X workflow that produces a risk-to-test-coverage matrix - maps each risk in the product/release register to the tests / cases / monitoring that mitigate it. Walks the author through ingesting risks (from risk-matrix / product-risk-register-builder), inventorying test coverage (test cases via traceability-matrix-builder, automated tests via repo scan, production monitoring via observability dashboards), and computing per-risk coverage depth + identifying orphan risks (no coverage) + orphan tests (not linked to risks). Output is a Markdown matrix + executive summary. Use before a release sign-off or compliance audit, when the team must show which tests, cases, or monitors back each registered risk and which risks have nothing behind them.

risk-matrix

Produces the per-feature / per-release risk-matrix artifact itself: a structured intake (feature, category, impact 1-5 by likelihood 1-5, score), mitigations with owners and due dates, supporting both lightweight (impact by likelihood) and heavyweight (FMEA / Cost of Exposure) methods per RBT canon, output as a Markdown / spreadsheet the team reviews each sprint. Use when building the matrix artifact; to facilitate the live risk-storming meeting use risk-storming-facilitator, to calibrate scores across raters use risk-matrix-calibration, and to map the resulting risks onto test coverage use risk-coverage-mapper.

risk-matrix-calibration

Checks an already-written risk matrix against what actually happened. Maps each row's likelihood rating to observed defect density, test failure rate and code churn, maps its impact rating to the severity mix and escape rate, then classifies each row as over-stated, under-stated, in-agreement, or not calibrated, using stated reporting thresholds so small differences are not treated as findings. Every proposed rating change carries the observation that produced it, and every proposal is handed to the matrix owner rather than applied. Also emits candidate new entries for areas that show up in defect data but have no row. Owns calibration only: choosing a scoring methodology, designing the matrix structure, picking risk categories, mapping risks to test types, FMEA scoring, review cadence and file storage are all out of scope. Use when a matrix has been driving test decisions for at least three releases and nobody has yet checked whether its ratings match the defects, escapes and incidents that followed.

risk-storming-facilitator

Reference guide for planning and facilitating a risk-storming session yourself - meeting structure, participant roster, per-category brainstorm prompts (categories from risk-matrix), affinity grouping, impact by likelihood scoring, and mitigation assignment. Static reference only, not an active runner that writes the matrix file. Use to learn or teach the facilitation pattern, or to run a feature-kickoff session without agent assistance. For the matrix artifact itself use risk-matrix, to calibrate its ratings against real defect data use risk-matrix-calibration, and to map the resulting risks onto test coverage use risk-coverage-mapper.

smoke-suite-gate

Build-an-X workflow for a critical-path smoke suite that runs in <5 minutes - picks the 5-15 highest-business-value journeys (login, hero flow, checkout, payment, primary read), implements as fast E2E or API tests, gates per-deploy, retries on transient failures with quarantine. Use as the canary-precursor or per-deploy verification gate; the team's "if this fails, the build can't proceed" floor.

tdd-stuck-pattern-resolver

Pattern catalog for "I can't write the test first" moments - recognizes common testability blockers (singletons / static dependencies, network in constructors, time / random as hidden inputs, deeply nested construction, untestable boundaries) and proposes the refactor that unblocks TDD (extract interface, dependency injection, seam, ports-and-adapters). Use as TDD coaching when an engineer is stuck on a class of code. For a catalog of what-to-test heuristics with no story use heuristic-test-design-reference, to label a change's shape before planning test effort use code-change-shape-classifier, and for conventions on writing the test well once the code is testable use test-code-conventions.

test-case-from-live-feature

Build-an-X workflow that produces a test-case matrix from a **live, undocumented feature** - running app at a URL, screen recording, screenshot, or verbal brief - by combining structured exploration (Playwright trace / DevTools / accessibility tree) with the heuristic models in `heuristic-test-design-reference` (SFDPOT, Whittaker attacks, FEW HICCUPPS, ISO 25010). Output is a structured case matrix, not an exploratory session charter. Use when there is no story, no AC, and no documentation - only a live feature.

test-case-ideation-from-story

Takes a user story or feature spec and emits a markdown test-case matrix - one row per case (id, title, precondition, steps, expected, tier) covering happy path, alternate paths, boundaries, and negative paths - before any test code is written. Output is the human-reviewable matrix that goes into TestRail / Qase / Xray. Emits the human-reviewable case matrix itself - not Gherkin scenarios written against locked acceptance criteria, and not executable test code. Use as the first artifact a manual tester or three-amigos session produces from a story, ahead of automation.

test-effort-estimation

Turns a list of testable areas plus a change-shape distribution into a PERT three-point test effort estimate, reporting every row as a range around the expected value rather than a single number, requiring a named assumptions ledger across six mandatory categories, and recommending a per-layer ownership split across developer, automation, and exploratory roles. Owns the hours and the ownership recommendation only: it consumes a change-shape distribution rather than producing one, and it does not choose which tests to run or how deep coverage should go. Use when an epic or release has been broken into testable areas and someone is about to commit test capacity for a sprint.

test-pyramid-balancer

Build-an-X workflow that analyzes a repo's test mix (unit / integration / E2E counts + runtimes) and recommends rebalancing toward the test pyramid ratios per the change-set shape - pure-logic-heavy repo wants ~80/15/5; UI-heavy repo wants ~60/25/15. Detects 'ice-cream cone' (E2E-heavy) and 'hourglass' (integration-thin) anti-patterns. Use when the user asks about test distribution, test strategy, test balance, too many E2E tests, slow CI caused by tests, testing best practices, or rebalancing their test suite; also suitable for quarterly calibration of the test mix to codebase reality.

test-strategy-author

Authors a test strategy document (a master test plan) for a project, release, or feature - covers scope, in/out, test types per layer (unit / integration / contract / E2E / perf / security / a11y), risk-based test prioritization that maps top risks to test investment (per `risk-matrix`), tooling stack, environments, exit criteria, and ownership. Use when a team needs the release-readiness artifact stakeholders sign off on before significant test investment, and the reference engineering teams return to when scope or quality questions arise.

tool-selection-decision-record

Defines the output contract for writing down a chosen developer tool as a portable decision record: the observed project signal, exactly one primary recommendation, rationale that names the rejected alternative, what to read next, and a mandatory list of the conditions that would flip the choice. Adapts Architecture Decision Record conventions (context, decision, consequences, status, supersede rather than edit) to tool selection, and refuses any recommendation inferred from a README or a folder name instead of a manifest, lockfile, config file, or existing test directory. Distinct from a catalog or advisor that compares candidate tools on their merits: this owns the shape of the written record, not the comparison. Use when a tool has just been chosen (test framework, build tool, linter, package manager, migration tool) and the choice needs to be recorded so a later reader can see the signal, the rejected alternative, and what would reverse it.