Testland
Browse all skills & agents

qa-hiring-kit

End-to-end structured hiring for QA / SDET / automation / test-lead / quality-manager roles - one chain from job description through interview question bank, competency-anchored scoring rubric, interviewer calibration guide, and post-interview panel debrief, to the 30-60-90 day onboarding plan. Implements the canonical Levashina 2014 et al. structured-interview methodology with ISTQB-aligned competency vocabulary, so the JD a candidate reads, the rubric the panel scores, and the ramp plan the manager runs all describe the same role. Use when opening a QA requisition, authoring any artifact of the hiring loop (JD, questions, rubric, calibration guide, onboarding plan), or running the panel debrief to a defensible hire / no-hire decision.

Install with skills.sh (any agent)

npx skills add testland/qa --skill qa-hiring-kit
View source

qa-hiring-kit

Overview

Hiring for QA roles is calibration-heavy: the same question scored by two interviewers without a rubric produces high noise. The remedy, per the structured-interview research (Levashina et al. 2014, Personnel Psychology (opens in new window)), is a structured interview - same questions, same order, same scoring rubric, calibrated interviewers. This kit runs the whole chain as one workflow, with one competency vocabulary (drawn from the ISTQB Foundation Level v4.0 syllabus (opens in new window)) threaded through every artifact:

StageArtifactDeep reference
1. Open the roleJob description + recruiter screening notereferences/jd-author.md
2. AskRole + seniority question bank (STAR, Bloom's K1-K4)references/interview-questions.md
3. Score4-level competency-anchored rubricreferences/rubric.md
4. CalibrateGold-standard answers + panel session scriptreferences/calibration.md
5. DecidePanel debrief to hire / no-hire (workflow below)this file, "Running the debrief"
6. Ramp30-60-90 day onboarding planreferences/onboarding.md

Stages 2-4 are the structured-interview tripod - what we ask, how to score, what each score looks like. None of the three is sufficient alone; running structured interviews requires all three.

When to use

  • A QA / SDET / test-lead / quality-manager requisition is approved and nothing is written yet - start at stage 1 and work down.
  • One artifact of an existing loop needs authoring or recalibrating (a stale question bank, a rubric with drifting anchors, a missing calibration guide) - jump to that stage's reference.
  • An interview round is complete and the panel must converge on a documented hire / no-hire recommendation - run the debrief workflow below.
  • A hire has accepted and the manager needs the first-90-days plan - stage 6.

Do not use this kit to:

  • Map the current team's capability or decide train-vs-hire - that is skill-matrix-author (this plugin); its gap report is the ideal "why this role is open" input to stage 1.
  • Author generic (non-QA) engineering interview loops - the competency model is QA-specific.

The chain, end to end

  1. JD first (references/jd-author.md). Responsibilities derive from ISTQB CTFL §1.4.5's testing-role vs test-management-role split; must-haves reuse the rubric's competency dimensions so the posting and the scoring describe the same role. Ships with a recruiter screening note (signal vs noise per must-have).
  2. Question bank (references/interview-questions.md). 6-8 slots per interview, mixed by role (technical / behavioral-STAR / scenario / system-design), difficulty tuned by seniority via Bloom's K1-K4, with pre-authored follow-up probes. Lock the bank at round start.
  3. Rubric (references/rubric.md). 5-8 competency dimensions per role; four behavioural anchors per (competency × question) cell describing what the candidate said or did; summary by per-dimension floors, never averages.
  4. Calibration guide (references/calibration.md). Four worked gold-standard answers per question (one per score level), common interviewer pitfalls, and the mandatory 90-minute panel calibration session before the first real candidate.
  5. Debrief - the workflow below, run after the final interview.
  6. Onboarding plan (references/onboarding.md). 30-60-90 ramp anchored to the same rubric axes; borderline-scored axes get targeted phase-2 development.

Lock artifacts at the start of a hiring round; mid-round changes invalidate prior candidates' scores.

Running the debrief

After the final interview, the panel converges on a decision in a facilitated calibration loop. Required inputs: role + seniority, the filled-in rubric with each interviewer's per-dimension scores, the calibration guide, and the panel list. Refuse to start if any interviewer's scores were not submitted independently before the debrief - per anchoring-bias research (https://en.wikipedia.org/wiki/Anchoring_effect), the first opinion revealed in a group setting has disproportionate influence on subsequent scorers; independent pre-submission is the primary structural mitigation.

  1. Collect independent scores. Confirm every panelist has scored all dimensions before any scores are shared. If a submission is missing, halt: emit INCOMPLETE_PANEL_SUBMISSION with the missing interviewer's name and a deadline.
  2. Compute per-dimension agreement. Flag any dimension with a spread greater than 1 point. Per the employment-interview research showing anchored rating scales are the load-bearing mechanism for acceptable inter-rater reliability (https://en.wikipedia.org/wiki/Employment_interview), a >1 spread means the anchor was applied differently - the discussion must focus on the anchor text, not general impressions.
  3. Surface evidence, not impressions. For each flagged dimension, each panelist quotes the specific candidate utterance or action that drove their score (the rubric's anchor principle: anchors describe what the candidate said or did, not what the interviewer felt). Impressions not traceable to a quoted behavioural observation are set aside.
  4. Apply the calibration guide to disagreements. Read the relevant gold-standard answers and ask each dissenting panelist: "Which of the four worked examples does this candidate's answer most resemble?" Concrete comparison is the resolution mechanism - not majority vote or seniority deference.
  5. Flag bias language. Watch for the calibration guide's pitfall categories: scoring on tone or confidence; halo-effect generalisation ("the candidate is clearly senior"); anchor drift ("I never give 4s"). When flagged, name the category and re-score against the anchor.
  6. Compute the summary recommendation by the rubric's per-dimension floor rules (any dimension at 1 is a no-hire regardless of totals; two or more dimensions at 2 is a no-hire; one dimension at 2 with all others ≥ 3 is borderline). Never average across dimensions.
  7. Write the decision document: per-dimension score table with evidence anchors, disagreements resolved (scores before discussion, gold-standard comparison used, bias flags raised), the HIRE / NO HIRE / BORDERLINE summary with a rationale traced to specific scores, and required next steps. A BORDERLINE recommendation escalates to the hiring manager with the document - never a committee re-vote. A debrief with no rubric present reverts to an unstructured discussion (MISSING_RUBRIC - halt); that validity loss is what structured interviews exist to prevent (Levashina et al. 2014, https://en.wikipedia.org/wiki/Structured_interview).

Anti-patterns (chain level)

Anti-patternWhy it failsFix
Running interviews with questions but no rubric (or vice versa)Half a structured interview still drifts; the pair is the unit.Stages 2 + 3 travel together; stage 4 before the first candidate.
JD vocabulary differing from the rubricCandidate applies to one role, gets scored on another; debriefs derail.Stage 1 reuses the rubric's competency dimensions.
Skipping the calibration sessionInter-rater agreement does not appear without practice.The stage-4 session is mandatory before the first real candidate.
Deciding by committee re-vote on a borderlineReverts to unstructured group dynamics.Escalate BORDERLINE to the hiring manager with the decision document.
Shelving the rubric after the offerThe hire's weak axes are known and then ignored.Stage 6 converts borderline axes into targeted phase-2 development.

Limitations

  • No compensation benchmarking, legal review, or fairness audit. Salary bands, employment-law phrasing, and bias review of questions / anchors / criteria are HR and legal's call per jurisdiction.
  • Not a technical screen. SDET / automation roles still warrant a take-home or coding round; the kit produces interview artifacts, not coding exercises.
  • Artifacts age. Re-derive the JD per opening and re-calibrate per hiring round; the anchor-drift log in the calibration guide carries refinements forward.

Hand-off targets

  • Why this role is open / train-vs-hireskill-matrix-author (this plugin) - its capability-gap report feeds stage 1.
  • Ongoing development beyond day 90 → the team's career-progression and performance process; out of scope here.

References

  • Structured interview research - Levashina et al. 2014 (Personnel Psychology (opens in new window)) - the validity uplift from same questions / same order / anchored scoring; methodological basis for the whole chain.
  • ISTQB Certified Tester Foundation Level v4.0 syllabus - the competency vocabulary threaded through JD, questions, rubric, and onboarding: https://www.istqb.org/certifications/certified-tester-foundation-level
  • STAR behavioral interviewing method: https://en.wikipedia.org/wiki/Situation,_task,_action,_result
  • Anchoring effect - the bias that independent score pre-submission mitigates: https://en.wikipedia.org/wiki/Anchoring_effect
  • Employment interview - anchored rating scales and inter-rater reliability: https://en.wikipedia.org/wiki/Employment_interview
  • PractiTest 2026 State of Testing Report - hiring-artifact authoring as a high-adoption, low-risk AI use case for QA managers: https://www.practitest.com/state-of-testing/

Interviewer calibration guide authoring

View source (opens in new window)

Interviewer calibration guide authoring

Deep reference for qa-hiring-kit SKILL.md - producing the calibration guide, the what each score looks like leg of the structured-interview tripod.

Why a calibration guide

A rubric describes how to score; a calibration guide demonstrates what a score looks like. The structured-interview research is consistent that anchor descriptions reduce inter-rater noise but do not eliminate it - interviewers still drift on edge cases without worked examples of the same question scored across the whole 1-4 range.

The calibration guide is the third leg of the structured-interview tripod:

Without all three, an interview loop is structurally incomplete - the team will produce an inconsistent signal regardless of how good the questions and rubric are individually. The guide is panel-internal: sharing the gold-standard answers with candidates pollutes the signal.

Step 1 - Capture the inputs

InputNotes
Question bankPer interview-questions.md (opens in new window)
RubricPer rubric.md (opens in new window) - defines the 4-level anchors per dimension
Sample transcripts2 - 5 anonymised candidate transcripts (or recordings the panel has scored). The minimum is 2 (one strong, one weak); 5 is plenty. Halt with INSUFFICIENT_TRANSCRIPTS below 2 - calibration without real examples is theoretical, not load-bearing.
Panel sizeNumber of interviewers who will calibrate; informs the calibration-session timing

If the team is launching a new role with no prior transcripts, use synthesised transcripts (emit transcripts that demonstrate each score level) but flag explicitly that the panel must replace them with real anonymised transcripts after the first 5 candidates.

Step 2 - Per-question gold-standard answers

For each question in the bank, the calibration guide emits four worked answers demonstrating each rubric level (1 = no hire, 2 = borderline, 3 = hire, 4 = strong hire):

### Q3 - Behavioral (STAR): late-defect catch - gold standards

#### Score 1 - no hire

> "Yeah, last release we caught a really bad bug right before launch. The whole team was upset. We were lucky we caught it. After that I'm just always more careful."

**Why score 1:** No technique named; no STAR structure (S/T/A/R undifferentiated); attributes the catch to luck and emotional response rather than process. Anchor: "Cannot articulate a partition / boundary / decision-table technique" - matches.

#### Score 2 - borderline

> "Last release I caught a defect where the cart total wrapped around at $99.99 to a negative number. I added a boundary test for the max-price limit. It was fine after that."

**Why score 2:** Names a technique (boundary value analysis, implicitly); STAR is partial - Situation and Action are clear; Task is implicit ("I was the QA lead"); Result is "it was fine after that" - no measurable outcome, no retrospective learning. Probe for the Task and Result; if the candidate fills both, score moves to 3.

**Probe to use:** "What was your specific role at that point? And how did the team know the fix worked beyond that one test?"

#### Score 3 - hire

> "On the v3.4.0 release we caught a defect where the cart total wrapped around at $99.99 to a negative number. As the QA lead for the release, I was running the standard regression suite when I noticed our test cases all used round-dollar amounts - $50, $100. I added boundary-value tests at the upper limit (one cent under, exactly at, one cent over) and the bug reproduced at exactly $99.99. We fixed the cents-precision rounding in the price-display component, added the boundary tests to the regression suite, and updated our test-data conventions to use Faker's `randomFloat({ precision: 0.01 })` rather than round numbers."

**Why score 3:** Specific technique (boundary value analysis) named **and** applied. STAR complete: situation (v3.4.0), task (QA lead, running regression), specific action (noticed pattern, added boundary tests), measurable result + retro learning (fix shipped, regression added, conventions updated).

#### Score 4 - strong hire

> "[Same situation as score 3, plus:] After we shipped the fix, I noticed our test-data conventions had no rule against round-number-only test data. I authored a one-page conventions update that mandated Faker boundary inputs for any numeric field with a documented constraint, and got the rest of the QA team to sign off. Over the following two releases, we caught two more rounding-precision bugs at the same boundary class - proving the convention change was load-bearing, not just paperwork."

**Why score 4:** Generalises beyond the specific defect to a systemic process change, ties the change to a measurable downstream outcome (two more bugs caught), and documents organisational influence (got the team to sign off).

The four-level demonstration is the load-bearing artifact. An interviewer who reads all four can grade ambiguous transcripts by asking which of these does this candidate's answer most resemble - concrete comparison rather than abstract scoring.

Step 3 - Common pitfalls

For each question, the guide emits a "common pitfalls" section - the typical mistakes interviewers make scoring this question:

### Q3 - common interviewer pitfalls

| Pitfall | Why it produces noise | Correction |
|---|---|---|
| Scoring on tone of voice / "confidence" | Tone is interviewer-specific; the same candidate scored differently by interviewer A and B on tone alone. | Score on what the candidate said, not how. |
| Awarding score 3 because the candidate "is clearly senior" | Halo effect; the question's anchor doesn't mention seniority. | Strict adherence to the anchor's behavioural description. |
| Probing too aggressively after a partial answer | Different interviewers probe to different depths; a thoroughly-probed candidate looks stronger than a lightly-probed one with the same competence. | Use only the pre-authored probes from the question bank; one probe per missing component, no more. |
| Stopping at score 3 because "I never give 4s" | Anchor calibration drift; the rubric's level 4 exists for a reason. | If the candidate matches the level-4 anchor, give level 4; review your prior interviews for under-scoring. |
| Awarding score 4 because "the answer was great" without checking the systemic-gap and measurable-outcome anchors | Halo effect on level 4. | Level 4 requires both anchor sub-conditions (systemic + measurable); level 3 is the default for excellent answers that don't cross both. |

Step 4 - Calibration-session script

The guide's last section is the script for the panel's calibration meeting - typically 90 minutes for a 6-question loop:

  1. Pre-read (the panel reads the guide's gold-standards individually, ~30 minutes before the meeting).
  2. Per-question discussion (~10 minutes per question): the panel scores the same anonymised real transcript independently, reveals scores, discusses any disagreement until the panel converges on a score.
  3. Anchor refinement (~10 minutes): if the panel discovers the rubric's anchor is ambiguous, edit the rubric (and document the change). Calibration sessions are the canonical moment to refine the rubric.
  4. Pitfall review (~10 minutes): the panel walks through Step 3's common pitfalls; each interviewer flags one they recognise from their own prior interviews.

The script is mandatory before the first real candidate. Per the structured-interview research, calibration is the dominant variable in inter-rater agreement - more important than rubric quality alone.

Step 5 - Emit the guide

The output is a single markdown document with:

  1. Header: role, seniority, source question bank + rubric references, panel size, recommended session length.
  2. Per-question gold standards (Step 2 - four worked answers per question).
  3. Per-question pitfalls (Step 3 - common interviewer mistakes).
  4. Calibration-session script (Step 4 - agenda + timing).
  5. Anchor-drift log (initially empty; the panel records every rubric anchor that was tightened during the session, so future hiring rounds inherit the refinement).
  6. Hand-off block: run the calibration session before any real candidate; after the first 5 real candidates, retro the questions and replace synthesised transcripts with real anonymised ones; re-author the guide whenever the role description, seniority, or rubric change.

Anti-patterns

Anti-patternWhy it failsFix
Sharing gold-standard answers with candidatesPollutes the signal; the candidate parrots the answer back.Panel-internal only; tag the file accordingly.
Skipping the calibration session because "we'll just use the rubric"Inter-rater agreement does not appear without practice; the rubric alone is insufficient.Step 4 calibration session is mandatory.
Calibrating on synthesised transcripts onlyThe transcripts are a best-guess; real candidates score in unpredictable shapes.Replace synthesised with real anonymised transcripts after the first 5 candidates.
Authoring the guide without the rubricAnchors drift from the rubric's level definitions.Step 1 hard-requires the rubric.
Locking the guide for the whole hiring round with no anchor-drift logAnchor refinements during real interviews are lost; the guide becomes outdated mid-round.Step 5's anchor-drift log is part of the artifact.
Treating the calibration session as a one-time eventCalibration drifts over months; new interviewers join.Re-calibrate per hiring round, or at least quarterly.

Limitations

  • Synthesised transcripts are educated guesses, not real data. They demonstrate the rubric's intent but do not reflect actual candidate variance. Replace ASAP.
  • Calibration is bounded by panel discipline. A panel that nods through the script without genuinely scoring transcripts together will not achieve inter-rater agreement.
  • The guide is panel-internal and should not leave the hiring loop. Storing it in the team's wiki is fine; sending it to candidates is not.
  • Anchor refinement is a living artifact. The rubric and this guide co-evolve; the team's wiki should hold the latest version, and prior versions should be archived per hiring round.
  • No fairness / bias audit. The gold-standard answers are not checked for class bias - that is the team's HR / legal review.

References

  • Structured interview research - Levashina et al. 2014 (Personnel Psychology (opens in new window)) on the validity uplift from calibrated structured interviews; the methodological basis for the "gold-standards + calibration session" defaults.
  • STAR behavioral interviewing method - Situation / Task / Action / Result framework, the structure of the gold-standard answers for behavioural questions: https://en.wikipedia.org/wiki/Situation,_task,_action,_result
  • ISTQB glossary - the canonical QA vocabulary the gold-standard answers reference (defect, defect-density, escaped-defect, test-design-technique, etc.): https://glossary.istqb.org/
  • Bloom's taxonomy - the K1 - K4 cognitive-difficulty levels used to ensure gold-standard answers match the question's intended depth: https://en.wikipedia.org/wiki/Bloom%27s_taxonomy
  • PractiTest 2026 State of Testing Report - hiring rubric / interviewer calibration cited as a high-adoption, low-risk AI use case for QA managers: https://www.practitest.com/state-of-testing/
  • interview-questions.md (opens in new window), rubric.md (opens in new window) - the two upstream references that produce the questions and the rubric this guide demonstrates.

Interview question bank authoring

View source (opens in new window)

Interview question bank authoring

Deep reference for qa-hiring-kit SKILL.md - producing the QA-role-specific question bank, the what we ask leg of the structured-interview tripod.

Why structured questions

Hiring for QA roles is calibration-heavy: the same question scored by two interviewers without a rubric produces high noise (well-documented in structured-interview research (opens in new window)). The remedy is a structured interview - same questions, same order, same scoring rubric across candidates. This reference produces the questions half of that pair; the scoring rubric (rubric.md (opens in new window)) and the calibration guide (calibration.md (opens in new window)) are the other legs.

The bank is QA-specific by design: (a) ISTQB-aligned competency framing, (b) role-specific question depth (manual QA vs SDET vs test lead require different technical / behavioral mixes), (c) STAR-format anchoring on behavioral questions per the canonical STAR method (opens in new window), and (d) a structured output ready to drop into a hiring loop.

Step 1 - Capture the role inputs

InputNotes
Role titleOne of: manual-qa-engineer, qa-automation-engineer, sdet, test-lead, quality-manager. Each has different default depth weights (Step 2).
SeniorityOne of: junior, mid, senior, staff+. Drives Bloom's-taxonomy difficulty mix in Step 4.
Domain contextThe product area / regulated industry (e.g., "fintech payments", "healthcare EHR", "consumer mobile") - drives scenario-based questions.
Required competenciesOptional; if absent, default to ISTQB Foundation Level chapters relevant to the role (test design, test management, test process, defect management, tools).
Forbidden topicsOptional; topics already covered elsewhere in the loop or out-of-scope for legal / compliance reasons.

If the role title is not one of the five recognised QA roles, halt with UNRECOGNISED_ROLE: supply a role from the recognised list, or run with role=qa-generic to use a flat default mix.

Step 2 - Allocate question slots

A typical 60-minute interview holds 6 - 8 questions; default to a six-question shape. The mix shifts per role:

RoleTechnical depthBehavioral (STAR)Scenario-basedSystem / framework design
manual-qa-engineer2220
qa-automation-engineer3111
sdet2112
test-lead1311
quality-manager0411

The mix is configurable; the table is the default. Behavioral count grows with seniority and people-leadership scope per the structured-interview research; technical depth grows with hands-on coding scope.

Step 3 - Author per-slot questions

For each slot, emit one question with the metadata reviewers need:

### Q3 - Behavioral (STAR) | Senior | Bloom: K3 (Apply)

**Question:** Tell me about a release where you caught a critical defect late - after the test cycle but before production. Walk me through the situation, what your role was, what you did, and what the team learned.

**ISTQB competency:** Defect management (defect → failure distinction, escape-defect lifecycle).
**STAR cues:** Listen for: (S) the release context, (T) the candidate's specific responsibility, (A) the diagnostic and communication actions taken, (R) measurable outcomes + retro learnings.
**Time budget:** 8 min.
**Follow-up probes** (use only if the answer is shallow):
- "Was the defect found through automated tests, manual exploration, or a customer report?"
- "What changed in your team's process after this incident?"
- "How did you handle the stakeholder communication?"

Each question carries:

  • Type: technical / behavioral / scenario / system-design.
  • Seniority anchor: which level the question is calibrated for; senior questions remain useful at staff+ but with deeper probes.
  • Bloom's level (K1 - K4): K1 (remember) for fundamentals, K2 (understand) for explanation, K3 (apply) for situational execution, K4 (analyse) for synthesis. Bloom's levels are used by ISTQB to anchor question difficulty across the Foundation Level syllabus.
  • ISTQB competency the question maps to (cite by stable ID; the ISTQB Foundation Level v4.0 syllabus (opens in new window) is the canonical competency reference, though its specific URL drift means cite by syllabus version).
  • STAR cues (for behavioral questions only) - per the STAR method (opens in new window), the interviewer listens for all four components; missing components prompt follow-up probes.
  • Time budget in minutes.
  • Follow-up probes - pre-authored so different interviewers don't drift into ad-hoc probing (the dominant source of interview noise per structured-interview research).

Step 4 - Tune the difficulty distribution

Bloom's taxonomy mix per seniority (default; configurable):

SeniorityK1K2K3K4
junior30%40%25%5%
mid15%35%35%15%
senior5%25%40%30%
staff+0%15%35%50%

Flag questions whose Bloom's level is too far from the role's centre of gravity (e.g., a K1 fundamental question for a staff+ candidate is a wasted slot).

Step 5 - Emit the bank

The output is a single markdown document with:

  1. Header: role, seniority, domain, total slots, mix summary.
  2. Question slots 1 - N: per the format in Step 3.
  3. Reading list for the candidate (optional): canonical references the candidate can prepare against - ISTQB Foundation Level syllabus, the team's testing handbook, etc. Always optional; avoid if the role advertises "no prep needed".
  4. Hand-off block: pair with rubric.md (opens in new window) for the per-question scoring rubric (the rubric and the questions must travel together; otherwise the loop reverts to unstructured), then calibration.md (opens in new window) once the rubric exists. Lock the question bank at the start of the hiring round; if the bank changes mid-round, every prior candidate's score is no longer comparable.

Anti-patterns

Anti-patternWhy it failsFix
Generic behavioral questions ("Tell me about a time you faced a challenge")Drains the slot; the answer is unscorable because the question lacks specificity.Behavioral questions must name the QA-specific context (release, defect, framework, regulation).
Asking the same question across all seniority levelsThe signal is wasted - a K1 fundamentals question reveals nothing about a staff+ candidate.Step 4 difficulty tuning per seniority.
Including a question already covered in the take-home / coding screenDouble-coverage at the cost of a slot.The forbidden topics input excludes those areas.
Including "puzzle" questions ("estimate the number of QA engineers in your city")Validity is documented to be near zero per structured-interview meta-analyses; the question signals interviewer preference, not candidate competence.Refuse to emit Fermi / puzzle questions. Cite structured-interview research (opens in new window) as the basis.
Behavioural questions without STAR cues for the listenerDifferent interviewers listen for different things; scoring drifts.Step 3 STAR cues are mandatory for behavioural questions.
Letting interviewers free-form their own follow-upsThe dominant source of interview noise.Step 3 pre-authored follow-up probes.
Authoring the question bank without the rubricHalf a structured interview - questions without scoring still drift.Hand-off insists on rubric.md (opens in new window) next.

Limitations

  • The bank is not a substitute for technical screening. A take-home or coding round is still recommended for SDET / automation roles where hands-on code is load-bearing. This produces interview questions, not coding exercises.
  • STAR cues are heuristics, not auto-scoring. The interviewer still has to listen and score. The cue list reduces variance but does not eliminate it.
  • Bloom's levels are the ISTQB-canonical proxy, not a perfect difficulty measure. A K3 question can be more or less difficult than another K3 depending on the topic. Use the level as a coarse mix knob.
  • Locale / language localisation is the team's responsibility. Behavioral questions translated literally can lose nuance; the defaults are English-language.
  • Legal compliance varies by jurisdiction. Some questions (age, family, health) are illegal in some jurisdictions and merely poor practice in others. The default forbidden-topics list covers age, family status, religion, sexual orientation, disability, national origin, and citizenship except where role-required, but the hiring team is responsible for jurisdiction-specific compliance.

References

  • ISTQB Certified Tester Foundation Level v4.0 syllabus - competency areas (test design, test management, test process, defect management, tools); cite by stable syllabus version (v4.0, 2023): https://www.istqb.org/certifications/certified-tester-foundation-level
  • ISTQB glossary - defect, defect-density, escaped-defect, test design technique (the canonical competency vocabulary): https://glossary.istqb.org/
  • STAR behavioral interviewing method - Situation / Task / Action / Result framework: https://en.wikipedia.org/wiki/Situation,_task,_action,_result
  • Structured interview research - Levashina et al. 2014 (Personnel Psychology (opens in new window)) on the validity uplift from structured employment interviews; the methodological basis for the "same questions, same order, same scoring" defaults.
  • Bloom's taxonomy - K1 remember / K2 understand / K3 apply / K4 analyse - ISTQB-adopted cognitive-difficulty levels: https://en.wikipedia.org/wiki/Bloom%27s_taxonomy
  • PractiTest 2026 State of Testing Report - hiring rubric / interview question generation cited as a high-adoption, low-risk AI use case for QA managers: https://www.practitest.com/state-of-testing/
  • rubric.md (opens in new window), calibration.md (opens in new window) - the sibling references that complete the structured-interview triple.

Job description authoring

View source (opens in new window)

Job description authoring

Deep reference for qa-hiring-kit SKILL.md - authoring the QA job description, the upstream-most artifact of the hiring chain.

Why the JD comes first

The JD is the first scoring instrument in the hiring chain, applied by candidates to themselves: a vague posting screens nobody, and an everything-list screens out exactly the people the team wants. The JD's responsibilities come from the role's actual test activities and its skills section uses the same competency vocabulary the downstream rubric (rubric.md (opens in new window)) will score - so the role a candidate applies to is the role the panel evaluates.

Two grounding sources. For what the role does: ISTQB CTFL v4.0 section 1.4.5 defines two principal roles in testing - the testing role, which "takes overall responsibility for the engineering (technical) aspect of testing" and is "mainly focused on the activities of test analysis, test design, test implementation and test execution", and the test management role, which "takes overall responsibility for the test process, test team and leadership of the test activities" and is "mainly focused on the activities of test planning, test monitoring and control and test completion". The same section notes one person may take on both roles at the same time (ISTQB CTFL Syllabus v4.0, §1.4.5 (opens in new window); syllabus text verified 2026-06-10 from the published PDF). For how the document reads: Workable's JD guidance - clear, standard job titles (creative titles like "Rockstar Engineer" read as "unrealistic and potentially discriminatory"), 300 - 660 words total, bulleted duties that show a typical workday, and requirements split into must-have versus nice-to-have (Workable, "How to write a good job description" (opens in new window), fetched 2026-06-10).

Step 1 - Capture the inputs

InputNotes
Role + senioritySame axis the whole kit uses: manual QA / automation / SDET / test lead / quality manager × junior / mid / senior / staff+
Role balanceWhere this role sits between CTFL's testing role and test management role; a senior SDET is nearly all testing role, a test lead carries a documented share of the management role (CTFL §1.4.5 (opens in new window))
Team contextStack and toolchain, domain, test levels in scope, why the role is open (a team capability-gap report from skill-matrix-author is the ideal version of this input)
ConstraintsLocation/remote, compensation-disclosure rules in the posting jurisdictions, non-negotiables

Step 2 - Derive responsibilities from the role's test activities

Write 5 - 8 responsibility bullets, each traceable to a test activity, phrased as a typical-workday action (Workable's guidance: duties should show what a normal day looks like, not an aspirational mission). Map per role balance:

  • Testing-role bullets draw on test analysis, test design, test implementation, test execution (CTFL §1.4.5 (opens in new window)): "design and automate API-level regression tests for the payments services", "run and extend exploratory charters on new checkout features".
  • Test-management bullets draw on test planning, monitoring and control, completion: "own the test strategy section for your product area", "report quality status to engineering leadership each sprint".

A role that is 100% one kind needs no bullets from the other; a mixed role states the split rather than hiding it ("~70% hands-on automation, ~30% process ownership").

Step 3 - Split skills into must-have vs nice-to-have

The split is the JD's main screening mechanism (Workable: be upfront about non-negotiables; separate must-have from nice-to-have so candidates self-assess accurately). Rules:

  • Must-haves: 4 - 6 items maximum, each one the team would genuinely reject an otherwise-strong candidate for lacking. Use the competency dimensions from rubric.md (opens in new window) Step 2 for the role (e.g., for automation: test analysis and design, test code conventions, tooling depth, communication) so JD and rubric stay one vocabulary.
  • Nice-to-haves: 3 - 5 items, explicitly labeled as such.
  • Include the generic tester skills only when they will be screened for: CTFL 1.5.1 lists testing knowledge, thoroughness, curiosity, attention to detail, being methodical, and communication among skills particularly relevant for testers - but as JD boilerplate they screen nothing; tie them to an observable ("bug reports that developers can reproduce first try") or leave them out.
  • Certifications (ISTQB included) are nice-to-have evidence of knowledge, not must-have proxies for skill, unless a client or regulator mandates them.

Step 4 - Define screening signals for the recruiter

The JD ships with a one-page screening note (internal, not posted): for each must-have, what in a CV or portfolio counts as signal vs noise. Example for "tooling depth, Playwright":

SignalNoise
Public repo or described project with Playwright specs they authored"Playwright" in a skills word-cloud
Describes flake-debugging or CI-stabilization workLists every test tool released since 2015
Automation framework decisions they can own ("migrated from X because...")Certification list with no applied work

This note is what keeps the recruiter's screen consistent with the panel's rubric - the same chain-of-custody idea the structured-interview triple applies after the screen.

Step 5 - Assemble and length-check the JD

Order: title, one-paragraph role summary (which includes why the role is open), responsibilities, must-haves, nice-to-haves, team and stack, process and timeline, compensation per jurisdiction rules. Target 300 - 660 words total per Workable's guidance; bulleted lists for mobile readability. Worked example (condensed):

# Senior QA Automation Engineer - Payments

We run weekly releases for a payments product used by 40k merchants. This role
is open because our capability review found one engineer covering performance
testing for the whole group; you will broaden and own that coverage. ~80%
hands-on engineering, ~20% strategy input for your area.

## What you will do
- Design and automate regression tests for payment-retry and reconciliation flows
  (TypeScript + Playwright, k6 for load profiles).
- Extend the CI quality gates: flake quarantine, suite budget, pass-rate reporting.
- Run risk-based test analysis on new payment features with the product trio.
- Coach two mid-level engineers on test design through review.

## Must have
- Test analysis and design: you can derive tests from risks and requirements,
  not only from acceptance criteria handed to you.
- Production-grade test code in TypeScript or a near language; you have owned
  a suite others contribute to.
- Load or performance testing on at least one real system (k6, Gatling, or similar).
- Written communication: bug reports and strategy notes that stand alone.

## Nice to have
- Payments or other regulated-domain background; ISTQB CTAL-TA; CI ownership
  (GitHub Actions); accessibility testing exposure.

(~420 words in full form, inside the 300 - 660 band.)

Anti-patterns

Anti-patternWhy it failsFix
Unicorn JD (every tool, every level, "QA ninja")Screens out honest strong candidates; Workable flags creative titles as unrealistic and potentially discriminatoryStandard title; 4 - 6 must-haves with the Step 3 reject-test
Years-of-experience as must-havesYears measure exposure, not competence, and import biasPhrase must-haves as capabilities with observable evidence
JD vocabulary differing from the rubricCandidate applies to one role, gets scored on another; debriefs derailStep 3 reuses the rubric's competency dimensions
Responsibilities copied from a templateThe workday described matches no actual workday; early attrition followsStep 2 derives bullets from the team's real test activities
Hiding the management share of a lead roleCandidates discover the meeting load after signingState the split explicitly (Step 2)
Posting with no screening noteRecruiter invents their own filter; the funnel disconnects from the rubricStep 4 note ships with the JD

Limitations

  • No compensation benchmarking. Salary bands come from the org's compensation source; only enforce that disclosure rules per jurisdiction are respected in the posting.
  • No legal review. Employment-law phrasing (at-will clauses, accommodation language) varies by jurisdiction and is HR/legal's call.
  • Channel strategy out of scope. Where to post and how to source is recruiting strategy; this produces the artifact, not the campaign.
  • The JD ages. Re-derive from current team context per opening; a reposted two-year-old JD advertises a team that no longer exists.

References

  • ISTQB CTFL Syllabus v4.0, section 1.4.5 "Roles in Testing" (testing role vs test management role definitions quoted above) and 1.5.1 "Generic Skills Required for Testing": https://istqb.org/wp-content/uploads/2024/11/ISTQB_CTFL_Syllabus_v4.0.1.pdf - syllabus text verified 2026-06-10 from the published v4.0 PDF (ISTQB resource CDN copy at https://d288qud2qgn4l3.cloudfront.net/media/resources/ISTQB_CTFL_Syllabus-v4.0.pdf).
  • Workable, "How to write a good job description" - title guidance, 300 - 660 word target, bulleted typical-workday duties, must-have vs nice-to-have split: https://resources.workable.com/tutorial/how-to-write-a-good-job-description (fetched 2026-06-10).
  • interview-questions.md (opens in new window), rubric.md (opens in new window), calibration.md (opens in new window), onboarding.md (opens in new window) - the downstream hiring chain.

Onboarding plan authoring (30-60-90)

View source (opens in new window)

Onboarding plan authoring (30-60-90)

Deep reference for qa-hiring-kit SKILL.md - the post-hire ramp artifact. It starts at offer acceptance, not during the interview loop.

Why a structured ramp

Hiring ends at offer acceptance; onboarding is where the competency signal from the rubric is converted into a development plan. The 30-60-90 day framework - originally popularized by Michael Watkins' The First 90 Days (2003, Harvard Business Review Press) as the canonical structured-transition framework for new organizational members - divides the ramp into three equal phases: learn (days 1-30), integrate (days 31-60), and contribute independently (days 61-90). Each phase has a distinct focus, observable exit criteria, and a handoff to the next. (Notion blog on the 30-60-90 day plan framework: https://www.notion.com/blog/30-60-90-day-plan; Wikipedia, "Onboarding": https://en.wikipedia.org/wiki/Onboarding.)

The plan anchors each phase's competency targets to the six rubric axes from rubric.md (opens in new window): test analysis and design, defect lifecycle, test code conventions (automation roles), tooling depth, communication, and domain reasoning. Rather than treating onboarding as a generic HR checklist, the plan treats it as a continuation of the structured-interview signal - a hire scored 2 on "test code conventions" during the loop gets targeted development investment in exactly that axis during phase 2.

The PractiTest 2026 State of Testing Report found that nearly 40% of individual contributors feel "test strategy" is underdeveloped in their teams, and that practitioners who pivot toward strategy earn a +10.6% income premium versus those who remain in pure technical execution (https://www.practitest.com/state-of-testing/). The day-61-90 phase directly addresses this gap by pushing the hire toward strategy ownership proportional to seniority.

Step 1 - Capture the inputs

InputNotes
Seniorityjunior / mid / senior / lead - the same axis used in the hiring rubric
Role variantmanual-qa-engineer / qa-automation-engineer / sdet / test-lead - determines which rubric axes are load-bearing
Rubric scoresThe per-dimension hire scores from rubric.md (opens in new window) - axes scored 2 ("borderline") get targeted phase-2 development plans, not just the default milestones
Team contextStack, CI toolchain, domain (fintech / consumer / B2B SaaS / etc.) - informs phase-1 environment setup and phase-2 tooling targets
Mentor availabilityWhether a dedicated senior QA mentor is available, or whether the new hire will pair with an engineering team member

If rubric scores are not available (e.g., an informal hire), treat all six axes as equally weighted development targets and flag this assumption explicitly in the output.

Step 2 - Determine the seniority multiplier

Seniority changes the pace and scope of the plan, not its three-phase structure. The table below describes the expected exit state for each seniority at day 90 - the plan's success criteria derive from this:

SeniorityDay-30 exitDay-60 exitDay-90 exit
JuniorEnvironment set up; first test case authored and passing in CI; has read the team's test conventions docOwns one test suite with no supervision; contributing to defect triage meetingsParticipating in test planning for a sprint; mentor cadence reduces to fortnightly
MidEnvironment set up; first test authored and reviewed; independently diagnosed one CI failureOwns a feature's test coverage from planning through execution; identified at least one test-architecture gapLeads test planning for a sprint; peers consult them on test design questions
SeniorEnvironment set up; independent on toolchain; has reviewed one teammate's test PR and provided actionable feedbackOwns test strategy for a component or service; running the weekly test review meetingDriving test strategy for the team; mentoring junior or mid QA hires; identifiable as the quality signal for a release
LeadIndependent on toolchain and team norms; has mapped the team's test coverage gaps; has met all stakeholders (product, engineering, support)Owns the test strategy document for the quarter; running test-planning ceremoniesQuality-gate ownership at the release level; proposals for tooling or process changes with team sign-off

[author opinion] These exit criteria follow the general "learn - integrate - contribute" arc described in the 30-60-90 day literature, adapted to QA-specific observable outputs. The pace should be recalibrated if the team's domain or compliance environment (e.g., healthcare, finance) has an unusually long environment-access ramp.

Step 3 - Author the three-phase plan

For each seniority, emit three sections. The structure below is for a mid qa-automation-engineer; adjust the target values per the multiplier table in Step 2 and the rubric scores from Step 1.

Phase 1 (days 1-30): Learn

Focus: environment, context, and first contribution. Per the 30-60-90 day framework (https://www.notion.com/blog/30-60-90-day-plan), this phase centres on acquiring the knowledge and role clarity needed to begin contributing - not on producing output.

Milestones:

  1. Development environment running and verified against the team's setup doc by day 3.
  2. All repository access, CI pipelines, and test-infrastructure credentials provisioned by day 5.
  3. Read the team's test conventions document and the project's domain README by day 7.
  4. Paired with mentor for a 1-hour walkthrough of the existing test suite by day 7.
  5. First test case authored (existing coverage gap identified by the mentor) passing in CI by day 14.
  6. Attended one sprint planning, one defect triage, and one retrospective by day 21.
  7. Written a bug report for a defect found during exploratory testing, reviewed by a senior QA or engineer by day 28.
  8. Completed ISTQB Foundation Level Chapter 1 self-study (Fundamentals of Testing) if not already certified (ISTQB CTFL v4.0 syllabus: https://astqb.org/certifications/foundation-level-certification/) by day 30.

Competency targets (rubric axes - first contact):

Rubric axisPhase-1 target
Test analysis and designCan name the technique used when authoring the first test (boundary, equivalence partition, or decision-table)
Defect lifecycleWrites a bug report that includes steps to reproduce, expected vs. actual, environment, and severity - reviewed and accepted by a senior
Test code conventionsFirst test passes lint and follows the team's naming convention
Tooling depthCan run the test suite locally, filter by tag, and read a CI run log
CommunicationAttends ceremonies and asks one clarifying question per session
Domain reasoningCan describe the product's user journey at a feature level

Mentor cadence (phase 1): 30-minute daily check-in for the first 2 weeks; 1-hour weekly thereafter. Mentor reviews all authored tests before merge.

Success criteria for phase-1 exit: First test merged to the main test suite; bug report accepted without revision; mentor confirms independent environment operation.

Phase 2 (days 31-60): Integrate

Focus: applying skills under supervision and owning a bounded scope. Per the 30-60-90 day framework, this phase shifts toward integration with the team and application of the new hire's existing skills, moving from supervised contribution to scoped ownership.

Milestones:

  1. Takes ownership of one feature's test suite - additions, maintenance, and CI failures - by day 35.
  2. Attends all sprint test-planning sessions and proposes at least one test case per story by day 40.
  3. Reviews one teammate's test PR and provides written feedback by day 45.
  4. Identifies and documents one test-architecture gap (e.g., a boundary class with no negative tests) by day 50.
  5. Runs a 30-minute knowledge-share session for the QA or engineering team (format: what I learned in month 1) by day 55.
  6. Completes ISTQB Foundation Level Chapters 2-4 self-study (Test Activities and Roles; Static Testing; Test Analysis and Design) if not certified by day 60.

Competency targets (rubric axes - applied):

Rubric axisPhase-2 target
Test analysis and designApplies at least two ISTQB techniques (from CTFL Ch. 4: equivalence partitioning, boundary value analysis, decision table, state transition - https://astqb.org/certifications/foundation-level-certification/) independently to a real feature
Defect lifecycleOwns the defect-triage process for the owned feature suite; can distinguish defect vs. failure per ISTQB glossary terminology
Test code conventionsTest PRs pass review without convention-related comments on second PR and after
Tooling depthCan diagnose a flaky test (identify root cause from the CI log) and open a fix PR without mentor assistance
CommunicationDelivers the knowledge-share session without material feedback from the mentor
Domain reasoningCan map a new user story to the affected test boundary without prompting

Targeted development for rubric-axis score of 2 (borderline at hire): For each axis where the hire scored "borderline" during the loop, the plan adds a specific paired-learning task: e.g., a score-2 on "test code conventions" triggers a required code-review pairing with a senior on two consecutive PRs; a score-2 on "defect lifecycle" triggers a defect-retrospective exercise where the hire re-examines three historical bug reports and identifies the point of detection vs. the point of failure.

Mentor cadence (phase 2): 1-hour weekly; mentor shifts from prescriptive reviewer to code-review approver. Mentor stops blocking on approvals by day 45 (the hire uses the team's standard PR review process).

Success criteria for phase-2 exit: Owned feature suite green in CI with no unreviewed failures; architecture-gap document shared and acknowledged by the team; knowledge-share delivered.

Phase 3 (days 61-90): Contribute independently

Focus: independent contribution and the beginning of strategy ownership proportional to seniority. Per the 30-60-90 day framework, this phase prepares the hire to lead their first project or initiative - moving from "applying skills" to "driving decisions."

Milestones:

  1. Leads test planning for one sprint end-to-end (story estimation, test-case authorship, risk identification) by day 70.
  2. Produces a one-page test-coverage assessment for an assigned component, with a risk-ranked gap list, by day 75.
  3. Proposes one process or tooling improvement with a concrete rationale and presents it at the team's retro by day 80.
  4. Completes ISTQB Foundation Level Chapters 5-6 self-study (Managing the Test Activities; Test Tools) if not certified by day 85.
  5. First 30-minute solo 1:1 with hiring manager at day 90: hire drives the agenda using the success criteria below.

Competency targets (rubric axes - owned):

Rubric axisPhase-3 target
Test analysis and designAuthors a test strategy note for a sprint (not just individual test cases) - identifies which techniques apply and why
Defect lifecycleTracks defect escape rate for the owned suite over two sprints and presents the trend to the team
Test code conventionsTest PRs pass review first-pass on average (no convention-related revision comments)
Tooling depthHas extended or configured at least one existing test-infrastructure component (e.g., added a new test tag, extended a shared fixture, or updated a CI step)
CommunicationWritten sprint test-coverage summary shared with product and engineering without prompting
Domain reasoningCan identify the highest-risk user flows for a new feature with no rubric reference - asks product or engineering for confirmation, not instruction

Mentor cadence (phase 3): Bi-weekly 30-minute check-in; mentor role shifts from reviewer to consultant. New hire sets the agenda. Formal mentor relationship closes at day 90; standard peer-review process applies thereafter.

Success criteria for phase-3 (day-90) exit - the success criteria double as the hiring manager's 90-day review inputs:

  1. Owns at least one test suite end-to-end with no unresolved CI failures.
  2. Has led at least one sprint's test planning with documented coverage decisions.
  3. Has produced at least one written artifact (coverage assessment, process proposal, or knowledge-share) that other team members reference.
  4. Rubric-axis score improvements: any axis scored "borderline" at hire has a documented development event (pairing session, retro exercise, PR review) with a mentor note on progress.
  5. Peer feedback (from one engineer and one QA reviewer, gathered by the hiring manager): the hire is "someone I can hand a feature to."

Step 4 - Emit the plan

The output is a single markdown document with:

  1. Header: hire name (or placeholder), role, seniority, start date, hiring manager, mentor.
  2. Rubric-score summary: the per-axis hire scores from the structured interview, copied from the rubric output - these anchor the targeted-development items in phase 2.
  3. Three phase sections (Step 3 structure per seniority).
  4. Milestone tracker table: a flat checklist of all milestones with day targets and owner (hire / mentor / hiring manager) for easy status tracking.
  5. Hand-off block: share the plan with the new hire on or before day 1 (it is not confidential - the hire should know the success criteria); hiring manager reviews phase-exit criteria at day 30, 60, and 90; any axis still "borderline" at day 60 gets a second targeted-development pairing before day 75; at day 90, archive the plan alongside the rubric output in the team's hiring record - the question bank, rubric, and onboarding plan together form the hiring-to-ramp provenance trail.

Anti-patterns

Anti-patternWhy it failsFix
Identical onboarding plans regardless of seniorityA lead who spends 30 days on environment setup and basic test authoring is wasted and will disengage.Step 2's seniority multiplier sets phase-exit expectations; senior and lead plans front-load ownership.
Ignoring the rubric scoresThe structured interview identified the hire's weak axes; not acting on them is the most common onboarding gap.Phase 2's targeted-development section is mandatory when any axis scored "borderline" at hire.
Treating "success criteria" as aspirationalIf the 90-day criteria are not measurable, the hiring manager and hire will disagree on performance at day 91.Step 3's success criteria are observable artifacts (merged tests, written documents, peer feedback), not feelings.
Over-loading phase 1New hires cannot absorb toolchain, domain, team norms, and first contribution simultaneously.Phase 1's milestones are sequenced: access and environment first, first contribution second, ceremonies third. Defer strategy discussions to phase 2.
Closing the mentor relationship before day 90Phase-3 mentoring shifts to consulting cadence (bi-weekly), not zero. Early closure removes the safety net for the hire's first strategy-ownership attempt.Mentor cadence is explicit per phase; the formal relationship closes only at the day-90 review.
Re-running the onboarding plan as a performance planThis plan covers the ramp to independent contribution, not ongoing performance management.If the hire meets the day-90 criteria, transition to the team's standard career-progression process. If they do not, that is a separate conversation outside this artifact's scope.

Limitations

  • The plan assumes rubric scores exist. Without per-axis scores, the targeted-development section of phase 2 defaults to generic guidance and loses its main differentiation from a generic onboarding checklist.
  • Domain-specific ramp times vary. In regulated industries (healthcare, finance), environment access and compliance training can extend phase 1 by 2-4 weeks. Flag this assumption and adjust milestones accordingly.
  • Mentor availability is a hard dependency. The phase-1 daily check-in and phase-2 PR review cadence require a named mentor. If no dedicated QA mentor is available, substitute a senior engineer and flag the capability gap.
  • The ISTQB self-study milestones are optional enrichment. The CTFL Foundation Level syllabus (https://astqb.org/certifications/foundation-level-certification/) covers test fundamentals, static testing, test analysis and design, test activities, and test tools. For hires who are already certified, replace these milestones with domain-specific reading or toolchain deep-dives.
  • No fairness review. The success criteria are not audited for differential treatment across protected classes - that is the team's HR review.

References

  • Michael Watkins, The First 90 Days (Harvard Business Review Press, 2003 / 2013) - the foundational 90-day structured-transition framework. The learn / integrate / contribute phase arc follows Watkins' "orienting - learning - acting" progression.
  • Notion, "30-60-90 day plan" - three-phase framework description (learn / integrate / lead): https://www.notion.com/blog/30-60-90-day-plan
  • Wikipedia, "Onboarding" - organizational socialization definition and four adjustment dimensions (role clarity, self-efficacy, social acceptance, organizational culture knowledge): https://en.wikipedia.org/wiki/Onboarding
  • ISTQB Certified Tester Foundation Level v4.0 syllabus (ASTQB mirror) - competency chapters cited for self-study milestones (Ch. 1: Fundamentals; Ch. 2: Test Activities and Roles; Ch. 3: Static Testing; Ch. 4: Test Analysis and Design; Ch. 5: Managing Test Activities; Ch. 6: Test Tools): https://astqb.org/certifications/foundation-level-certification/
  • PractiTest 2026 State of Testing Report - test strategy underdevelopment (40% gap) and the strategy-vs-execution income premium (+10.6%) that motivates the phase-3 strategy-ownership milestones: https://www.practitest.com/state-of-testing/
  • rubric.md (opens in new window) - the upstream reference whose six competency axes are the target axes for this plan's per-phase competency tables.
  • calibration.md (opens in new window) - the calibration guide that closes the interview loop; the onboarding plan picks up where it ends (post-offer-acceptance).

Hiring rubric authoring

View source (opens in new window)

Hiring rubric authoring

Deep reference for qa-hiring-kit SKILL.md - producing the competency-anchored scoring rubric, the how to score leg of the structured-interview tripod.

Why anchored rubrics

Without a rubric, two interviewers asking the same question produce different scores; the literature on structured interviewing (opens in new window) is clear that the questions alone are not sufficient - the scoring rubric is what converts them into a comparable signal.

Anchored rubrics outperform free-form scoring because the anchor descriptions at each level (no-hire / borderline / hire / strong-hire) constrain what each score means. An interviewer who reads "level 3: candidate explains the AAA pattern with a worked example and identifies one of: assertion strength, mocking pitfalls, or fixture coupling" cannot drift the score on tone or rapport - the anchor is concrete.

Step 1 - Capture the inputs

InputNotes
Role + senioritySame as the upstream question bank - manual QA / SDET / automation / test lead / quality manager × junior / mid / senior / staff+
Question bankThe output of interview-questions.md (opens in new window). Each question's competency tag drives the rubric's competency-by-question matrix.
Team's competency modelOptional. If absent, default to the ISTQB-aligned model in Step 2.

If a question bank is not available (e.g., an ad-hoc loop, or an existing interview set that was never written down), author one anchor set per competency dimension rather than per (competency × question) cell, mark the rubric provisional in the header, and flag this assumption explicitly. A provisional rubric must be re-run against the bank once it exists - competency-general anchors drift from the questions actually asked, which is the failure the requirement exists to prevent.

Step 2 - Pick the competency dimensions

A QA hiring rubric scores against 5 - 8 competency dimensions. The default set (drawn from ISTQB Foundation Level v4.0 (opens in new window) competencies and adapted to interviewable behaviour) per role:

manual-qa-engineer / qa-automation-engineer

  1. Test analysis & design - partitioning, boundary, decision-table reasoning per ISTQB technique.
  2. Defect lifecycle - defect vs failure distinction; bug-report quality; reproducibility.
  3. Test code conventions (automation only) - AAA structure, assertion strength, mocking discipline.
  4. Tooling depth - fluency with the team's primary toolchain (Playwright / Cypress / Selenium / pytest / JUnit / etc.).
  5. Communication - written bug reports; verbal hand-off to engineering.
  6. Domain reasoning - applies QA techniques to the team's domain (fintech / healthcare / consumer mobile).

sdet

  1. Test analysis & design.
  2. Test code conventions.
  3. Test framework / tool architecture - how to extend the team's framework; CI integration; flake budget.
  4. Production-quality coding - AAA, refactoring, naming, fixture cleanliness.
  5. System reasoning - service boundaries; what to test at which layer.
  6. Communication & collaboration.

test-lead

  1. Test strategy authoring - risk-based testing; the test pyramid as an argument, not a template.
  2. Stakeholder management - engineering, product, support, leadership.
  3. Hiring & coaching of QA team members.
  4. Defect management at the team / cross-team layer.
  5. Tooling & CI ownership.
  6. Communication (written + verbal, exec-level).

quality-manager

  1. Quality strategy across releases / quarters.
  2. Risk-based prioritisation - data-informed decisions with traceability.
  3. Stakeholder communication, exec-level.
  4. Hiring & team development.
  5. Process / methodology fluency - agile, BDD, shift-left, shift-right, when each applies.
  6. Defect / escape management at the org layer.
  7. Compliance / regulated-industry framing (if applicable).

Emit the dimensions selected for the role; the team can add or remove dimensions before locking the rubric.

Step 3 - Author the 4-level anchors per dimension

For each (competency × question) cell, the rubric needs four behavioural anchors. The anchor describes what the candidate said or did, not what the interviewer felt - this is the load-bearing principle that reduces interviewer noise.

### Test analysis & design - Q3 (Behavioral, STAR: late-defect catch)

| Score | Anchor (what the candidate said / did) |
|---|---|
| **1 - no hire** | Cannot articulate a partition / boundary / decision-table technique. Describes the catch as "I just got lucky." Or attributes the catch to a tool ("the linter caught it"). |
| **2 - borderline** | Names one ISTQB technique correctly but cannot apply it to the catch they describe. STAR is partial: missing Result or missing the candidate's specific Action (says "we" throughout). |
| **3 - hire** | Identifies the specific technique that caught the defect (e.g., "we had no negative test for the empty-cart case - equivalence partitioning would have flagged it"). STAR complete: situation, task, the candidate's specific action, measurable result + retro learning. |
| **4 - strong hire** | Generalises beyond the specific defect: identifies a systemic gap (e.g., "we had no convention requiring a negative test per public method; I added that to our conventions doc"), and ties the change to a measurable downstream improvement. |

**Probe-trigger:** If the candidate scores 2 on STAR completeness, probe for the missing component; do not deduct further on the second pass.
**Time-budget impact:** A score of 4 typically takes 2 extra minutes; budget accordingly.

Each anchor is concrete enough that two interviewers reading the same transcript would arrive at the same score - that is the only test of the anchor's quality.

Step 4 - Compute the role-level summary score

The rubric outputs a per-dimension score and a summary recommendation. The summary is not a simple average:

Per-dimension scoring ruleSummary recommendation
All dimensions ≥ 3, ≥ 1 dimension at 4Strong hire
All dimensions ≥ 3Hire
1 dimension at 2, all others ≥ 3Borderline - debrief required
≥ 2 dimensions at 2, no 1sNo hire - competency gap
Any dimension at 1No hire - fundamental gap

The summary refuses to average across competencies - a candidate weak in defect lifecycle and strong in tooling depth is not "average"; the role demands both. Per-dimension floors are the load-bearing constraint.

Step 5 - Emit the rubric

The output is a single markdown document with:

  1. Header: role, seniority, source question bank reference, competency dimensions, summary-rule table.
  2. Per-question scoring sections (one per question in the bank, scoring against each competency the question targets - typically 1 - 2 competencies per question).
  3. Summary recommendation rules.
  4. Hand-off block: pair with calibration.md (opens in new window) for gold-standard model answers and common pitfalls per question - without those, the anchors here are aspirational. Run a calibration interview before the first real candidate (per the structured-interview research, calibration is the dominant variable in inter-rater agreement). Lock the rubric at the start of the hiring round; mid-round changes invalidate prior candidates' scores. After the round, retro the rubric: which competencies discriminated; which were noise; which scored everyone at 3 (a sign the anchor is too generous).

Anti-patterns

Anti-patternWhy it failsFix
Free-text "1 - 5 score" with no anchorsThe score is the interviewer's opinion, not a behavioural observation.Step 3 anchors are mandatory; no anchorless dimensions.
Anchors that describe the interviewer's feeling ("I was impressed", "the candidate seemed confident")Tone signals; not behaviour. Interviewer noise is the dominant source.Anchors describe what the candidate said or did verbatim.
Averaging dimension scores into a summaryHides the load-bearing competency gaps.Step 4's per-dimension floor; no averages.
Using the same rubric across seniority levelsA senior candidate at "score 3" is mid-level performance for that role; the absolute number means different things.Per-seniority anchors; junior-3 ≠ senior-3.
Rubrics with 10+ dimensionsInterviewer can't hold them all; scoring fragments.Cap at 5 - 8 dimensions.
Rubric authored without the question bankAnchors drift from the actual questions; scoring becomes generic.Step 1 hard-requires the question bank as input.
"Cultural fit" as a dimensionDocumented bias amplifier; legally fraught.Use the team's Definition of Done / engineering values translated into behavioural anchors instead.

Limitations

  • The rubric is only as good as its anchors. Vague anchors produce inter-rater drift; concrete behavioural anchors take time to author and refine.
  • Anchor-validation requires real candidate data. Until the rubric has been used through 5 - 10 interviews, its anchors are theoretical. Plan a calibration interview before the first real candidate.
  • Rubrics drift over time. A rubric may anchor on tools that are no longer the team's default. Re-author per hiring round, or at least review.
  • No fairness audit. The rubric is not checked for bias against protected classes - that is the team's HR / legal review.
  • Weighting is uniform per dimension. Some teams want to weight tooling depth higher than communication; emit unweighted scores and leave weighting to the hiring manager. Custom weights can be applied post hoc to the per-dimension scores.

References

  • ISTQB Certified Tester Foundation Level v4.0 syllabus - the competency model adapted into the default dimensions per role: https://www.istqb.org/certifications/certified-tester-foundation-level
  • ISTQB glossary - defect / failure distinction (load-bearing for the defect lifecycle dimension): https://glossary.istqb.org/en_US/term/defect
  • Structured interview research - Levashina et al. 2014 (Personnel Psychology (opens in new window)) on the validity uplift from structured rubrics + same questions / same order.
  • STAR behavioral interviewing method - Situation / Task / Action / Result framework, used in the behavioural-question anchors: https://en.wikipedia.org/wiki/Situation,_task,_action,_result
  • Bloom's taxonomy - K1 - K4 cognitive levels used to align the rubric's anchor depth with the question's intended difficulty: https://en.wikipedia.org/wiki/Bloom%27s_taxonomy
  • PractiTest 2026 State of Testing Report - hiring rubric authoring named as a high-adoption, low-risk AI use case for QA managers: https://www.practitest.com/state-of-testing/
  • interview-questions.md (opens in new window), calibration.md (opens in new window) - the sibling references that complete the structured-interview triple.

Related skills

exec-quality-narrative

Build-an-X workflow that turns already-computed quality data - weekly digests, KPI roll-ups, DORA delivery metrics, escape-defect trends, OKR grading - into an executive or QBR narrative structured by the Minto Pyramid Principle: governing answer first, MECE-grouped support beneath it, SCQA opening (Barbara Minto, The Pyramid Principle, ISBN 978-0273710516). Distinct from single-team digest computation (which computes the RAG digest from raw CI and tracker signals; this skill consumes such digests and writes the upward story), from portfolio-review aggregation (which aggregates teams into a portfolio review; this skill is the communication layer either output feeds), and from QA OKR authoring (forward-looking commitments; this skill narrates what happened and what it means). Use before a QBR, board update, or exec review when the data exists but the story does not.

quality-status-digest

Computes a recurring quality status digest from metrics that already exist: CI pass rate with an explicit denominator rule, escape-defect count, and a flake-debt score, assigns red / amber / green per area against stated thresholds, then rolls the same per-team rows into a portfolio view with a severity-by-blast-radius heatmap, STABLE / WATCH / INVEST tags, and a capacity flag. Keeps DORA delivery metrics separate from defect-leakage and flake measures instead of blending them under one label. Produces the status artifact only: it does not instrument anything, does not define SLOs or targets, and does not decide what gets fixed first. Use when a weekly quality review, sprint check-in, or quarterly portfolio review is due and the CI history, defect tracker, and quarantine list already hold the numbers but nobody has assembled them into one page.

skill-matrix-author

Build-an-X workflow that produces a QA team skill matrix - team members crossed with competency dimensions at explicit proficiency levels, each cell backed by observable evidence - then derives the full gap analysis: classify each gap (coverage / capability / bus-factor / surplus), rank the gaps against the team's roadmap, and recommend a closing move per gap (train / peer-learn vs hire vs external expert). Competency dimensions follow ISTQB CTAL-TM v3.0 chapter 3 (Managing the Team): professional, methodological, social, and personal competence. Maps the existing team on an ongoing basis - not a point-in-time score of external candidates and not one new hire's ramp plan. Use when a QA manager needs to know what the team can do today versus what its projects demand - before quarterly planning, a training-budget decision, or opening a requisition.