junit-xml-analysis
Parses JUnit-format XML reports (the de-facto interchange format every CI ingests - Jenkins, GitHub Actions, GitLab, Buildkite, CircleCI) into structured, machine-readable per-suite and per-case metrics tables (passed / failed / errored / skipped, time, classname, message, stack), groups failures by classname for trend analysis, and distinguishes "new failures vs flakes" by cross-referencing the `flakyFailure` and `rerunFailure` rerun elements. Use when the downstream consumer is a dashboard, script, or aggregator - not when the goal is a human-readable prose summary (use test-run-summary-author for that). Single-run, in-XML aggregation only; for cross-run cross-environment roll-ups, use a cross-run test-suite aggregator.
Install with skills.sh (any agent)
npx skills add testland/qa --skill junit-xml-analysisjunit-xml-analysis
Overview
The "JUnit XML" format is the de-facto schema every CI consumes, emitted by virtually every test runner (pytest, Jest, Vitest, Go test, Maven Surefire, Cypress, Playwright, and the rest).
Per llg-junit (opens in new window) (the community schema reference used by Jenkins's parser):
"Root element:
<testsuites>(optional if only one suite exists;<testsuite>can be the root instead)."
The hierarchy is testsuites → testsuite → testcase, with result child elements (<failure>, <error>, <skipped>) hanging off each testcase. This skill covers parsing the format, building per-suite + per-case metrics, and the flaky-vs-new distinction via the modern <rerunFailure> / <flakyFailure> extensions.
When to use
Step 1 - Schema overview
Per llg-junit (opens in new window):
| Level | Required attributes | Common attributes |
|---|---|---|
testsuites | (none required at root) | tests, failures, errors, disabled, time, name |
testsuite | name, tests | failures, errors, skipped, time, timestamp, hostname, id, package |
testcase | name, classname | time, assertions, status |
Each <testcase> contains at most one of:
Plus optional:
Critical distinction: per llg-junit (opens in new window), <failure> is an assertion failure (the test made a claim that came back false). <error> is an exception or crash before the assertion ran. Group them differently in dashboards - errors are usually environment / infra; failures are usually code or fixture drift.
Step 2 - Parse safely
Use a streaming parser for large files (multi-thousand-test suites are common). Python core:
# scripts/parse_junit.py
import xml.etree.ElementTree as ET
def parse_junit(path):
tree = ET.parse(path)
root = tree.getroot()
suites = root.findall('testsuite') if root.tag == 'testsuites' else [root]
for suite in suites:
for case in suite.findall('testcase'):
fault = case.find('failure')
if fault is None:
fault = case.find('error')
yield {
'suite': suite.get('name'),
'classname': case.get('classname'),
'name': case.get('name'),
'time': float(case.get('time') or 0),
'status': classify(case),
'failure_message': fault.get('message') if fault is not None else None,
}
def classify(case):
if case.find('failure') is not None: return 'failure'
if case.find('error') is not None: return 'error'
if case.find('skipped') is not None: return 'skipped'
return 'pass'An Element with no children is falsy, so case.find('failure') or case.find('error') would skip a childless <failure>; test the nodes with is None instead.
Always handle both root shapes: the root may be <testsuites> or a bare <testsuite>. The Node.js (fast-xml-parser) equivalent, which also has to undo single-element collapsing (one testcase = bare object, multiple = array), is in references/junit-xml-parsing.md.
Step 3 - Distinguish new failures from flakes
Per llg-junit (opens in new window), the schema "supports modern variants including <flakyFailure>, <flakyError>, <rerunFailure>, and <rerunError> elements for additional test run metadata."
When the runner does automatic retries (Maven Surefire's rerunFailingTestsCount, pytest-rerunfailures, etc.):
Classification:
def reliability(case):
has_flaky = case.find('flakyFailure') is not None or case.find('flakyError') is not None
has_rerun = case.find('rerunFailure') is not None or case.find('rerunError') is not None
has_final = case.find('failure') is not None or case.find('error') is not None
if has_flaky and not has_final: return 'flaky' # passed on retry
if has_rerun and has_final: return 'consistently_failing'
if has_final: return 'newly_failed'
return 'pass'Surface flaky tests in a separate report - they're noise to the PR author but signal to the test-suite owner.
Step 4 - Aggregate per-suite metrics
from collections import defaultdict
def per_suite(cases):
agg = defaultdict(lambda: {'pass': 0, 'failure': 0, 'error': 0, 'skipped': 0, 'flaky': 0, 'time': 0.0})
for c in cases:
agg[c['suite']][c['status']] += 1
agg[c['suite']]['time'] += c['time']
return aggStep 5 - Trend analysis (cross-run)
To detect "is this a new failure or has this test been failing for a week?", store every run's parsed metrics in a per-suite history file:
{"sha":"abc123","ts":"2026-05-05T14:00:00Z","suite":"checkout","failure":2,"flaky":1,"time":12.4}
{"sha":"def456","ts":"2026-05-05T14:30:00Z","suite":"checkout","failure":2,"flaky":0,"time":12.1}Compare by suite + classname:
| classname | name | last 5 runs result | first failed sha |
|---|---|---|---|
cart.CartTest | addItem_validatesStock | F F F F F | abc123 (5 days ago) |
checkout.PromoTest | applyPromo_caseInsensitive | P P P P F | this PR (suspected regression) |
The first row is a stale failure; the second is a probable regression.
Step 6 - Per-case slow-test list
Sort testcases by time descending. The top 1% is the fast feedback target - moving any one of them from 30s → 3s saves more than refactoring a hundred tests that already run in <100ms.
Step 7 - CI integration
Run tests with a JUnit reporter enabled, then parse and upload the results on if: always() - JUnit XML matters most on failed runs, so a step gated on success would drop exactly the data you need. The full GitHub Actions workflow (reporter env var, parse step, artifact upload) is in references/junit-xml-parsing.md.
Anti-patterns
| Anti-pattern | Why it fails | Fix |
|---|---|---|
Treating <error> and <failure> as the same | Errors are usually infra (DB connection lost), failures are usually code. Conflating hides root-cause patterns. | Group them separately. |
Dropping <flakyFailure> reports from the dashboard | Hidden flake budget; quality erodes silently. | Surface flaky tests on a separate panel; assign owner. |
Loading multi-MB XML with xml.dom.minidom.parseString | Whole-tree-in-memory. OOM on large suites. | xml.etree.ElementTree.iterparse for streaming. |
Failing the build on any <skipped> count > 0 | Many runners legitimately skip (platform-gated, conditional). | Skip is informational; only fail on failure / error. |
Hardcoding <testsuites> as the root | Some runners emit a single <testsuite> as the root. | Detect both shapes (Step 2). |
Trusting time for sub-millisecond tests | Some runners emit 0 for any test under their granularity; sort breaks. | Treat time = 0 as "not measured"; don't include in slow-test list. |
Cross-suite aggregation by name alone | Two suites can have a it('renders') each - merging false-flags both. | Always group by (classname, name) tuple. |
Limitations
References
JUnit XML: Node.js parser and CI wiring
View source (opens in new window)JUnit XML: Node.js parser and CI wiring
Companion detail for junit-xml-analysis. The Python parse_junit.py in SKILL.md is the runnable core; this file holds the Node.js equivalent and the CI workflow.
Node.js parser
import { XMLParser } from 'fast-xml-parser';
import { readFileSync } from 'node:fs';
const parser = new XMLParser({ ignoreAttributes: false, attributeNamePrefix: '@_' });
const xml = parser.parse(readFileSync(path, 'utf8'));
const suites = xml.testsuites
? (Array.isArray(xml.testsuites.testsuite) ? xml.testsuites.testsuite : [xml.testsuites.testsuite])
: [xml.testsuite];
for (const suite of suites) {
const cases = Array.isArray(suite.testcase) ? suite.testcase : [suite.testcase];
// ...
}Handle both root shapes (<testsuites> or a bare <testsuite>) and single-element collapsing (one testcase = bare object, multiple = array), which is common in JS XML libraries.
CI integration
# .github/workflows/test-analytics.yml
- name: Run tests (any framework, JUnit XML reporter enabled)
run: npm test -- --reporters=default,jest-junit
env:
JEST_JUNIT_OUTPUT_FILE: junit.xml
- name: Analyze JUnit XML
if: always()
run: python scripts/parse_junit.py junit.xml > analytics.json
- name: Upload analytics
if: always()
uses: actions/upload-artifact@v4
with:
name: junit-analytics
path: |
junit.xml
analytics.jsonif: always() is critical - JUnit XML matters most on failed runs.
Related skills
allure-reports
Configures Allure Report (test-runner adapter install, `allure-results` directory wiring, `categories.json` for failure classification, `history-trend.json` retention via the copy-history-between-runs pattern), runs the Allure CLI to convert `allure-results` to a static HTML site, and uploads the report as a CI artifact. Use when the team needs richer test reporting than JUnit XML - step-level attachments, per-test history, retry tracking, and severity / epic / feature labeling across framework-agnostic adapters (pytest, Jest, JUnit, TestNG, NUnit, Mocha). As a rich static HTML report generator, it is the open-source alternative to extentreports (JVM/.NET per-test HTML narrative); for hosted cross-run flakiness analytics rather than a static per-run report use currents-integration.
cobertura-analysis
Parses Cobertura XML coverage reports (the JVM-canonical format originally from the cobertura-cobertura tool, also emitted by JaCoCo `--coverage-xml`, coverage.py `--xml`, Istanbul / Jest `cobertura` reporter, gocover-cobertura, and dotnet's `coverlet`). Walks the coverage-04 DTD structure (coverage → packages → classes → methods → lines + conditions), computes per-file deltas, and emits PR-time gating verdicts. Use when the existing CI emits Cobertura XML - typical for JVM-heavy stacks and tools that ship Cobertura as a default reporter.
coverage-diff-reporter
Builds a per-PR coverage delta report from any pair of LCOV / Cobertura / JSON coverage outputs (current run + baseline from the merge target) - emits a per-file table with line% / branch% deltas, called-out new files, hidden drops (overall +0.1pp but one file -8pp), and a single-line PR-comment summary. Use when the team has coverage in CI but needs human-readable PR feedback that points at the specific file the reviewer should focus on, not just an aggregate number.
coverage-py-analysis
Configures coverage.py for Python projects - wires `coverage run` (replacing `python` for instrumentation), enables branch coverage via the `--branch` flag or `branch = True` config, manages the `.coverage` data file (single-process and `combine` for parallel pytest-xdist runs), authors `.coveragerc` with `source` / `omit` / `fail_under`, and emits the format the downstream tool needs (`coverage report` for terminal, `coverage xml` for Cobertura, `coverage html` for human review, `coverage lcov` for SaaS, `coverage json` for programmatic post-processing). Use for any Python test stack (pytest, unittest, nose) that needs PR-time coverage signal.
currents-integration
Wires Currents.dev cross-run test analytics into a Playwright suite: installs `@currents/playwright`, authors `currents.config.ts` (env-sourced `recordKey` + `projectId`), registers `currentsReporter()`, enables trace/video/screenshot artifacts, and runs via `npx pwc` so per-test traces stream to the Currents dashboard with over-time flakiness, slowest-test, and pass-rate trends. Use when a Playwright suite needs hosted cross-run suite-health analytics; for a static per-run report use extentreports or allure-reports, and to sync results into Jira test management use zephyr-integration or xray-integration.
extentreports
Configures ExtentReports v5 for a JVM (or .NET via `extentreports-dotnet`) test run: wires `ExtentSparkReporter`, `attachReporter`, `createTest`, the `info`/`pass`/`warning`/`skip`/`fail` log chain, screenshots via `MediaEntityBuilder`, hierarchical `createNode` parent/child tests, and category/author/device labels, emitting a static HTML report alongside JUnit XML for CI artifact upload. Use when a suite on the Aventstack ExtentReports stack wants a richer per-test HTML narrative than JUnit XML gives; for code-coverage reporting use jacoco-analysis, and for hosted cross-run flakiness analytics use currents-integration.
jacoco-analysis
Configures JaCoCo for JVM projects (Java / Kotlin / Scala / Groovy) - wires the runtime agent via `jacoco-maven-plugin` `prepare-agent`, generates per-build reports (HTML / XML / CSV) via the `report` goal, gates the build via the `check` goal with element / limit / minimum rules, parses the six native counters (instructions, branches, lines, methods, classes, cyclomatic complexity), and converts JaCoCo XML to LCOV / Cobertura when downstream tools need a different format. Use when the JVM build is Maven / Gradle and the team wants the canonical JVM coverage tool - or to convert JaCoCo output for cross-language coverage aggregation.
jest-coverage-analysis
Configures Jest's built-in coverage (Istanbul-instrumented `babel` provider or V8-native `v8` provider), wires the right `coverageReporters` for downstream consumption (`lcov` for SaaS / cross-tool, `cobertura` for Jenkins, `text-summary` for terminal, `html` for human review), authors per-file `coverageThreshold` rules that focus the gate on critical paths (vs the global-only foot-gun), and parses the per-file JSON output for PR-time deltas. Use when the project tests with Jest (or Vitest, which uses the same Istanbul/V8 provider) and the team needs PR-time coverage signal that's both local-runnable and CI-gateable.
lcov-analysis
Parses LCOV `.info` text files (the de-facto coverage interchange format produced by gcov, llvm-cov, Coverage.py via `py2lcov`, JaCoCo via `xml2lcov`, Devel::Cover, Jest via `lcov` reporter, NYC, and most others). Extracts per-file line / function / branch metrics from the canonical record keywords (TN/SF/FN/FNDA/FNF/FNH/BRDA/BRF/BRH/DA/LH/LF), computes the diff vs a baseline, and emits per-file gating verdicts. Use for PR coverage gates that don't depend on a specific language runtime.
test-coverage-targeter
Builds a "what to test next" recommendation by combining a coverage report (LCOV / Cobertura / coverage.py JSON / Jest JSON / JaCoCo XML) with the PR's `git diff`, ranking uncovered branches by risk × cost - risk weighted by McCabe cyclomatic complexity and code-churn frequency, cost weighted by the unit-test pyramid layer (unit tests cheaper than integration than E2E). Emits a prioritized list with concrete file:line targets and the test layer recommended for each. Use when a team has the budget to write 5 - 10 new tests and needs help picking which uncovered code to target first instead of blindly chasing 100% coverage.
test-run-summary-author
Build-an-X workflow that turns a structured test-run artifact (JUnit XML, Allure JSON, TestRail / Xray / Zephyr export) plus optional release context (version, build URL, deploy target) into a narrative markdown summary for release notes, an exec status update, or a stand-up Slack post. Distinct from the per-framework parsers junit-xml-analysis / allure-reports / coverage-diff-reporter, which emit structured tabular reports: this skill takes the same data and writes the human-readable narrative. Use when a manager needs a draft release note or stand-up summary from a single run; for cross-run trend analytics use currents-integration.
testrail-integration
Syncs test runs / results / cases between an automated test suite and TestRail (Gurock / Idera) - opens a Test Run for the build (`add_run`), batches per-case results back via `add_results_for_cases` (preferred over per-test `add_result_for_case` - N+1 API calls vs 1), maps the test framework's pass/fail/skip to TestRail status IDs, and attaches build URL + version + elapsed time. Use when the team's test management is standalone TestRail (not a Jira app) and automated suites must update it without a human copy-paste step; when the TCM is instead a Jira app use xray-integration (Xray) or zephyr-integration (Zephyr Scale), and for hosted cross-run flakiness analytics rather than TCM sync use currents-integration.
xray-integration
Imports CI test results into Xray for Jira - authenticates via the `client_id` + `client_secret` → JWT exchange (Cloud) or PAT / Basic (Server), posts to the format-specific `/api/v2/import/execution/*` endpoint (`/junit` for JUnit XML, `/cucumber` for Cucumber JSON, `/nunit` / `/testng` / `/robot` for the others), and maps automated results to existing Xray Test issues via the `xray-junit-extensions` `@XrayTest(key="...")` annotation. Use when the team uses the Xray Jira app to manage Test, Test Set, and Test Execution issue types and CI must keep those execution issues in sync; for the other Jira TCM app use zephyr-integration (Zephyr Scale), for standalone non-Jira TestRail use testrail-integration, and for hosted cross-run flakiness analytics rather than TCM sync use currents-integration.
zephyr-integration
Syncs automated test results to Zephyr Scale for Jira (formerly TM4J / SmartBear / Adaptavist): picks the product variant (Scale Cloud / Squad / Enterprise), authenticates with a long-lived API token as a Bearer header, opens a Test Cycle per build, posts executions via `POST /testexecutions` (or bulk JUnit via `/automations/executions/junit`), and maps test methods to Zephyr Test Cases via `@TestCaseKey`-style annotations. Use when the team's Jira test management is Zephyr Scale; for the Xray Jira app use xray-integration, for standalone TestRail use testrail-integration, and for over-time flakiness analytics rather than TCM sync use currents-integration.