visual-diff-summarizer
Triages a PR's visual regression diffs when there are too many changed screenshots to review one by one. Clusters snapshots (from Percy, Chromatic, Playwright `toHaveScreenshot`, Storybook, Loki) by component / route, separates changes that match PR intent from cascade / regression suspects, recommends which baselines to update, and emits one PR comment pointing the reviewer at the screenshots that need actual eyes. Use when a PR has 20+ visual diffs / changed screenshots and the reviewer needs help deciding which to open.
Install with skills.sh (any agent)
npx skills add testland/qa --skill visual-diff-summarizervisual-diff-summarizer
Overview
A 50-diff visual review causes diff blindness - reviewers rubber-stamp or skip. This skill turns "50 diffs" into "3 components changed as the PR intended; 1 component changed unexpectedly - focus here":
The intent-vs-diff classification mirrors the same logic used by snapshot golden-file management and visual-diff classification - a wrong-but-consistent visual baseline is worse than no baseline at all.
When to use
If the PR has 1 - 2 diffs in scoped files, this skill is overkill - the reviewer can open them directly. The value compounds at scale.
Step 1 - Pick the upstream tool's output
Each tool exposes per-snapshot diff data (a diff_ratio per snapshot) via API or local artifact - see references/tool-outputs.md for the source per tool. This skill is downstream of the qa-visual-regression per-tool wrappers and consumes their output.
Step 2 - Normalize to per-snapshot rows
interface SnapshotDiff {
tool: 'percy' | 'chromatic' | 'playwright' | 'storybook' | 'loki';
storyOrPage: string; // e.g. "Button/with-icon" or "/checkout"
componentOrRoute: string; // derived: "Button" or "checkout"
variant?: string; // e.g. viewport, theme, hover state
beforeImageUrl?: string;
afterImageUrl?: string;
diffImageUrl?: string;
diffRatio: number; // 0..1, fraction of pixels that changed
baselineSnapshotId?: string;
}Cluster keys:
function clusterKey(d: SnapshotDiff): string {
return d.componentOrRoute; // "Button", "checkout", etc.
}Step 3 - Read PR intent
The PR's title + description + labels is the stated intent signal. Pull via gh pr view:
gh pr view --json title,body,labels,files,headRefOidExtract structural intent:
interface Intent {
title: string;
body: string;
labels: string[];
changedFiles: string[]; // src paths
// Derived:
scopeKeywords: string[]; // e.g. ["Button", "checkout", "design-tokens"]
changedComponents: Set<string>; // mapped from changedFiles
}scopeKeywords extraction:
Step 4 - Classify each cluster against intent
type Classification = 'aligned' | 'adjacent' | 'unrelated';
function classify(cluster: string, diffs: SnapshotDiff[], intent: Intent): Classification {
if (intent.scopeKeywords.includes(cluster) ||
intent.changedComponents.has(cluster)) {
return 'aligned';
}
// Adjacent: cluster is a child / parent of a changed component (e.g. PR touches
// Modal; Button-inside-Modal also diffs).
if (intent.changedComponents.has(parentOf(cluster)) ||
intent.changedComponents.has(childOf(cluster))) {
return 'adjacent';
}
return 'unrelated';
}parentOf/childOf use the team's component-graph manifest (from Storybook or a hand-maintained JSON) to identify hierarchy.
Step 5 - Render the report
Emit a sticky PR comment grouping clusters aligned → adjacent → unrelated, each with a per-cluster recommendation and a "Quick actions" block that auto-approves only the aligned cluster. The full report template (headers, per-cluster tables, quick-action commands) is in references/report-format.md.
Step 6 - Cluster sort order
The reviewer scans top-down. Order:
Inside each group, sort by max diff ratio descending - biggest visual change first.
Step 7 - Verify, then auto-update only the aligned cluster
Checkpoint before approving: open one sample diff per aligned cluster and confirm it matches the PR's stated intent; only then run the auto-accept command. The summary includes the safe-to-run command for the aligned cluster only - adjacent and unrelated clusters never get an auto-update suggestion.
This matches the adversarial logic of per-snapshot visual-diff classification: aligned diffs go through; unrelated diffs are refused with a recommendation to escalate to a regression bisect.
Step 8 - CI integration (sticky comment)
- name: Fetch tool diff data
run: |
npx chromatic --dry-run --json > visual-diffs.json
- name: Generate summary
run: python scripts/visual_summary.py visual-diffs.json --pr ${{ github.event.pull_request.number }} > summary.md
- name: Post sticky comment
uses: marocchino/sticky-pull-request-comment@v2
with:
header: visual-diff-summary
path: summary.mdAnti-patterns
| Anti-pattern | Why it fails | Fix |
|---|---|---|
| Posting one comment per snapshot | PR conversation is buried; reviewer can't see the forest. | One sticky summary (Step 8); per-snapshot detail in linked tool UI. |
| No intent classification - just listing diffs | Reviewer has to figure out what's expected vs surprise; same problem as no summary. | Aligned / adjacent / unrelated buckets (Step 4). |
| Auto-approving every diff | Regressions silently become baselines. The point of visual tests is defeated. | Auto-approve only aligned cluster (Step 7); refuse unrelated. |
| Sorting by alphabetical name | High-impact diffs sit below low-impact alphabetical earlier-letter ones. | Sort by max diff ratio within each classification group (Step 6). |
| Cross-tool deduplication missing | Same snapshot appears in 3 reports (Percy + Playwright + Chromatic dual-instrumentation). | Deduplicate by (componentOrRoute, variant) cluster key. |
| Reporting unchanged snapshots | Hides the few that did change; reviewer scrolls past. | Report only changed; unchanged count in the summary header. |
| Treating "diff% = 0.1" as a real diff | Anti-aliasing / font rendering jitter; not a visual change. | Threshold (typically diff% > 0.5 or pixel count > 10) before counting. |
Limitations
References
Visual-diff summary report format
View source (opens in new window)Visual-diff summary report format
The sticky PR comment emitted by SKILL.md Step 5. Clusters are grouped aligned → adjacent → unrelated, sorted by max diff ratio within each group.
## Visual diff summary - `<sha>`
**Total snapshots:** 87 (12 changed, 75 unchanged)
**Verdict:** REVIEW (1 unrelated cluster suspects regression)
### ✅ Aligned with PR intent (3 clusters, 7 diffs)
The PR title says **"Refactor Button to use new design tokens"** - these clusters match.
| Cluster | Diffs | Max diff% | Recommendation |
|---------|------:|----------:|----------------|
| Button | 4 | 8.2% | Update baselines (accept new snapshots) |
| ButtonGroup | 2 | 3.1% | Update baselines |
| IconButton | 1 | 1.5% | Update baseline |
### ⚠ Adjacent (1 cluster, 3 diffs) - confirm intent
| Cluster | Diffs | Max diff% | Recommendation |
|---------|------:|----------:|----------------|
| Modal | 3 | 2.8% | Modal contains Button; check that the Button color change inside Modal is intended (it should be - but eyeball one). |
### ❌ Unrelated (1 cluster, 2 diffs) - DO NOT update without investigation
| Cluster | Diffs | Max diff% | Recommendation |
|---------|------:|----------:|----------------|
| Footer | 2 | 12.0% | The Footer component isn't mentioned in the PR. Suspected unintended cascade. Open the diffs and run a regression bisect if no obvious cause. |The report ends with a Quick actions block that auto-approves only the aligned cluster - adjacent and unrelated clusters are left for manual review:
# Update aligned baselines after eyeballing 1 sample per cluster:
chromatic --auto-accept-changes --only-changed --components Button,ButtonGroup,IconButton
# OR for Percy:
percy approve <build-id> --snapshots Button,ButtonGroup,IconButton
# Refused - Footer cluster needs investigation; do NOT auto-approve.Per-tool diff-data sources
View source (opens in new window)Per-tool diff-data sources
Where each visual tool exposes per-snapshot diff data for SKILL.md Step 1 to consume. This skill is downstream of the per-tool wrappers.
| Tool | Where the diff data lives |
|---|---|
| Percy (BrowserStack) | Build API: GET /api/v1/builds/<id>/snapshots; per-snapshot diff_ratio. |
| Chromatic | chromatic --dry-run JSON; --exit-zero-on-changes build summary. |
Playwright toHaveScreenshot | Test reporter output; failed expectations include attached image diffs. |
| Storybook test-runner | Per-story coverage diff via @storybook/test-runner's snapshot mode. |
| Loki / BackstopJS | JSON report with per-scenario misMatchPercentage. |
The upstream tool wrappers in qa-visual-regression (percy-visual-regression-testing, chromatic-visual-regression-testing, playwright-snapshots, storybook-visual-regression-testing) cover the per-tool integration and produce the output this skill consumes.
Related skills
chromatic-visual-regression-testing
Authors and runs Chromatic visual tests on Storybook, Playwright, or Cypress projects via the `chromatic` CLI; configures baselines, TurboSnap, UI Review, and CI gating; reads exit codes for change-vs-error classification. Use when the project ships visual regression coverage to Chromatic Cloud.
percy-visual-regression-testing
Authors Percy visual snapshot tests via the @percy/cli + framework SDK (Playwright, Cypress, Selenium, Storybook), runs them with `percy exec -- {test command}`, configures viewports / masking / ignored regions, and reviews diffs in the Percy build UI. Use when the project ships visual regression coverage to BrowserStack Percy.
playwright-snapshots
Authors Playwright `expect(page).toHaveScreenshot()` assertions, configures masks / clips / threshold / maxDiffPixels per test, manages the per-OS / per-browser snapshot directory, and runs the update flow with `--update-snapshots`. Use when the project ships self-hosted visual regression coverage in Playwright (no external snapshot service).
responsive-breakpoint-runner
Produces a single breakpoint-matrix report (rows = pages/stories, columns = viewports) across Percy, Chromatic, Playwright snapshots, or Storybook test-runner. Routes per-engine viewport syntax, runs each, and aggregates the results into one cross-breakpoint view. Use when the team needs one unified pass/fail view across three or more viewport widths instead of separate per-engine or per-breakpoint reports. The matrix-view output is the distinguishing trait: it dispatches to playwright-snapshots (and the other engines) rather than replacing any single-engine skill.
storybook-visual-regression-testing
Sets up visual regression coverage for a Storybook project - either via the official @chromatic-com/storybook addon (hosted) or via @storybook/test-runner with a postVisit hook that calls Playwright's toHaveScreenshot (self-hosted). Covers test-runner install, lifecycle hooks (setup / preVisit / postVisit), and CI integration. Use when a repo already has a working `.storybook/` config and the team wants per-story visual coverage rather than page-level snapshots.
visual-baseline-conventions
Reference catalog for visual regression coverage decisions - which Storybook stories or pages get baselines, how to choose breakpoints, when to mask vs adjust threshold, when to add or remove a baseline, and a decision matrix for picking among Percy / Chromatic / Playwright / Storybook test-runner. Use when designing visual coverage for a new project or auditing an existing baseline set.
visual-baseline-gate
Consumes pre-classified visual-diff JSON and a reviewer-signed acceptance log to produce a single go/no-go CI verdict for visual regression. Blocks when intentional baseline changes lack a non-author reviewer sign-off or when regressions are present, and emits a markdown + JSON artifact for the CI step. Use this skill when the gate's input is pre-classified diff data and the enforcement concern is reviewer approval, not when the goal is fanning out to multiple engines (use a multi-engine CI orchestrator for that).