Playwright · 10 min read

Playwright Visual Tests That Pass in CI for the Right Reasons

Fix failing Playwright visual regression checks in CI: a four-step triage with tolerance last, and CI-generated baselines approved by named code owners.

On Playwright's default 1280x720 viewport (921,600 pixels), a maxDiffPixels allowance of 5,000 is 0.54% of the image, against 100 in the docs' own example. A change smaller than the allowance passes wherever it is.

Screenshot baselines recorded on a laptop tend to fail Playwright visual regression checks (toHaveScreenshot) in GitHub Actions. The two cheapest ways to turn that check green prove nothing: loosen the tolerance until the diff disappears, or let the author of the UI change re-record the baseline. A visual check earns its CI minutes only when it passes without hiding the change it exists to catch. That takes a fixed triage order (match the CI environment, stabilize the page, mask or scope volatile regions, set tolerance last) and one review rule: a CI job in the tests' container image generates new baselines, and a code owner who didn't write the change approves them.

Table of contents

This post assumes:

  • A TypeScript Playwright Test suite running visual specs on pull requests in GitHub Actions (verified against Playwright 1.64; changed and pathTemplate need 1.50 or later).
  • At least one toHaveScreenshot assertion with baselines committed from a laptop.
  • Permission to add a workflow and a CODEOWNERS file (paths mapped to reviewers) on the default branch and to turn on "Require review from Code Owners", or an admin who will; check that your plan offers it.
  • A visible GitHub team of two or more people, with explicit write access, to own the screenshots folder.
  • The GitHub CLI (gh), installed and authenticated.

Why a laptop baseline fails in GitHub Actions

A baseline, the committed reference image toHaveScreenshot compares against, is only valid for the environment that rendered it. Playwright's snapshot docs warn: "Browser rendering can vary based on the host OS, version, settings, hardware, power source (battery vs. power adapter), headless mode, and other factors."

The default filename carries the platform: in example-test-1-chromium-darwin.png, chromium-darwin is "the browser name and the platform", and {platform} is "The value of process.platform." On a Linux runner, the failures to expect first are:

  • A Linux run doesn't look up your -darwin.png file, so the baseline is missing. The docs' sample reads "Error: A snapshot doesn't exist at example.spec.ts-snapshots/example-test-1-chromium-darwin.png, writing actual." With no flag (default mode), "Missing snapshots are created, but the tests that create them fail, so that the run does not silently pass in CI."
  • A Linux baseline rendered on another machine or an older image fails with a diff over tolerance. With {platform} dropped from the path (the setup adopted below), a baseline rendered on a Mac lands where the Linux run looks, so only the Playwright container image may write baselines here.

Read the diff in the Trace Viewer, then work four steps in order

Read the image diff first: with retries enabled and trace: 'on-first-retry', the HTML report's trace icon opens the Trace Viewer, where you can "compare screenshots by examining the image diff, the actual image and the expected image."

Steps 1 and 2 keep every region in the comparison; step 1 goes first because nothing else can be judged while the environment differs. Step 3 hides named regions that a reviewer can read in the spec. Step 4 hides small changes anywhere, including in the thing under test, with nothing in the diff showing where, so tolerance comes last. The order of steps 2 to 4 is this guide's.

These symptoms are heuristics, not Playwright error text.

StepSymptomFix
1. EnvironmentDifferences spread across the page, or a missing snapshotCompare in the Playwright container image; generate baselines there
2. PageAssertion times out; the page keeps changingWait for the final state; stop motion and live data
3. Mask or scopeOne region differs: a date, an avatar, an ad slotMask it, or screenshot the component
4. ToleranceA few stray pixels remain after steps 1 to 3Set a pixel count on that assertion

Step 1: run the test job in the Playwright container image

Run the test job in the Playwright container image (called the "CI image" from here on); the update job below uses the same one, so baselines are written and checked by one renderer. Laptop baselines won't match inside this image either, so regenerate them there (moving existing laptop baselines) before judging steps 2 to 4. The snapshot docs put it directly: "For consistent screenshots, run tests in the same environment where the baseline screenshots were generated."

The image tag must match the installed @playwright/test version, and the Docker guide covers why, along with the --init --ipc=host options that Playwright's Docker page recommends and its CI example omits. The report upload uses if: ${{ !cancelled() }}, so it runs even after a failed test step.

# .github/workflows/playwright.yml (verified against Playwright 1.64)
name: Playwright Tests
on:
  pull_request:
permissions:
  contents: read
jobs:
  playwright:
    runs-on: ubuntu-latest
    container:
      image: mcr.microsoft.com/playwright:v1.64.0-noble
      options: --init --ipc=host --user 1001
    steps:
      - uses: actions/checkout@v6
      - uses: actions/setup-node@v6
        with:
          node-version: lts/*
      - run: npm ci
      - run: npx playwright test
      - uses: actions/upload-artifact@v7
        if: ${{ !cancelled() }}
        with:
          name: playwright-report
          path: playwright-report/ # written by reporter: 'html'
          retention-days: 30

Step 2: stabilize the page until two screenshots match

A page that never holds still fails toHaveScreenshot with a timeout, not a pixel diff. The toHaveScreenshot reference says: "This function will wait until two consecutive page screenshots yield the same result, and then compare the last screenshot with the expectation." A screenshot counts as stable when two consecutive captures match.

Motion that never settles can't produce two identical captures before timeout runs out. caret defaults to "hide", and animations to "disabled", which "stops CSS animations, CSS transitions and Web Animations"; it names those three, so script-driven motion, video, and live data are the test's job: stop them at the source or wait for the final state. For timing failures that aren't pixel diffs, see the flaky-test guide.

Hover effects are captured too, so the snapshot docs move the mouse away before taking the screenshot:

await page.goto('/pricing');
await page.mouse.move(-1, -1);
await expect(page).toHaveScreenshot('pricing.png');

Step 3: mask volatile regions and scope the screenshot

mask and stylePath should hide only volatile regions, never the thing under test. mask takes locators: "Specify locators that should be masked when the screenshot is taken." Per the assertion reference, masked elements are "overlaid with a pink box #FF00FF", so a difference under the overlay can't fail the check. Masking the price cell to silence a diff also silences a pricing regression.

A locator-level toHaveScreenshot captures one component instead of the page, so unrelated regions can't fail it. stylePath applies a custom stylesheet while the screenshot is taken; the snapshot docs say "This allows filtering out dynamic or volatile elements, hence improving the screenshot determinism."

import path from 'path'; // in an ESM project, use import.meta.dirname instead of __dirname

await expect(page).toHaveScreenshot('dashboard.png', {
  mask: [page.getByTestId('build-timestamp')],
});
await expect(page.getByRole('navigation')).toHaveScreenshot('nav.png');
await expect(page).toHaveScreenshot({ stylePath: path.join(__dirname, 'screenshot.css') });
/* screenshot.css */
iframe { visibility: hidden; }

Step 4: set tolerance last, sized against the default viewport

Tolerance goes last because it hides small changes anywhere on the page. threshold works per pixel: "An acceptable perceived color difference in the YIQ color space between the same pixel in compared images, between zero (strict) and one (lax)", with a default of 0.2. maxDiffPixels counts pixels allowed to differ and is "Unset by default." maxDiffPixelRatio sets that allowance as a fraction of all pixels, between 0 and 1.

Size the allowance against the screenshot: the default viewport is 1280x720, so with the default scale ("css") and fullPage (false), 5,000 pixels is 0.54% of a 921,600-pixel image. A change touching fewer pixels than the allowance passes wherever it is, including in the thing under test. Measure the value from a diff you've looked at.

Set it per assertion, not suite-wide in expect.toHaveScreenshot: a suite-wide value loosens every screenshot from one config line, while a per-assertion value sits beside the one screenshot it excuses.

// 5,000 is illustrative; the docs' own example uses maxDiffPixels: 100.
await expect(page).toHaveScreenshot('report.png', { maxDiffPixels: 5000 });

Updating baselines without approving your own diff

Baselines should come from a job in the CI image, approved by a code owner who isn't the author: the job fixes the environment, and the approval keeps authors from approving their own diffs. That takes two pull requests: a workflow_dispatch workflow runs only if its file is on the default branch, and code owners get review requests only when CODEOWNERS is on the base branch.

Pull request 1 adds the update workflow (it runs @visual specs, tagged in pull request 2), the CODEOWNERS line, and the "Require review from Code Owners" setting. The job never pushes, so it needs no write permission, and GITHUB_TOKEN pushes don't trigger push workflows.

# .github/workflows/update-visual-baselines.yml
name: Update visual baselines
on:
  workflow_dispatch:
permissions:
  contents: read
jobs:
  update:
    runs-on: ubuntu-latest
    container:
      image: mcr.microsoft.com/playwright:v1.64.0-noble
      options: --init --ipc=host --user 1001
    steps:
      - uses: actions/checkout@v6
      - uses: actions/setup-node@v6
        with:
          node-version: lts/*
      - run: npm ci
      - run: npx playwright test --grep @visual --update-snapshots=changed
      - uses: actions/upload-artifact@v7
        if: ${{ !cancelled() }}
        with:
          name: visual-baselines
          path: tests/__screenshots__/**/*.png
          if-no-files-found: error

The CODEOWNERS line works because "Pull request authors cannot approve their own pull requests." and "an approval from any of the owners is sufficient", so a team of two or more always has a non-author approver. Repository owners and administrators are the exception: they "can merge a pull request even if it hasn't received an approving review".

# .github/CODEOWNERS
/tests/__screenshots__/ @your-org/visual-reviewers

It covers only tests/__screenshots__/, which the next section's path template creates. Specs, playwright.config.ts, and the workflow sit outside it, so a raised tolerance is an ordinary code change. To protect the repository fully, "you also need to define an owner for the CODEOWNERS file itself." Solo repos keep the CI half but can't meet the approval half.

Moving existing laptop baselines into the CI image

Pull request 2 does the move:

  1. Tag the visual specs so --grep @visual selects them.
  2. Add the path template, which sets where baselines are written. It deliberately drops {platform} from the default naming, since baselines are meant to come only from the CI image, and keeps one folder per project. If testDir isn't ./tests or the config isn't at the repository root, change tests/__screenshots__ to match in the CODEOWNERS line, the upload step, and the download command.
  3. Delete the laptop baselines (*-darwin.png, *-win32.png, the old *-snapshots/ folders).
  4. Dispatch the job on the branch, download the artifact, and commit the PNGs. Until they land, the test job fails on missing snapshots.
// In each visual spec
test('pricing page', { tag: '@visual' }, async ({ page }) => {
  // ...
});
// playwright.config.ts
import { defineConfig } from '@playwright/test';

export default defineConfig({
  testDir: './tests',
  expect: {
    toHaveScreenshot: {
      pathTemplate: '{testDir}/__screenshots__{/projectName}/{testFilePath}/{arg}{ext}',
    },
  },
});
# Run from the repository root
gh workflow run update-visual-baselines.yml --ref my-branch
gh run list --workflow update-visual-baselines.yml --branch my-branch --limit 1
# Wait for the run to finish (gh run watch <run-id>), then download
gh run download <run-id> -n visual-baselines -D tests/__screenshots__
git status

Use changed, not all: since 1.50, all rewrites every snapshot, so older advice assumes a different meaning.

The owner opens each changed PNG and checks it against the change the pull request describes.

With one file per screenshot, a laptop re-record overwrites the Linux file in place, shows as a modified PNG in the owned folder, and normally fails CI. Never run --update-snapshots locally, and exclude the visual specs from laptop runs, where they would compare against Linux renders.

# Laptop runs skip the visual specs
npx playwright test --grep-invert "@visual"

Every later UI change repeats the dispatch, download, and commit, then the owner's review, and a developer with a red check needs only the first three commands from the four-command block: dispatch, find the run ID, download.

Update modes and the screenshot option at a glance

Verified against Playwright 1.64; version notes inline.

Mode or optionBehaviorVersion note
--update-snapshots=changedRewrites mismatches, creates missingAdded in 1.50; current docs say a bare --update-snapshots means this
--update-snapshots=allRewrites every snapshotBefore 1.50: only failed or changed
--update-snapshots=missingCreates missing; those tests passPassing since 1.64
No flagWrites missing; those tests failNamed 'default' in 1.64
use: { screenshot }Screenshot after each test or failure; defaults to 'off'Not a comparison

Start with pull request 1: the update workflow, the CODEOWNERS line, and the Code Owners setting. To place visual specs in a wider suite, see how to organize regression testing for web apps; the Playwright topic page collects the rest.

Start with qa-starter.

10 plugins every tester needs on any stack: test planning, test data, environments, reporting, flaky tests, bug reports and exploratory testing.

/plugin install qa-starter@testland-qa