Testland
Browse all skills & agents

seed-data-curator

Builds a reproducible E2E seed dataset for the project's test environments - picks a representative user / org / data-product cross-section, generates the rows via the project's chosen factory library (FactoryBot / mimesis / Bogus / Faker + factory_boy), persists the dataset as a checked-in fixture (SQL dump / JSON / per-engine seed file), and wires it into the test bootstrap. Use when starting E2E coverage on a project that has no seed strategy, or when an existing seed has drifted.

Install with skills.sh (any agent)

npx skills add testland/qa --skill seed-data-curator
View source

seed-data-curator

Overview

E2E tests need repeatable starting state - the same set of users, orgs, products, etc. before every run. Generating fresh data per run produces flake; pulling from production raises PII / security concerns; hand-crafted SQL fixtures rot. This skill defines a workflow for building a curated seed set that:

  1. Lives in the repo as a checked-in artifact.
  2. Is regenerable on demand from the factory definitions.
  3. Provides representative coverage of business-relevant states (roles, tiers, status variants).
  4. Refreshes intentionally (PR-reviewed) rather than continuously.

When to use

  • A project is starting E2E coverage and has no seed strategy.
  • An existing seed.sql has accumulated cruft over years and reviewers can't tell what each row represents.
  • The team migrated from one factory library to another and the seed set should be regenerated to match.
  • A new role / tier / feature flag needs to be exercised by E2E tests; the seed needs an example user.

Step 1 - Define the coverage matrix

Enumerate the business-relevant states the seed must cover:

CategoryVariantsWhy
User roleAdmin / Manager / Standard / Read-onlyDifferent UI / permissions per role.
Account tierFree / Starter / Pro / EnterpriseFeature flags / quota differ.
Account stateActive / Suspended / Trial-expired / DeletedTest the lifecycle handlers.
Localeen-US / ja-JP / de-DE / ar-SAi18n / RTL coverage.
Data volumeEmpty / 1 item / 10 items / 100 itemsPagination, empty-states, truncation.
TimeNew (created today) / Old (created 1y ago) / AnniversaryDate-relative business logic.

Pick the minimum cross-product that covers your test surface - not every combination. A typical sweet spot is 12-20 seed users.

Step 2 - Pick the persistence format

FormatWhen to use
SQL dump (seeds/seed.sql)Fastest restore for E2E suites; database-specific.
JSON / YAML fixtures (seeds/users.json)Database-agnostic; consumed by the factory library at boot.
Factory script (scripts/seed.rb / seed.py)Most flexible; runs the factories at boot time.
Per-engine seed file (seeds/snowflake.sql + seeds/postgres.sql)Multi-warehouse projects.

Default: factory script - runs the team's factory library to build the dataset every time the test environment starts. SQL dumps are faster but harder to review; reserve them for very large seed sets.

Step 3 - Author the factory script

Example with FactoryBot (Ruby):

# scripts/seed.rb
require 'factory_bot'
require_relative '../db/factories/all'

# Deterministic seed for reproducibility
Faker::Config.random = Random.new(42)

# Coverage cross-section
roles  = %i[admin manager standard read_only]
tiers  = %i[free starter pro enterprise]
states = %i[active suspended]

users = []
roles.each_with_index do |role, ri|
  tiers.each_with_index do |tier, ti|
    user = FactoryBot.create(
      :user,
      role: role,
      tier: tier,
      state: ri == 0 && ti == 0 ? :suspended : :active,   # one suspended for coverage
      email: "#{role}-#{tier}@example.com",                # predictable for E2E tests to reference
      created_at: ri.zero? ? 1.year.ago : Time.current,
    )
    users << user
  end
end

puts "Seeded #{users.count} users"

The predictable email is intentional - E2E tests reference admin-pro@example.com rather than guessing a Faker-generated email. The factory still uses Faker for non-identifying fields (name, address, phone).

For Python equivalent with factory_boy + mimesis, see the mimesis-data examples; the same deterministic-seed-plus-predictable-email pattern applies.

Step 4 - Wire into test bootstrap

Local development

# Reset DB, run migrations, run seed
bundle exec rake db:test:reset
bundle exec ruby scripts/seed.rb

Expose this as a single command (make seed, npm run seed, yarn seed) so contributors don't memorize the chain.

Verify: before wiring the seed into the suite, assert the row count and PII safety - the factory writes 4 roles × 4 tiers = 16 users:

bundle exec ruby scripts/seed.rb | grep -q "Seeded 16 users" || { echo "seed count wrong"; exit 1; }

If the count is off, fix the factory script (a coverage cell changed) and re-run before proceeding; if any generated email falls outside @example.com, route that field through synthetic-pii-generator and re-seed.

CI and ephemeral-env (Docker Compose) bootstrap wiring: references/seed-scripts.md.

Step 5 - Refresh intentionally

The seed dataset is a maintained artifact, not a one-shot. Refresh when:

  • A new role / tier / feature flag is introduced that needs E2E coverage.
  • A factory definition changes shape (new required column, type change).
  • An NFR review surfaces a missing-coverage gap (e.g. "we never test the suspended-account path").

Refresh process:

  1. Update the factory script.
  2. Run locally; verify the seed produces a sensible output.
  3. Commit the factory script change, not a re-generated seed.sql. Reviewers can re-run from the script if needed.
  4. Update CI's seed command if the script's interface changed.

Output format

When this skill runs, it emits:

## Seed dataset for `<project>`

**Persistence format:** factory script | SQL dump | JSON fixtures
**Total rows:** N (across M tables)
**Coverage matrix:**

| Category | Variants | Count |
|---|---|---:|
| Role | admin / manager / standard / read_only | 4 |
| Tier | free / starter / pro / enterprise | 4 |
| State | active / suspended | 2 |
| **Cross-product covered** | role × tier (16 cells) | 16 users |

**Files added/modified:**
  - scripts/seed.rb (new / modified)
  - db/factories/users.rb (extended for new role variant)

**Re-run command:** `bundle exec ruby scripts/seed.rb`

**Re-generation cadence:** on-demand only; no CI auto-refresh.

### Validation

- Local seed run: success (N users created).
- E2E suite passing against the seed: confirmed.
- No PII in seed (all emails `*@example.com`, all names Faker-generated).

Anti-patterns

Anti-patternWhy it failsFix
Production data dump as the seedPII; legal / compliance risk; data drifts.Always generated; never copied from prod.
Random fresh data per runNon-reproducible failures; "it passed last time" debugging.Deterministic seed (Random.new(42)); predictable identifiers.
10k-row seed for unit testsTest setup time dominates; suite slows linearly.Seed set is for E2E only; unit tests use per-test factories.
Editing seed.sql by handBypasses the factory; drift creates inconsistent state.Never hand-edit; always re-run the factory script.
Seed grows unboundedlyOld role variants no longer used; unclear which rows the tests need.Annual review; remove rows whose tests have been deleted.
One mega-seed for all environmentsDifferent envs need different scale; one monolithic file is wrong for all.Tier the seed: seed:minimal, seed:standard, seed:perf (each loads a different superset).

Limitations

  • Not for performance testing. Perf tests need different volume profiles; use a separate seed (seed:perf with 100k rows) rather than reusing the E2E seed.
  • Locale coverage adds rows fast. 4 roles × 4 tiers × 4 locales = 64 users; review whether all 64 are needed.
  • Seed-vs-migration ordering. If the seed depends on migrations, the migration order must be deterministic; otherwise a seed run on a fresh DB can hit "column not found" errors.

References

Related skills

  • All four factory libraries: faker-data, factory-bot-data, mimesis-data, bogus-data.
  • synthetic-pii-generator - for PII fields in the seed; ensures the seed never carries real-looking PII.
  • golden-file-conventions - sibling reference for snapshot fixtures (similar long-lived-fixture concerns).

Seed bootstrap wiring

The seed runs the same factory script in every environment; only the surrounding orchestration differs. Keep the seed command a single named target (make seed) and call it from each environment.

CI

# .github/workflows/e2e.yml (excerpt)
- name: Set up DB
  run: |
    bundle exec rake db:test:reset
    bundle exec ruby scripts/seed.rb

- name: Run E2E tests
  run: bundle exec rspec spec/system

Reset between test suites - never share state across suites unless the team explicitly designed for it (and accepted the flake risk; see flake-pattern-reference Pattern 2).

Ephemeral env (Docker Compose)

# docker-compose.test.yml (excerpt)
services:
  app:
    build: .
    depends_on:
      db:
        condition: service_healthy
    command: |
      sh -c "
        rake db:migrate &&
        ruby scripts/seed.rb &&
        bundle exec rspec
      "

Related skills

bogus-data

Authors .NET test fixtures using the Bogus library - fluent typed `Faker` builders with `.RuleFor` per property, generation via `Generate()` / `GenerateBetween(min, max)` / `GenerateLazy()`, and `UseSeed()` for reproducibility. Provides the Bogus equivalent of Python's Faker / Ruby's FactoryBot. Use when the project is C# / F# / VB.NET and the team needs typed fixture creation.

boundary-value-generator

Generates boundary-value test cases from typed input specifications - for each input field, produces the canonical 6-point set (one below, at, and above the lower bound; one below, at, and above the upper bound) plus equivalence-class representatives. Emits cases as parameterized test inputs (pytest @parametrize / Jest test.each / xUnit InlineData / etc.). Use when a function or endpoint has numeric / string-length / collection-size constraints and the team needs systematic edge-case coverage.

e2e-test-narrative-builder

Assembles a multi-step end-to-end user-journey test from a list of high-level user intents - translates each intent ("user signs up", "user adds product to cart", "user completes checkout with promo code") into the corresponding test-runner step (Playwright / Cypress / Selenium / Karate), wires shared state across steps via test fixtures, and emits the resulting test as a single Scenario in the project's E2E framework. Use when scaffolding an E2E test that exercises a complete user flow rather than a single page.

factory-bot-data

Authors Ruby FactoryBot factories with traits, associations, sequences, and the three build strategies (build / create / build_stubbed); integrates with RSpec / Minitest test suites; pairs with Faker for randomized field values. Use when the project is Ruby / Rails and needs structured fixture creation with referential integrity.

faker-data

Authors test-data factories using Faker: the Python `faker` library, the `@faker-js/faker` JS port, and the `faker-ruby` gem. Owns the library mechanics end to end: install per language, the provider catalogue (person / internet / location / date / finance / lorem), locale selection and multi-locale mode, and seed-based determinism for reproducible runs. Scope is generating fresh values for tests that start from nothing, not replacing values inside an existing dataset that already holds real records, which raises referential-integrity and re-identification concerns this skill does not address. Prefer this skill when the codebase already uses the Faker family or when cross-language consistency across Python, JS, and Ruby matters; use mimesis-data only when deeper Python locale coverage is the primary requirement. Use when authoring fixtures or factories that need realistic-looking field values.

golden-file-conventions

Reference catalog for snapshot / golden file management - naming conventions, directory layout, when to add / update / remove a baseline, sanitization (timestamps, IDs, PII), per-OS / per-runtime variant strategy, and review workflow for snapshot diffs in PRs. Use when designing a snapshot-testing convention or auditing an existing one for drift.

malicious-payload-bank

Reference catalog of curated adversarial input payloads keyed by attack class - SQL injection, XSS, SSRF, path traversal, command injection, XXE, prototype pollution, regex DoS, Unicode confusables, header injection - plus per-context guidance for which payloads apply (URL parameter / form input / JSON body / file upload). Use when authoring negative-test cases for input validation, fuzz targets, or a security-focused test suite that needs to exercise the OWASP Top 10 attack surface.

mimesis-data

Authors Python test fixtures using mimesis - a fast, type-hinted, locale-aware test-data generator with 46 locales - covering Person / Address / Internet / Datetime providers and the Schema/Field pattern for typed-dict generation. Pairs with factory_boy when referential integrity is needed. Use when the project is Python and the team values speed, type hints, or strong locale coverage over Faker's larger ecosystem.

mountebank-imposters

Authors Mountebank imposters (multi-protocol mock servers - HTTP, HTTPS, TCP, SMTP, LDAP, gRPC, WebSockets, GraphQL, and more) by POSTing JSON definitions to the Mountebank control API on port 2525, configures stubs with predicates and responses, and uses record-playback proxy mode to capture upstream traffic. Use when the project needs a multi-protocol mock server beyond HTTP-only tools like WireMock or MSW.

msw-handlers

Authors Mock Service Worker (MSW) request handlers for both browser and Node.js test environments using the `http.get` / `http.post` / `HttpResponse.json` API, wires them via `setupWorker` (browser) or `setupServer` (Node), and manages the test lifecycle (`server.listen` / `resetHandlers` / `close`). Use when the project uses JavaScript / TypeScript and needs to mock fetch / XHR at the network layer for both Vitest / Jest unit tests and Cypress / Playwright integration tests.

negative-test-generator

Generates negative / error-path test cases that mirror happy-path tests - for each happy-path test, produces companions exercising input validation rejection, missing required fields, type mismatches, authorization failures, rate-limit errors, and adversarial payloads from the malicious-payload-bank. Emits cases as parameterized tests in the project's runner format. Use when a feature has happy-path coverage but the rejection / error / unauthorized paths are untested.

pairwise-test-case-generator

Generates parameterized test inputs combining boundary-value, equivalence-class, and pairwise-combinatorial cases from a typed multi-input specification - produces the cross-product of cases up to a configurable strength (1-wise / 2-wise / N-wise) using all-pairs reduction so the test surface stays tractable. Emits cases in the project's test-runner-native parametrize format. Use when a function or endpoint takes 3+ inputs whose interactions matter and full Cartesian product would explode.

synthetic-data-tool-selector

Chooses between the four mainstream synthetic test-data generators - Faker (JavaScript), FactoryBot (Ruby), mimesis (Python), Bogus (.NET) - picks the right tool by language and use case (raw value generation vs. typed factory orchestration), shows side-by-side equivalents for the same fixture across all four, and emits the language-appropriate code. Use when starting test-data work on a project and the team wants the "which tool should I use" decision documented.

synthetic-pii-generator

Generates realistic-but-fake personally identifiable information (PII) - emails, phone numbers, SSNs / national IDs, addresses, names, credit-card numbers (test BIN ranges), date-of-birth - for non-production environments. Wraps Faker / mimesis with PII-aware constraints so generated values match real format expectations (Luhn-valid card numbers, region-valid phone formats, ITIN/SSN format) without ever generating real-person data. Use when seeding test environments, building demo data, or replacing real PII in copied datasets.

test-data-patterns

Pure reference catalog of the cross-language object-construction patterns for test data - Test Data Builder (Pryce/Freeman), Factory (with traits and associations), Object Mother, Fixture composition (per-test / per-describe / shared), Snapshot (defers to `golden-file-conventions` for the operational details), and Production-Data Anonymisation. Distinct from per-language data wrappers (`factory-bot-data` Ruby, `faker-data` JS, `mimesis-data` Python, `bogus-data` .NET) which document tool-specific configuration; this catalog is the architecture-tier reference for choosing **which pattern** before reaching for the tool. Use when choosing a test-data construction pattern for a new suite, or auditing an existing suite whose fixtures have drifted into shared mutable state.

wiremock-stubs

Authors WireMock stub mappings for HTTP service mocking - `stubFor` with verb/path/header matchers + `willReturn` response shaping, lifecycle via `WireMockServer` (start / stop) or JUnit `WireMockExtension`, request verification via `verify()`, and dynamic-port allocation for parallel tests. Use when the project is JVM-based and tests need to mock HTTP dependencies (third-party APIs, internal microservices) at the network layer.