rum-to-synthetic-gap-analyzer
Reads Real User Monitoring data (Datadog RUM, Sentry Performance, GA4 Core Web Vitals / CrUX) to identify high-traffic user journeys that have no synthetic monitor coverage: ranks journeys by session volume times business value, diffs the ranked list against existing synthetic monitors, and emits a prioritized gap list ready to feed into synthetic-monitor-author. Use when an observability stack has RUM instrumented but the team suspects synthetic coverage is sparse, biased toward low-traffic paths, or was never systematically derived from real usage data.
Install with skills.sh (any agent)
npx skills add testland/qa --skill rum-to-synthetic-gap-analyzerrum-to-synthetic-gap-analyzer
Overview
Synthetic monitors verify journeys that the team chose to script. Real User Monitoring records journeys that users actually take. The gap between the two sets is where production breakage goes undetected: a journey with 40 k sessions per day but no synthetic monitor can fail silently for hours before on-call is paged.
This skill closes that gap. It reads RUM journey data, scores each journey by traffic volume times business value, diffs the result against the team's existing synthetic monitor inventory, and emits a ranked gap list that synthetic-monitor-author can consume directly.
How to use
Step 1 - Collect the RUM journey inventory
Pull the top-N view paths (or transaction names) by session volume from the active RUM source. Aim for the top 50 to avoid chasing long-tail pages that carry negligible traffic.
Per-source query instructions - Datadog RUM Explorer grouping, Sentry throughput sorting, and GA4 + CrUX field data (LCP/INP/CLS at p75) - are in references/rum-source-queries.md.
Step 2 - Score each journey
Assign each journey a coverage-priority score:
coverage_priority = session_volume_score x business_value_scoreSession volume score (1-5): rank by daily session or view count.
| Daily sessions | Score |
|---|---|
| > 10 k | 5 |
| 1 k - 10 k | 4 |
| 100 - 1 k | 3 |
| 10 - 100 | 2 |
| < 10 | 1 |
Business value score (1-5): assign by journey type. Adjust to your domain.
| Journey type | Score |
|---|---|
| Revenue-generating (checkout, upgrade) | 5 |
| Authentication (login, SSO, MFA) | 5 |
| Primary feature (core read/write action) | 4 |
| Onboarding (sign-up, first-run wizard) | 4 |
| Support / self-service (docs, status) | 3 |
| Informational (marketing pages, help) | 2 |
| Admin / internal tooling | 1 |
Score range: 1 (low-traffic, low-value) to 25 (highest-traffic, revenue-critical).
For public-facing sites, supplement volume score with CrUX visit share where available: higher CrUX weight indicates broader real-user exposure.
Step 3 - Build the existing-monitor inventory
Collect the names or URL patterns of every active synthetic monitor. Most platforms expose this via API or config file:
Normalize each monitor to a canonical URL path pattern (strip query strings, replace ID segments with {id}, lowercase). Store as a set.
Step 4 - Diff: rank the gap list
For each journey in the scored list (Step 2), check whether the normalized path matches any pattern in the monitor inventory (Step 3).
gap_list = [j for j in scored_journeys if not matches_any_monitor(j.path)]Sort gap_list descending by coverage_priority. The output is the gap list.
Step 5 - Emit the gap report
Output format (Markdown table, one row per gap):
| Rank | Journey path | Sessions/day | Biz value | Score | Recommended monitor type |
|------|--------------------|--------------|-----------|-------|--------------------------|
| 1 | /dashboard | 12 k | 4 | 20 | Browser (multi-step) |
| 2 | /onboarding/step1 | 5 k | 4 | 16 | Browser (multi-step) |
| 3 | /reports/{id} | 1.2 k | 4 | 12 | Browser (read + assert) |Recommended monitor type heuristic:
Pass the gap list to synthetic-monitor-author as the journey input for Step 1 of that skill.
Worked example
Scored journeys (top 5):
/checkout vol=5, biz=5 -> score 25 [MONITOR EXISTS: checkout-journey.spec.ts]
/dashboard vol=5, biz=4 -> score 20 [NO MONITOR] <-- gap rank 1
/login vol=4, biz=5 -> score 20 [MONITOR EXISTS: auth-flow.spec.ts]
/onboarding/step1 vol=4, biz=4 -> score 16 [NO MONITOR] <-- gap rank 2
/reports/{id} vol=3, biz=4 -> score 12 [NO MONITOR] <-- gap rank 3
Gap list (ready for synthetic-monitor-author):
1. /dashboard score=20 sessions/day=12k business=primary feature
2. /onboarding/step1 score=16 sessions/day=5k business=onboarding
3. /reports/{id} score=12 sessions/day=1.2k business=primary feature/checkout and /login carry the highest scores but already have monitors, so they drop out of the diff. /dashboard surfaces as gap rank 1: 12 k sessions/day, a primary feature, and no monitor - exactly the silent-failure risk this skill exists to catch. Hand these three rows to synthetic-monitor-author.
Hard-reject rule
If no RUM source is available (no Datadog RUM, no Sentry Performance data, no CrUX data for the target site), halt and return:
HALT: no RUM data source available.
Supply at least one of: Datadog RUM Explorer access, Sentry Performance
transaction list, or a CrUX-eligible public origin.
Synthetic coverage gap analysis requires real usage data as input.Do not estimate journey volume from gut feel or static sitemap inspection. Gap prioritization without usage data produces a monitor list biased by developer assumptions rather than actual user behavior.
Anti-patterns
| Anti-pattern | Why it fails | Fix |
|---|---|---|
| Deriving monitor list from sitemap alone | Sitemap contains every URL, not the ones users visit. High-traffic gaps get buried under low-traffic pages. | Use RUM session volume (Steps 1-2). |
| Treating all uncovered paths equally | A gap on /checkout and a gap on /legal/privacy are not the same risk. | Apply the coverage-priority score (Step 2). |
| Matching monitor URLs by exact string | /reports/123 and /reports/456 are the same journey pattern. Exact match leaves parameterized paths always "uncovered." | Normalize paths before diffing (Step 3). |
| Using CrUX for authenticated pages | CrUX only captures publicly discoverable pages per CrUX methodology (opens in new window). Authenticated journeys (dashboards, checkout) are invisible. | Use Datadog RUM or Sentry for post-login journeys. |
| Generating monitors for score < 10 | Creates monitor sprawl; low-traffic paths are not worth the maintenance and on-call noise. | Defer to a backlog; re-evaluate when traffic grows. |
Limitations
References
RUM source queries
View source (opens in new window)RUM source queries
Per-source instructions for pulling the top-N journey inventory (Step 1 of rum-to-synthetic-gap-analyzer). Aim for the top 50 view paths (or transaction names) by session volume to avoid chasing long-tail pages with negligible traffic.
Datadog RUM
In the RUM Explorer (https://app.datadoghq.com/rum/explorer):
Query syntax shorthand: @view.url_path:* | count by @view.url_path | sort desc. RUM Explorer supports key:value pairs where custom attributes require a created facet first (Datadog RUM Search (opens in new window)).
Sentry Performance
Open the Performance module and use the Trace Explorer to slice by transaction name. Per Sentry Transaction Summary docs (opens in new window), the platform surfaces throughput as TPM (transactions per minute) and TPS (transactions per second) per named transaction. Sort by Total throughput to surface highest-volume journeys. Export the table.
GA4 + CrUX (public-facing sites)
For public pages, the Chrome User Experience Report provides origin-level and URL-level field data. Per CrUX methodology (opens in new window), pages must be publicly discoverable (HTTP 200, no noindex) and meet a minimum visitor threshold for statistical confidence; exact threshold is undisclosed. Access via:
Per web.dev Core Web Vitals (opens in new window), the three stable metrics are:
All thresholds apply at the 75th percentile of page loads (web.dev CWV (opens in new window)). CrUX field data is "the Google dataset of the Web Vitals program" (CrUX docs (opens in new window)).
Related skills
cutover-sequence-author
Sequences a multi-team release cutover into dependency-ordered gates: builds the cross-service dependency graph, converts it into a numbered gate list where every gate carries exactly one named owner, a hard timebox, and a written rollback trigger, then derives the reverse-order rollback path and the window hard-stop rule. Emits one cutover plan document with an authority table and a runtime log. Use when two or more teams must cut over interdependent services inside one shared release window and nobody has yet written down the order, who calls each gate, or what reverses it.
feature-flag-experiment-validator
Validates the statistical significance of an A/B / feature-flag experiment result - computes per-metric effect size + p-value (chi-square for proportions, Welch's t-test for continuous metrics), applies a multiple-comparison correction (Bonferroni / Benjamini-Hochberg) when N>1 metric, surfaces practical-vs-statistical-significance distinction, and emits a ship/don't-ship verdict per metric. Use when an experiment has finished and someone is about to ship the winning variant off a dashboard readout, when a result rests on a small sample, or when more than one metric was compared - the rigorous version of "the variant looks better in the dashboard."
prod-canary-validator
Builds a canary-validation workflow that compares a canary deploy's metrics against the baseline (current main) - picks the metric set (error rate, p50/p95/p99 latency, business KPIs like checkout-completion), defines per-metric thresholds (absolute + relative-to-baseline), runs a statistical-comparison check (effect size + significance) over the canary's observation window, and emits a promote/rollback verdict. Use as the gate between canary deploy and full rollout - the deterministic version of "the on-call eyeballs the dashboard for 30 min.
release-runbook-author
Turns one service's release into a written six-phase runbook: pre-flight checks, a smoke gate, a canary observation window, a named human promote gate, progressive rollout, and post-release verification. Fixes each phase's pass criteria as a delta against a recorded baseline rather than a bare absolute number, gives canary and rollout separate windows and separate thresholds, and emits a per-phase evidence table that becomes the release record. Use when a single service is about to ship and its release steps exist only as tribal knowledge or a chat thread, so nobody can say in advance what evidence promotes it, what evidence halts it, or who decides.
synthetic-monitor-author
Drafts a synthetic monitor configuration for one critical user journey - picks the platform (Datadog Synthetics, Pingdom, Checkly, New Relic, etc.), authors the scripted-transaction body (Playwright-style for browser checks; HTTP-step for API checks), wires the cadence (typical 1-15 min), defines per-step assertions (DOM presence, API status, response shape) and aggregate alert thresholds (consecutive-failure count + on-call routing). Use when a critical journey needs continuous-in-production verification per ISTQB-canonical shift-right ("a test approach to test a system continuously in production").