Testland
Browse all skills & agents

jmeter-load-testing

Authors Apache JMeter `.jmx` test plans (Thread Groups + HTTP samplers + assertions + listeners) in the JMeter GUI, runs them headlessly via `jmeter -n -t plan.jmx -l results.jtl`, generates an HTML dashboard with `-e -o`, and gates CI on JTL parsing. Use when the project has an existing JMeter investment, needs JVM-native load tooling, or works in domains with strong JMeter community support (banking, telecom, enterprise).

Install with skills.sh (any agent)

npx skills add testland/qa --skill jmeter-load-testing
View source

jmeter-load-testing

Overview

Apache JMeter is the long-running JVM-native load-testing tool. Tests are XML .jmx files authored in JMeter's GUI; CI runs them headless via the jmeter CLI (jmeter-getstarted (opens in new window)).

The official guidance is explicit:

"GUI mode should only be used for creating the test script, CLI mode (NON GUI) must be used for load testing." (jmeter-getstarted (opens in new window))

This skill covers the CLI / CI side. Authoring is GUI-driven and out of scope here - see the JMeter user manual for the Thread Group / HTTP Sampler / Assertion authoring flow.

When to use

  • The project already has .jmx test plans.
  • The team is on the JVM and prefers JMeter's mature ecosystem (BlazeMeter, Taurus, plugin manager, distributed mode).
  • A specific protocol is needed that JMeter has first-class support for: JDBC, JMS, MQTT, FTP, LDAP, SOAP - k6's plugin landscape is weaker for niche protocols.

If the team is starting fresh, evaluate k6-load-testing (developer-friendly JS) or the Gatling (JVM DSL) and Locust (Python) references in load-testing-overview before adopting JMeter - XML authoring has a steep learning curve.

Install

JMeter requires Java 8 or higher with JAVA_HOME set. Download the latest release from jmeter.apache.org (opens in new window) and extract - there is no installer (jmeter-getstarted (opens in new window)).

For Docker-based CI, official images are at apache/jmeter. Pin to a specific tag rather than latest.

Running

Canonical CLI invocation

Per jmeter-getstarted (opens in new window):

jmeter -n -t test.jmx -l results.jtl
FlagPurpose
-nNon-GUI mode. Required for load tests.
-tPath to the .jmx test plan.
-lOutput file for raw sample results (JTL - CSV-shaped).

With HTML dashboard

Per jmeter-getstarted (opens in new window):

jmeter -n -t test.jmx -l results.jtl -e -o report_folder
FlagPurpose
-eGenerate the HTML dashboard report automatically.
-oOutput folder (must be empty or non-existent).

The dashboard shows percentiles, throughput, error rates, response- time graphs, and per-sampler breakdowns - the canonical JMeter output for human review.

Other useful flags

FlagPurpose
-J<property>=<value>Override a JMeter property at the JVM level.
-G<property>=<value>Override a property in distributed (remote) mode.
-Jjmeter.save.saveservice.output_format=csvForce CSV JTL (vs. XML).
-q <props-file>Additional properties file.
-rStart the test on remote slaves (distributed mode).
-XExit JMeter when the test finishes.

For multi-environment runs, parameterize the test plan with ${__P(api.base.url, default)} and override at the CLI:

jmeter -n -t orders.jmx -l results.jtl \
  -Japi.base.url=https://staging.example.com \
  -Japi.token=$API_TOKEN

Parsing results

The JTL file is the structured artifact for CI gating. Default JTL columns (CSV): timeStamp, elapsed, label, responseCode, responseMessage, threadName, dataType, success, failureMessage, bytes, sentBytes, grpThreads, allThreads, URL, Latency, IdleTime, Connect.

Quick-and-dirty pass/fail with awk:

# Count errors
awk -F',' 'NR>1 && $8=="false"' results.jtl | wc -l

# Compute mean response time
awk -F',' 'NR>1 { sum+=$2; n++ } END { print sum/n }' results.jtl

For richer parsing, use the JMeter HTML report's statistics.json under the report folder - it contains the percentiles per sampler in JSON form.

CI integration

# .github/workflows/jmeter.yml
name: load-test

on:
  pull_request:
    paths: ['tests/load/**.jmx']
  schedule:
    - cron: '0 4 * * *'

jobs:
  jmeter:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v5

      - name: Run JMeter via Docker
        env:
          API_TOKEN: ${{ secrets.STAGING_API_TOKEN }}
        run: |
          docker run --rm \
            -v "$PWD:/work" \
            -w /work \
            apache/jmeter \
              -n -t tests/load/orders.jmx \
              -l results.jtl \
              -e -o report \
              -Japi.token=$API_TOKEN \
              -Japi.base.url=https://staging.example.com

      - name: Pass/fail gate
        run: |
          ERRORS=$(awk -F',' 'NR>1 && $8=="false"' results.jtl | wc -l)
          if [ "$ERRORS" -gt 10 ]; then
            echo "::error::Got $ERRORS errors (>10 threshold)"
            exit 1
          fi

      - name: Upload report
        if: always()
        uses: actions/upload-artifact@v4
        with:
          name: jmeter-report
          path: |
            report/
            results.jtl
          retention-days: 14

The Docker invocation pattern keeps the CI runner clean - no Java / JMeter install on the runner.

Anti-patterns

Anti-patternWhy it failsFix
Running tests in GUI mode "just for this one quick run"GUI runner adds overhead; metrics are skewed; per jmeter-getstarted (opens in new window) this is explicitly forbidden for load tests.Always -n. GUI is for authoring only.
Hard-coding URLs / tokens in the .jmx XMLTest plan binds to one environment.Parameterize: ${__P(api.base.url, ...)}; override with -J.
Saving JTL in XML formatXML JTL files are 5-10x larger than CSV; slow to parse.Force CSV with -Jjmeter.save.saveservice.output_format=csv.
Listeners enabled in CI runsUI listeners (View Results Tree, Aggregate Report) consume RAM proportional to sample count; OOMs at scale.Disable all listeners in .jmx; emit JTL only; generate HTML report post-run via -e -o.
One mega-test-plan with 100 samplersFailure attribution is hard; runtime dominated by one slow endpoint.Split into per-domain .jmx files; run as separate CI jobs.
Threshold gates only on error rateA 30-second response that succeeds passes the rate gate but breaks UX.Pair error-rate with percentile gates parsed from statistics.json.

Limitations

  • XML authoring. .jmx files are not human-friendly; reviews of test-plan changes are awkward.
  • Memory hungry. JMeter's per-sample memory overhead is high vs. k6 / Gatling.
  • Single-machine VU limits. ~1000 threads per machine before JVM contention; for higher loads, use distributed mode (master + remote slaves) or move to a tool with a smaller per-VU footprint.
  • Plugin ecosystem fragmentation. Many JMeter plugins are abandoned; vet plugin maintenance status before depending on one.

References

  • jmeter-getstarted (opens in new window) - canonical CLI flags, GUI-vs-CLI guidance, dashboard generation.
  • Apache JMeter user manual - https://jmeter.apache.org/usermanual/
  • k6-load-testing and the Gatling / Locust references in load-testing-overview - alternatives by stack.
  • perf-budget-gate - downstream gate aggregating multiple runner verdicts.

Related skills

db-query-plan-analyzer

Reads `EXPLAIN` / `EXPLAIN ANALYZE` output from PostgreSQL, MySQL, or SQLite - identifies the dominant cost (sequential scan, nested loop, sort spill, missing index, type-cast preventing index use), proposes the specific index or query rewrite to fix it, and emits the candidate `CREATE INDEX` statement. Use when load testing or production telemetry shows the database as the bottleneck and the team needs targeted query-level remediation.

flame-graph-analyzer

Reads CPU flame-graph output from py-spy (Python), async-profiler (JVM), Go pprof, or Node.js `perf_hooks` / clinic.js: identifies the hot path (top sample-time frames), classifies the bottleneck (CPU-bound vs lock contention vs allocator pressure), and proposes the next investigation step. Use when a perf regression is bisected to a commit but the hot path inside it is unclear; for tail-latency percentiles use the latency-percentiles reference in k6-load-testing, and for a slow SQL hot path use db-query-plan-analyzer.

k6-load-testing

Authors k6 JavaScript load-test scripts (VU loops + checks + sleeps), configures the `options` block with `stages` (ramp-up patterns) and `thresholds` (p(95) latency, error rate), runs via `k6 run script.js` or `--vus / --duration` ad-hoc flags, and uses thresholds as the CI pass/fail signal. Includes a latency-percentile interpretation reference: tail ratio (p99/p50), bimodal-distribution detection, coordinated omission and why naive p99 is optimistic, and constant-vus vs constant-arrival-rate executors. Use when the project ships HTTP / WebSocket / gRPC load tests and the team wants developer-friendly JavaScript authoring, or when a k6 threshold passes but the system still feels slow.

lighthouse-perf

Configures Lighthouse CI (`@lhci/cli`) to audit Web Vitals (LCP, INP, CLS) on every PR, asserts against canonical thresholds (LCP ≤2.5s, INP ≤200ms, CLS ≤0.1 at the 75th percentile), uploads Lighthouse reports as build artifacts, and posts deltas as PR comments. Includes a budget-authoring reference: per-route LCP/INP/CLS thresholds by traffic class (cached / dynamic / api-heavy / form-heavy / media-heavy) via `assertMatrix`, plus `budget.json` resource-size caps (JS / CSS / images / total bytes). Use when the project ships a web frontend and the team needs continuous Web Vitals monitoring tied to PR gating, or needs its first Lighthouse budgets drafted.

load-testing-overview

Teaches load and performance testing from zero: a tool-selection table choosing between k6, JMeter, Gatling, Locust, and Artillery from observable project facts; the six load profiles (smoke, average-load, stress, spike, soak, breakpoint); open vs closed workload models; why percentiles beat averages; turning a run into a pass/fail CI gate with a first runnable k6 script; a performance-incident triage workflow (confirm with a k6 smoke run, flame-graph the hot path, check slow queries, localize the cause); and full Gatling (Simulation DSL, injectOpen/injectClosed, setUp().assertions()) and Locust (HttpUser + @task locustfile, headless / distributed runs, CSV gating) deep dives in references. Use when a service needs performance coverage and the tool, load profile, or pass/fail threshold has not been decided yet, or when a live performance incident needs cause localization.

perf-budget-gate

Builds a unified release-readiness gate that aggregates verdicts from any combination of k6 / JMeter / Gatling / Locust load runners and Lighthouse CI Web Vitals, applies severity-aware pass/fail thresholds, and emits a single go / no-go decision with per-metric deltas vs the main-branch baseline. Posts the delta as a PR comment when the team has the integration set up. Use when authoring a CI step that gates a deployment on cross-runner perf compatibility.

slo-load-test-plan

Turns a service's SLOs and endpoint traffic mix into a named scenario matrix: one scenario per SLO boundary condition, a load profile (smoke, average-load, stress, soak, spike, breakpoint) per scenario, an open or closed workload injection model, a threshold expression derived from the SLO the scenario guards, and an error-budget calculation that sets the soak run's failure allowance. Stays runner-agnostic and fixes the pass/fail line before any tool is configured. Use when an SLO document and an endpoint list both exist but nobody has decided which load runs to make, what shape of load each carries, or what number would count as a failure.

web-vitals-inp-deep

Deep INP (Interaction to Next Paint) testing: decomposes input delay, processing duration, and presentation delay via the web-vitals/attribution build, asserts per-interaction INP budgets in Playwright using PerformanceObserver plus the web-vitals visibilitychange flush, and identifies long tasks blocking the main thread. Use when a page feels unresponsive while LCP and CLS are green, or to gate key interactions (form submit, modal open, route change) under an INP budget in CI. Covers interactions only - for page-load Web Vitals gating use lighthouse-perf; for service-worker cache-strategy latency use the qa-pwa plugin's service-worker skills.