Testland
Browse all skills & agents

jvm-gc-tuning

Diagnoses JVM garbage-collection behaviour under load: reads and interprets unified GC logs (-Xlog:gc*), selects the right collector (G1 vs ZGC vs Parallel vs Serial), tunes heap sizing and pause-time targets, quantifies allocation rate, and traces the GC-pause-to-latency-tail link using GCViewer and Java Flight Recorder (JFR). Use when a load test reveals p99/p999 latency spikes that correlate with GC activity, or when heap sizing and collector selection need justification before a performance baseline is locked.

Install with skills.sh (any agent)

npx skills add testland/qa --skill jvm-gc-tuning
View source

jvm-gc-tuning

Overview

JVM garbage collection is a common hidden cause of latency tail in load tests. An application that looks CPU-bound or I/O-bound in a flame graph may instead be stalling on stop-the-world (STW) pauses. This skill walks the full diagnostic workflow from log capture through collector selection to heap sizing and tool analysis.

Upstream: pair with flame-graph-analyzer when GC frames appear as wide leaves in the async-profiler output. Downstream: hand tuned JVM flags to the load runner (k6-load-testing, gatling-load-testing) to re-run the baseline after changes.

Step 1 - Enable unified GC logging

Java 9+ replaced the legacy -XX:+PrintGCDetails family with the unified logging framework. The canonical capture command, per the Java 21 java man page (opens in new window):

# Minimum useful capture (info level, timestamped, rotating files)
java -Xlog:gc*:file=gc.log:time,uptime,level,tags:filecount=5,filesize=20M \
     -Xms4g -Xmx4g \
     -jar app.jar

Recommended tag set for GC diagnostics:

TagWhat it adds
gcEvery collection event + pause time
gc+heapHeap occupancy before/after each GC
gc+phasesPhase breakdown inside each pause (G1/ZGC)
gc+ageObject age histograms (tenuring threshold)
gc+alloc+profilingTop allocation call sites (requires -Xlog:gc+alloc+profiling=debug)

For a single diagnostic capture that covers all these tags:

java -Xlog:gc*,gc+phases,gc+heap,gc+age,gc+alloc+profiling=debug:file=gc_full.log:time,uptime,pid,level \
     -jar app.jar

Legacy flags (-XX:+PrintGCDetails, -XX:+PrintGCTimeStamps) still compile in Java 21 but are deprecated. Migrate to -Xlog:gc* before Java 25.

Step 2 - Read the GC log: key signals

A G1 GC log line at info level looks like:

[2.415s][info][gc] GC(3) Pause Young (Normal) (G1 Evacuation Pause) 512M->320M(2048M) 12.453ms

Fields:

FieldMeaning
2.415sUptime at pause start
GC(3)Monotonic GC event counter
Pause Young (Normal)Pause type (Young / Mixed / Full)
512M->320M(2048M)Heap before -> after (heap capacity)
12.453msWall-clock pause duration (STW time)

Red flags to scan for first:

  • Pause Full - a full-heap STW collection; almost always a misconfiguration or heap exhaustion signal.
  • Pause durations consistently above MaxGCPauseMillis target - G1 is missing its goal; see Step 4.
  • Heap after GC is consistently above 70-80% of capacity - heap is undersized or there is a live-set growth (possible leak).
  • (G1 Humongous Allocation) triggers - objects larger than half a heap region are being allocated in the old generation directly, bypassing the young gen.

Step 3 - Select the right collector

Per the HotSpot GC Tuning Guide, Available Collectors (opens in new window):

SituationCollectorFlag
Small dataset (<= 100 MB heap) or single CPUSerial-XX:+UseSerialGC
Batch / offline job; throughput dominant; pauses >= 1 s acceptableParallel-XX:+UseParallelGC
Server app; pause target 50-500 ms; default on modern JVMsG1-XX:+UseG1GC (default)
Latency-critical service; pause target < 1 ms; heap 100 MB to 16 TBZGC (Generational)-XX:+UseZGC -XX:+ZGenerational

Oracle's guidance: "Allow the VM to select automatically unless you have strict pause-time requirements. Adjust heap size first before changing collectors." (Available Collectors (opens in new window))

Pause-time vs throughput tradeoff

Per the Ergonomics section (opens in new window): GC throughput overhead is calculated via GCTimeRatio:

max GC time fraction = 1 / (1 + GCTimeRatio)

-XX:GCTimeRatio=19 targets at most 5% of CPU in GC (default for Parallel collector). Setting -XX:MaxGCPauseMillis too aggressively on G1 reduces young-gen size and increases GC frequency: the core tradeoff is that shorter pauses require more frequent collections.

ZGC does nearly all work concurrently, trading slightly higher CPU and memory overhead for near-removal of that tradeoff (ZGC documentation (opens in new window)).

Step 4 - Tune G1 for a pause-time target

From the G1 GC Tuning section (opens in new window):

Baseline low-latency profile (target < 50 ms pauses)

-XX:+UseG1GC
-Xms4g -Xmx4g                    # fixed heap avoids resize pauses
-XX:MaxGCPauseMillis=50           # pause-time target (statistical goal, not hard cap)
-XX:GCPauseIntervalMillis=200     # measure target over 200 ms window
-XX:G1NewSizePercent=10           # smaller young gen -> shorter pauses
-XX:G1MaxNewSizePercent=30
-XX:+G1UseAdaptiveIHOP            # let G1 learn IHOP from observed marking duration

See the key G1 flags table (defaults and effects for MaxGCPauseMillis, InitiatingHeapOccupancyPercent, G1HeapRegionSize, and more) in references/gc-reference.md.

Important: G1 is not a real-time collector. The pause target is a statistical goal. Setting it below 10 ms without also switching to ZGC typically causes G1 to thrash with very frequent tiny collections.

Step 5 - Tune ZGC for sub-millisecond latency

From the ZGC section (opens in new window): ZGC performs all expensive work concurrently; pauses remain below 1 ms regardless of heap size. The primary tuning lever is -Xmx.

-XX:+UseZGC -XX:+ZGenerational   # generational mode recommended
-Xms8g -Xmx8g
-XX:+AlwaysPreTouch               # pre-page memory at startup (reduces first-request jitter)
-XX:SoftMaxHeapSize=7g            # soft ceiling; GC keeps heap under this when possible

ZGC does NOT benefit from manual tuning of pause-time targets - those flags are G1-specific. If ZGC pauses grow, increase -Xmx or reduce allocation rate.

Step 6 - Analyse allocation rate

High allocation rate is the primary driver of GC frequency. Allocation rate (MB/s) = heap consumed between GC events / time between events.

Compute from GC log heap lines:

# Extract heap-before values from gc.log
grep "Pause Young" gc.log \
  | awk -F'[-> ()ms]+' '{print $5, $7, $9, $10}' \
  # fields: before_MB after_MB capacity_MB pause_ms

Allocation rate signals:

RateImplication
< 100 MB/sNormal for most services
100-500 MB/sWatch for frequent young-gen collections
> 500 MB/sObject pooling, streaming, or escape-analysis fixes needed
Sudden spike during load testRequest path creating large short-lived objects

To identify the allocation call sites, enable gc+alloc+profiling:

-Xlog:gc+alloc+profiling=debug:file=alloc.log:time

Or use JFR (see Step 7) which captures allocation sites with lower overhead.

Step 7 - Use Java Flight Recorder (JFR) for deep analysis

JFR provides structured GC event data with nanosecond resolution and captures allocation stack traces. Per the JFR API documentation (opens in new window), JFR provides more context, better timing control, and richer event information than text log files.

Enable JFR during a load test:

java -XX:StartFlightRecording=duration=120s,filename=recording.jfr,settings=profile \
     -jar app.jar

The key JFR event categories for GC analysis (jdk.GarbageCollection, jdk.ObjectAllocationInNewTLAB, jdk.ObjectAllocationOutsideTLAB, and more) are listed in references/gc-reference.md.

Analyse the .jfr file with JDK Mission Control (GUI) or the jfr CLI:

# Print GC pause summary from recording
jfr print --events jdk.GarbageCollection recording.jfr | head -100

# Print allocation call sites (top allocators)
jfr print --events jdk.ObjectAllocationInNewTLAB recording.jfr \
  | sort -t'=' -k2 -rn | head -20

Step 8 - Analyse with GCViewer

GCViewer (opens in new window) parses GC log files and computes summary metrics (throughput, pause time statistics, allocation rate, promotion rate) from the same -Xlog:gc* output.

# Run GCViewer (requires Java)
java -jar gcviewer-1.36.jar gc.log

The key metrics GCViewer surfaces (throughput %, avg/max pause, allocation and promotion rate) and how to read them are in references/gc-reference.md.

Step 9 - Trace the GC-pause-to-latency-tail link

A GC pause does not appear in CPU flame graphs (threads are suspended). To correlate GC pauses with HTTP latency spikes:

  1. Align GC log timestamps with load runner request timing. Both must use wall clock (-Xlog:gc*:decorators=time; k6 uses Unix epoch).
  2. In k6 output or Gatling simulation.log, find requests whose response time exceeds p99 baseline. Note their timestamps.
  3. In GC log, find Pause events that overlap those windows. A 50 ms STW pause causes all concurrent requests to stall for 50 ms.
  4. If the spikes co-occur with Pause Full or Pause Young (Concurrent Start), the collector is failing to keep up with allocation rate; see Steps 4-6.
  5. If spikes occur without matching GC pauses, GC is not the cause; escalate to flame-graph-analyzer for CPU profiling.

Output format

After completing the diagnostic, emit a findings block:

## JVM GC diagnosis - <service-name>

**Collector:** G1 / ZGC / Parallel / Serial
**Heap:** Xms=Ng / Xmx=Ng (fixed or dynamic)
**Log capture window:** Ns during load test at Nrps

### Pause summary
| Metric | Observed | Target |
|--------|----------|--------|
| Mean pause | Nms | MaxGCPauseMillis |
| p99 pause | Nms | - |
| Max pause (cause) | Nms (Pause Full / Humongous / ...) | - |
| Throughput (% non-GC) | N% | >95% |

### Allocation rate
N MB/s average; peak N MB/s at N rps load.
Top allocator: <class>.<method> (<evidence: JFR or alloc log>)

### Finding
<One sentence: what is wrong and why it causes latency tail.>

### Recommended changes
1. <Flag change with before/after value>
2. <Heap resize or allocation fix>

### Verification
Re-run load test with the same scenario; confirm p99 drops below threshold and
throughput % improves. Feed results to perf-budget-gate.

Anti-patterns

Anti-patternWhy it failsFix
Tuning pause target without fixing allocation rateShorter target forces smaller young gen; collections become more frequent; throughput dropsReduce allocation rate first; then lower pause target
Setting -Xms much smaller than -XmxHeap resize pauses during warm-up inflate early p99Set Xms = Xmx in production
Switching to ZGC without increasing heapZGC needs more headroom for concurrent reclamationAdd 20-30% extra heap when migrating from G1 to ZGC
Reading only Pause Full lines and ignoring Young pausesFull GC is the most visible event but frequent young pauses add up to more cumulative stallCompute cumulative GC time from all pause types
Diagnosing GC from flame graph aloneSTW pauses do not appear in on-CPU profilers; threads are suspendedCorrelate GC log timestamps with latency percentiles

Limitations

  • GC log timestamps use wall-clock by default; under NTP corrections they can drift relative to request timestamps. Use monotonic uptime (uptime decorator) for precise GC-to-request correlation.
  • -Xlog:gc+alloc+profiling=debug adds overhead (10-15% CPU). Use JFR allocation profiling in production-representative environments instead.
  • ZGC sub-millisecond claim applies to STW pauses only. Concurrent GC threads compete for CPU with application threads; under heavy load this can increase latency via CPU contention even without STW pauses.
  • GCViewer does not support ZGC log format for all ZGC-specific events; use JFR for ZGC deep analysis.

References

  • HotSpot GC Tuning Guide (Java 21) - introduction and goals: https://docs.oracle.com/en/java/javase/21/gctuning/introduction-garbage-collection-tuning.html
  • Available Collectors + selection criteria: https://docs.oracle.com/en/java/javase/21/gctuning/available-collectors.html
  • G1 GC tuning flags and mixed GC mechanics: https://docs.oracle.com/en/java/javase/21/gctuning/garbage-first-g1-garbage-collector1.html
  • ZGC characteristics and configuration: https://docs.oracle.com/en/java/javase/21/gctuning/z-garbage-collector.html
  • Ergonomics and GCTimeRatio formula: https://docs.oracle.com/en/java/javase/21/gctuning/ergonomics.html
  • Unified logging syntax (-Xlog:gc*) in the java man page: https://docs.oracle.com/en/java/javase/21/docs/specs/man/java.html
  • JFR API programmer's guide (event richness, timing, custom events): https://docs.oracle.com/en/java/javase/21/jfapi/why-use-jfr-api.html
  • GCViewer open-source GC log analyser: https://github.com/chewiebug/GCViewer
  • Upstream skill for CPU hot-path analysis: flame-graph-analyzer

JVM GC reference tables

View source (opens in new window)

JVM GC reference tables

Lookup tables for jvm-gc-tuning: G1 tuning flags, JFR event categories, and the metrics GCViewer surfaces. The runnable capture and analysis commands live in the main SKILL.md; these are the reference details each step points to.

Key G1 flags (Step 4)

From the G1 GC Tuning section (opens in new window):

FlagDefaultEffect
-XX:MaxGCPauseMillis200Pause-time target; G1 sizes young gen to meet this with high probability
-XX:InitiatingHeapOccupancyPercent45%Triggers concurrent marking when old gen occupancy reaches this level
-XX:G1HeapRegionSizeergonomic (~2048 regions)Region size 1-512 MB (power of 2); objects >= half this become humongous
-XX:G1MixedGCLiveThresholdPercent85%Regions with live occupancy above this are skipped during mixed GC
-XX:G1HeapWastePercent5%Space-reclamation phase ends when reclaimable space drops below this

G1 is not a real-time collector; the pause target is a statistical goal. Setting it below 10 ms without switching to ZGC typically causes G1 to thrash with very frequent tiny collections.

Key JFR event categories (Step 7)

Event categoryWhat to look for
jdk.GarbageCollectionDuration, cause, GC ID per collection
jdk.G1HeapSummaryEden/survivor/old occupancy before and after
jdk.ObjectAllocationInNewTLABAllocation call stacks (sampled)
jdk.ObjectAllocationOutsideTLABLarge object allocations (humongous)
jdk.GCHeapSummaryHeap used/committed at collection start

Key GCViewer metrics (Step 8)

MetricGC health interpretation
Throughput %% of time NOT in GC; target > 95% for most services
Avg pause timeMean STW duration; should be below MaxGCPauseMillis
Max pause timeWorst-case STW; the main driver of p99/p999 latency
Allocation rate (MB/s)Matches the Step 6 allocation-rate analysis
Promotion rate (MB/s)Objects escaping young gen; high rate fills old gen faster

Related skills

db-query-plan-analyzer

Reads `EXPLAIN` / `EXPLAIN ANALYZE` output from PostgreSQL, MySQL, or SQLite - identifies the dominant cost (sequential scan, nested loop, sort spill, missing index, type-cast preventing index use), proposes the specific index or query rewrite to fix it, and emits the candidate `CREATE INDEX` statement. Use when load testing or production telemetry shows the database as the bottleneck and the team needs targeted query-level remediation.

flame-graph-analyzer

Reads CPU flame-graph output from py-spy (Python), async-profiler (JVM), Go pprof, or Node.js `perf_hooks` / clinic.js: identifies the hot path (top sample-time frames), classifies the bottleneck (CPU-bound vs lock contention vs allocator pressure), and proposes the next investigation step. Use when a perf regression is bisected to a commit but the hot path inside it is unclear; for tail-latency percentiles use latency-percentile-analyzer, for GC pauses specifically use jvm-gc-tuning, and for a slow SQL hot path use db-query-plan-analyzer.

gatling-load-testing

Authors Gatling simulations in Java / Kotlin / Scala (or JS / TS) using the Simulation class plus http() / scenario() / exec() DSL builders, ramps virtual users via injectOpen (arrival rate) or injectClosed (concurrent count), runs via Maven / Gradle / sbt with the Gatling plugin, and gates CI on assertions defined in setUp(). Use when the project is on the JVM and the team prefers code-first load tests over JMeter's XML or k6's JavaScript-only authoring.

jmeter-load-testing

Authors Apache JMeter `.jmx` test plans (Thread Groups + HTTP samplers + assertions + listeners) in the JMeter GUI, runs them headlessly via `jmeter -n -t plan.jmx -l results.jtl`, generates an HTML dashboard with `-e -o`, and gates CI on JTL parsing. Use when the project has an existing JMeter investment, needs JVM-native load tooling, or works in domains with strong JMeter community support (banking, telecom, enterprise).

k6-load-testing

Authors k6 JavaScript load-test scripts (VU loops + checks + sleeps), configures the `options` block with `stages` (ramp-up patterns) and `thresholds` (p(95) latency, error rate), runs via `k6 run script.js` or `--vus / --duration` ad-hoc flags, and uses thresholds as the CI pass/fail signal. Use when the project ships HTTP / WebSocket / gRPC load tests and the team wants developer-friendly JavaScript authoring.

latency-percentile-analyzer

Interprets latency distributions from k6 load tests beyond the p95/p99 gate: reads percentile summaries and JSON exports to identify tail shape, computes the tail ratio (p99/p50) as a distribution-spread signal, detects bimodal distributions, explains coordinated omission and why naive p99 values are optimistic under sustained load, and distinguishes request-rate from concurrency models. Use when a k6 threshold passes but the system still feels slow, when p99 is suspiciously low during ramp-up, or when the team needs to explain why tail latency is high rather than just observing that it is.

lighthouse-budget-author

Drafts a `lighthouserc.js` (or `budget.json`) at design time - picks Web Vitals thresholds (LCP / INP / CLS) per route based on traffic class (cached / dynamic / API-heavy / form-heavy) and the team's NFRs, plus resource-size budgets (JS / CSS / images / total bytes). Emits the config file ready for the lighthouse-perf runner. Use when starting Lighthouse coverage on a project that has no budgets yet, or when the existing budgets need a redesign.

lighthouse-perf

Configures Lighthouse CI (`@lhci/cli`) to audit Web Vitals (LCP, INP, CLS) on every PR, asserts against canonical thresholds (LCP ≤2.5s, INP ≤200ms, CLS ≤0.1 at the 75th percentile), uploads Lighthouse reports as build artifacts, and posts deltas as PR comments. Use when the project ships a web frontend and the team needs continuous Web Vitals monitoring tied to PR gating.

load-testing-overview

Teaches load and performance testing from zero: how to choose between k6, JMeter, Gatling, Locust, and Artillery based on observable project facts (team language, tests-as-code vs GUI authoring, protocols beyond HTTP, CI gating needs); the six load profiles (smoke, average-load, stress, spike, soak, breakpoint) and the question each one answers; the difference between open workload models that hold arrival rate constant and closed models that hold concurrent users constant; why percentiles rather than averages are the unit of measurement; and how to turn a run into a pass/fail CI gate, with a first runnable k6 script. Use when a service needs performance coverage and the tool, the load profile, or the pass/fail threshold has not been decided yet.

locust-load-testing

Authors Locust load tests as Python classes - HttpUser with @task-decorated methods plus on_start hooks and between() wait_time - runs via `locust -f locustfile.py` headless mode (or distributed via `--master` / `--worker`), and exports CSV / JUnit reports for CI gating. Use when the project's primary stack is Python and the team wants load tests in the same language as the application.

perf-budget-gate

Builds a unified release-readiness gate that aggregates verdicts from any combination of k6 / JMeter / Gatling / Locust load runners and Lighthouse CI Web Vitals, applies severity-aware pass/fail thresholds, and emits a single go / no-go decision with per-metric deltas vs the main-branch baseline. Posts the delta as a PR comment when the team has the integration set up. Use when authoring a CI step that gates a deployment on cross-runner perf compatibility.

slo-load-test-plan

Turns a service's SLOs and endpoint traffic mix into a named scenario matrix: one scenario per SLO boundary condition, a load profile (smoke, average-load, stress, soak, spike, breakpoint) per scenario, an open or closed workload injection model, a threshold expression derived from the SLO the scenario guards, and an error-budget calculation that sets the soak run's failure allowance. Stays runner-agnostic and fixes the pass/fail line before any tool is configured. Use when an SLO document and an endpoint list both exist but nobody has decided which load runs to make, what shape of load each carries, or what number would count as a failure.