Vicedomini Softworks

Software Development

Profile First: Optimize Java Performance With JFR, jcmd, and JVM Flags

28 September 2026

Decorative Java performance title card

Profile first, fix the code that costs the most, then tune the JVM with measured, single-variable changes. The handful of levers that produce the largest return are consistent across most workloads: capturing accurate diagnostics through JFR and jcmd, reducing unnecessary allocation, setting the correct maximum heap size, and selecting the garbage collector that matches the application’s latency and throughput goals. Every subsequent decision should be justified by data, not intuition.


TL;DR:

  • Setting -Xmx close to available memory and increasing heap size often reduces GC-induced latency more effectively than changing collectors.
  • Using JFR and jcmd for baseline recording under real traffic helps identify hot methods, allocations, and GC events before making configuration changes.
  • G1 remains the default for most enterprise workloads, while ZGC is suitable only for latency-sensitive systems with strict pause-time requirements.
  • Algorithmic improvements, such as reducing object allocations and reusing buffers, deliver higher performance gains than JVM flag tuning.
  • Profiling should be continuous and data-driven, with changes tested individually and documented to ensure reversibility and maintainability.

Vicedomini Softworks
Build Java Systems That Stay Fast
Vicedomini Softworks applies engineering expertise across Java, performance, cloud infrastructure, and ongoing application maintenance.
Explore software engineering services

Table of Contents

How to measure and diagnose performance before changing anything

Tuning without measurement produces guesses dressed up as engineering decisions. The correct starting point is a clear, quantifiable definition of success: a target p95 latency, a throughput floor, or a maximum acceptable GC pause. Without such a figure, it becomes impossible to determine when a change has helped, hurt, or made no meaningful difference, a point emphasised in established Java performance literature that argues for setting quantifiable success criteria before any tuning begins.

The modern diagnostic toolchain in the JDK is built around jcmd and Java Flight Recorder (JFR), which together allow engineers to capture detailed runtime data with a manageable overhead compared with older, more invasive profiling agents. According to Oracle’s own troubleshooting documentation, jcmd can start and dump JFR recordings, trigger heap dumps and print performance counters, provided it runs as the same user on the same machine as the target JVM. This makes it suitable for production diagnostics without deploying a separate profiling stack.

A practical, low-risk workflow looks like this:

  • Start a baseline JFR recording with jcmd <pid> JFR.start and let it run under real traffic before making any change.
  • Dump the recording with jcmd <pid> JFR.dump and inspect Hot Methods, allocation profiles and GC pause events in a viewer.
  • Capture a heap dump with jcmd <pid> GC.heap_dump only when investigating a specific memory concern, since dumps pause the JVM briefly.
  • Reserve JMH for isolated microbenchmarks once the profiler has identified a specific hot method or allocation site.

One detail matters more than most engineers expect: continuous production recordings should avoid enabling heap statistics, because Oracle’s guidance notes that this setting can trigger old garbage collections at the start and end of a recording, which skews the very latency figures being measured. A longer recording window, on the order of an hour, tends to produce a more representative sample without repeatedly forcing collections.

JMH deserves separate mention because it is frequently misused. Microbenchmarks written as ad-hoc timing loops in a main method are vulnerable to dead-code elimination, insufficient warm-up and JIT inconsistencies. The JMH project exists specifically to control these variables, and the correct approach is to scaffold a benchmark module using the Maven archetype rather than running benchmarks from within an IDE, where class loading and background processes introduce noise that invalidates results.

When reading a JFR recording, three views carry most of the diagnostic weight. Hot Methods shows where CPU time actually accumulates, often revealing that a method assumed to be cheap is called far more often than expected. TLAB allocation metrics expose which call sites generate the most short-lived objects, a strong predictor of GC pressure. GC pause events, plotted against the recording’s timeline, show whether pauses correlate with specific request patterns or batch jobs, which is often more informative than the raw pause duration alone.

Pro Tip: Take a JFR baseline before any change, however small; without it, every later comparison is a guess dressed up as a measurement.

Choosing and tuning a garbage collector for your workload’s goals

Garbage collector choice is a trade-off between pause-time and throughput, and the JVM’s own ergonomics already make a reasonable default choice. As of JDK 26, Oracle’s tuning guidance confirms that the JVM selects collectors by ergonomics, with G1 typically active on server-class machines, Serial reserved for small heaps under roughly 100 MB, and ZGC positioned for workloads with strict low-latency requirements. Changing the collector should be a response to a measured problem, not a starting assumption.

G1 balances pause-time and throughput reasonably well for most enterprise workloads, including typical Spring Boot or Jakarta EE services handling web traffic with moderate heap sizes. It divides the heap into regions and performs incremental collection, which keeps pauses generally short and predictable without demanding manual region sizing in most cases.

ZGC becomes relevant when an application has strict, sub-millisecond pause requirements, commonly seen in trading systems, real-time bidding platforms or latency-sensitive APIs where a single long pause causes a visible service degradation. Oracle’s documentation on the Z Garbage Collector identifies maximum heap size as its most important tuning knob, and notes that ZGC supports -XX:SoftMaxHeapSize to guide heap-growth heuristics without forcing a hard ceiling that could cause allocation stalls.

Increasing the maximum heap size resolves a substantial share of GC-induced latency problems, according to Oracle’s Z Garbage Collector guidance, because concurrent collectors need headroom above the live data set to avoid falling behind allocation rate. A collector starved of heap space will collect more frequently regardless of which algorithm it runs.

The practical knobs worth knowing, in rough order of impact:

  • -Xmx and -Xms set the maximum and initial heap size and remain the single most consequential parameters for GC behaviour.
  • -XX:MaxGCPauseMillis sets a soft pause-time goal for G1, which the collector attempts to honour by adjusting region collection sets.
  • -XX:GCTimeRatio expresses the ratio of application time to GC time the JVM should aim for, influencing how aggressively the heap grows.
  • Min/MaxHeapFreeRatio controls how much free heap space triggers growth or shrinkage, shaping how the JVM reacts to changing load.

The correct sequence is to measure, adjust heap sizing first, retest under the same workload, and only change the collector itself if heap adjustments fail to close the gap. Oracle’s own tuning guide frames this explicitly as an iterative process: establish a baseline, change one variable, measure, and prioritise application-level fixes such as reduced allocation over JVM-level tweaks, which should be applied only once profiling confirms the JVM itself is the bottleneck.

Heap sizing and generation sizing in practice

Heap sizing decisions have an outsized effect on collector behaviour, more so than most of the flags engineers reach for first. Oracle’s tuning documentation stresses that the JVM grows and shrinks the heap according to ergonomic targets governed by Min/MaxHeapFreeRatio, bounded between -Xms and -Xmx, with -XX:GCTimeRatio setting the underlying goal for GC time versus application time, as described in the HotSpot garbage collection tuning guide. Fixing sizes by hand before understanding this behaviour often works against the JVM rather than with it.

A few rules hold across most production deployments:

  • Set -Xmx close to available physical memory without approaching the point where the operating system starts swapping, since swapping turns a GC pause into a multi-second stall.
  • Set -Xms equal to -Xmx when predictable latency matters more than a smaller memory footprint at startup, avoiding repeated heap resizing under load.
  • Use -XX:NewRatio to control the ratio between young and old generation size, raising it when the workload allocates many short-lived objects that should be collected quickly.
  • Use -XX:MaxNewSize to cap young generation growth directly when NewRatio alone produces an awkward split.
  • Adjust -XX:SurvivorRatio when objects that should die young are instead being promoted to the old generation prematurely, forcing more expensive major collections.

Oscillating heap behaviour, where the heap repeatedly grows and shrinks under steady load, usually signals that the free ratio targets are fighting the actual allocation pattern. A JFR recording spanning several of these cycles typically shows the pattern clearly on the heap usage timeline, and the fix is almost always to widen the gap between MinHeapFreeRatio and MaxHeapFreeRatio or to set Xms closer to Xmx so the JVM stops chasing a moving target.

JIT and runtime compiler behaviour: warm-up and when to intervene

The Java runtime compiler works in tiers, starting with the fast but unoptimised C1 compiler and progressively promoting hot methods to C2, which applies deeper optimisations at the cost of longer compilation time. This staged approach explains why a service often performs noticeably worse in its first minutes after startup than it does once traffic has run long enough for hot paths to reach full C2 optimisation. Benchmarks and load tests that ignore this warm-up period routinely produce misleading conclusions.

Illustration of tiered JIT warm-up

Reaching for compiler flags before profiling is one of the more common mistakes in Java tuning. Flags such as forcing tiered compilation off, or manually adjusting compilation thresholds, address a symptom that is usually better explained by insufficient warm-up time in a benchmark or an allocation problem elsewhere. Oracle’s own guidance on collector tuning applies equally here: manual overrides that disable JVM ergonomics often make matters worse, since the runtime’s adaptive heuristics are generally well tuned for typical workloads unless a specific measurement says otherwise.

For teams with unusually strict startup-time requirements, GraalVM’s native image compilation and profile-guided optimisation represent a more advanced path, trading a longer and more complex build process for near-instant startup and a smaller memory footprint. This trade-off suits serverless functions and command-line tools far better than long-running services, where the JIT’s ability to specialise against actual runtime behaviour over time tends to win out on sustained throughput.

Pro Tip: Never judge a service’s steady-state performance from its first few minutes of traffic; the JIT has not finished its job yet.

  • Let a service run under representative load for several minutes before drawing conclusions from any benchmark or dashboard.
  • Treat compiler flags as a last resort, applied only after profiling rules out allocation, algorithmic and GC causes.
  • Reserve native image and PGO for workloads where startup latency, not sustained throughput, is the binding constraint.

Highest ROI code changes: algorithms, allocations, strings and collections

Fixing algorithmic complexity beats every micro-optimisation available at the JVM level. A method that scans a list in a loop, turning an O(n) operation into O(n squared) when nested inside another loop, will dominate a profile regardless of which garbage collector or heap size sits underneath it. Profiling data pointing at a specific hot method is the signal to check its complexity before reaching for anything else.

Once algorithmic issues are ruled out, the following changes tend to produce the largest measurable gains, in the order most teams should apply them:

  1. Reduce short-lived object allocation in hot paths, since fewer allocations directly reduce TLAB pressure and the frequency of young-generation collections.
  2. Reuse buffers and collections across iterations rather than constructing them fresh inside loops, particularly for byte arrays and StringBuilder instances used repeatedly.
  3. Prefer primitive types over boxed wrapper types in performance-sensitive code, since autoboxing creates hidden allocations that rarely show up until a profiler surfaces them.
  4. Replace repeated string concatenation with a single reused StringBuilder, since the plus operator on strings inside a loop compiles to a new StringBuilder on every iteration.
  5. Compile regular expressions once as a static Pattern field rather than calling String.matches or Pattern.compile inside a loop, where recompilation cost accumulates quickly.
  6. Avoid String.format in tight loops, since its locale-aware formatting machinery is considerably slower than direct string building for simple cases.
  7. Choose collection implementations and initial capacities deliberately, since a HashMap resized repeatedly during population wastes cycles that a correctly sized constructor call avoids entirely.
  8. Consider primitive-specialised collections, available in libraries built for that purpose, when a workload stores large numbers of primitive values and boxing overhead is measurable in a profile.

Streams deserve a specific note because they are often applied for readability without checking the performance cost. A stream pipeline with several intermediate operations can allocate more objects than an equivalent loop, particularly when boxing primitives, and the difference only matters when profiling shows the code path is hot. Where it is hot, a plain loop is frequently both clearer and faster than a heavily chained stream. Parallel streams introduce their own risk: they borrow threads from the common fork/join pool, and using them for anything other than genuinely CPU-bound, independent work can starve other parts of the application that rely on the same pool.

Diagnosing and reducing thread contention

Contention rarely announces itself directly; it shows up as CPU time spent waiting rather than working, and JFR is the most reliable way to see it. A recording’s thread activity and lock contention views reveal which monitors or synchronised blocks are creating queues, often in code that looked harmless during a code review because the contention only appears under concurrent production load. Traditional thread dumps, taken at the moment of a suspected stall, complement this by showing exactly which threads are blocked and on what.

Once a contended lock is identified, the fix depends on what the lock protects. Fine-grained locking, splitting one broad lock into several narrower ones guarding independent data, often removes contention without changing the underlying logic. Non-blocking structures, such as those built on atomic variables or concurrent collections, suit cases where the contended operation is simple enough to express without a full mutual-exclusion lock.

  • Size thread pools against the actual bottleneck, whether CPU cores or an external dependency’s concurrency limit, rather than an arbitrary round number.
  • Choose a bounded queue strategy for thread pools handling bursty traffic, since an unbounded queue converts back-pressure into a slow memory leak.
  • Avoid unbounded thread creation per request, which exhausts native memory and context-switching capacity well before it exhausts the heap.
  • Watch TLAB allocation metrics per thread, since a thread allocating far more than its peers is often the one driving young-generation collection frequency for the whole process.

Thread-local allocation buffers exist precisely so that most allocations avoid contention entirely, each thread claiming its own buffer from the heap. When allocation-heavy threads exhaust their TLAB rapidly, the resulting refill requests become a source of contention themselves, which is one more reason reducing allocation at the code level tends to help concurrency metrics as a side effect.

How to design and run reliable benchmarks with JMH

A benchmark that cannot be trusted is worse than no benchmark, since it produces confident wrong answers. The correct starting point is scaffolding a dedicated benchmark module using the JMH Maven archetype, then running it from the command line as a packaged jar rather than inside an IDE, where class loading and background threads distort measurement, a distinction the JMH project itself was built to enforce.

  1. Configure warm-up iterations generously enough that the JIT has reached steady state before measurement iterations begin.
  2. Use multiple JVM forks for each benchmark so that JIT compilation quirks in one fork do not bias the overall result.
  3. Compare runs using the same hardware, the same JVM version and the same flags, changing only the single variable under test.
  4. Feed representative data into the benchmark rather than trivial inputs that let the JIT optimise away the very work being measured.

Microbenchmarks answer a narrow question well, but they should never stand alone as evidence for a production decision. Pairing them with soak tests and load tests that exercise the whole request path catches interactions, such as GC pressure from concurrent traffic, that an isolated microbenchmark cannot see.

Pro Tip: A microbenchmark that shows a ten percent gain in isolation can still lose to GC pressure once ten thousand requests a second are allocating alongside it; validate with a load test before shipping the change.

Concrete JVM option examples and a stepwise tuning checklist

Copying flags from a blog post without measuring their effect is how JVM configurations accumulate cruft over years. The safer pattern is to start from a minimal, ergonomically sound baseline and add flags only when a specific measurement justifies each one, keeping the resulting configuration in version control alongside the application code.

A checklist worth following on every tuning pass:

  • Capture a baseline JFR recording under representative load before changing anything.
  • Change exactly one JVM flag, never several at once, so any effect can be attributed correctly.
  • Record again under the same workload and compare the same metrics: p95 latency, throughput, GC pause count and duration.
  • Revert immediately if the change makes any tracked metric worse, rather than layering a second change on top to compensate.
  • Enable -Xlog:gc alongside JFR for a lightweight, human-readable log of collection events that is easy to grep during follow-up analysis.
Goal Starting flags
Throughput-oriented server -Xms and -Xmx set to equal values with UseG1GC and an appropriate MaxGCPauseMillis setting
Low-latency service -Xms8g -Xmx8g -XX:+UseZGC -XX:SoftMaxHeapSize=7g
Small-memory process -Xms and -Xmx set to modest sizes suited for small-memory processes with UseSerialGC

These are starting points, not prescriptions; each must be validated against the application’s own workload with the profiling workflow described earlier. Flags that disable ergonomics outright, such as fixing generation sizes far from what the allocation pattern needs, or forcing a compilation tier off without a measured reason, tend to remove the JVM’s ability to adapt and should be treated as a last resort rather than a default.

Practical workflow and how Vicedomini Softworks applies performance engineering

Vicedomini Softworks approaches Java performance work as a lifecycle rather than a one-off fix: baseline the current behaviour with production-safe profiling, identify the highest-impact code and configuration issues, agree a remediation plan with measurable targets, implement changes one variable at a time, and leave monitoring in place so regressions surface before they reach customers. Consulting engagements include written recommendations, reflecting a commitment to no vendor bias and transparent delivery.

This lifecycle applies whether the underlying stack is Spring Boot, Quarkus or Jakarta EE, since the profiling and tuning principles described above are largely stack-agnostic. Teams with an established platform engineering function often handle this cycle internally once the workflow is in place. Organisations without dedicated capacity, or facing a performance issue with a deadline attached, tend to benefit from bringing in engineers who work directly with the code rather than through an account-management layer.

Balancing performance gains and maintainability

Performance work only pays off long-term when it is measurable and reversible. A JVM flag added to fix a symptom two years ago, with no record of why, becomes a constraint nobody dares remove. Every tuning change deserves a documented baseline, a stated goal and a place in version control, alongside a performance check in continuous integration and real observability in production, so a regression is caught in minutes rather than discovered by a customer.

— Pepe F.

How Vicedomini Softworks can help with Java performance work

Diagnosing a slow Java service under production load, without disrupting the traffic that revenue depends on, is where most in-house teams run short on time rather than skill. Vicedomini Softworks works directly with client engineers, using the same profiling-first workflow described throughout this article, to turn a vague complaint about slowness into a written, prioritised remediation plan.

Vicedomini Softworks

Relevant services for a performance engagement include:

  • A code audit that identifies the specific allocation, algorithmic or concurrency issues driving latency or GC overhead.
  • Production profiling using JFR and jcmd to gather evidence without disruptive instrumentation.
  • A remediation plan with measurable targets for latency, throughput and GC pause frequency.
  • Fractional CTO or CTO Advisory support for teams that need retained technical leadership through the fix and beyond.

Engagements typically start with an initial assessment, priced from €3,500 as a one-off, which establishes the baseline and scope before any remediation work begins. Teams wanting ongoing technical leadership can explore the CTO Advisory, Fractional CTO and CTO Partner plans on the same CTO advisory page, or review the full range of engineering services on the services overview.

Authoritative documentation and tool references

The following official references are worth bookmarking for deeper technical work beyond this guide. Oracle’s introduction to garbage collection tuning explains collector ergonomics and default selection logic in full. The G1 garbage collector tuning guide covers region sizing and pause-time goals for the default collector on most server-class machines. The Z Garbage Collector documentation details heap sizing and SoftMaxHeapSize behaviour for low-latency workloads. Oracle’s guide to troubleshooting performance issues using JFR is the primary reference for jcmd commands and recording strategy. The JMH project repository remains the canonical source for scaffolding correct, reproducible microbenchmarks.

FAQ

What is the single most effective way to improve Java performance?

Profiling the application under real load with JFR and jcmd before making any change is the most reliable starting point, since it points directly at the code or configuration actually causing the slowdown. From there, fixing allocation and algorithmic issues in the code usually delivers more gain than JVM flags, and collector or heap changes come last, guided by Oracle’s iterative tuning guidance.

Which garbage collector should I use for a typical Java service?

G1 is the ergonomic default on server-class machines and suits most enterprise workloads without manual intervention, according to Oracle’s collector documentation. ZGC is worth considering only when a workload has strict low-latency requirements that G1 cannot meet, since its main tuning lever is heap size rather than pause-time flags.

How do I safely profile a Java application in production?

Use jcmd to start and dump a Java Flight Recorder recording on the running process, which Oracle documents as running with low overhead compared with older invasive profilers. Avoid enabling heap statistics in continuous recordings, since this can trigger extra garbage collections that distort the very latency figures being measured.

Why do my JMH microbenchmark results not match production behaviour?

Microbenchmarks answer a narrow, isolated question and can miss interactions such as GC pressure from concurrent traffic that only appear under real load. The JMH project itself is built to avoid measurement artefacts like insufficient warm-up, but its results should still be validated against soak or load tests before informing a production change.

When should I increase the JVM’s maximum heap size?

Increase -Xmx when profiling shows GC pauses or frequency driven by insufficient headroom above the application’s live data set, which Oracle’s ZGC guidance identifies as the most impactful single parameter for concurrent collectors. The ceiling should stay below available physical memory to avoid swapping, which turns a manageable GC pause into a far longer stall.