JDK Flight Recorder in production: record, read, automate
JDK Flight Recorder (JFR) is part of the JDK. There is nothing to install. It records what happens inside the JVM: GCs, locks, exceptions, reads and writes, and thread stacks. It is designed to run in production.
It is often used only after an incident, with a GUI. But it can be used from the command line, with two JDK tools: jcmd to start and stop a recording, jfr to read it.
To show how to use JFR, this article relies on a small application that contains four problems. All tests were run on JDK 25. The recordings of the application were made on Temurin 25.0.4, in a Linux container with 4 CPUs. The cost measurement, the shutdown tests, JFR.dump, the custom events, RecordingStream and jfr scrub ran on OpenJDK 25.0.1, on a Mac.
What JFR records
JFR records events. Each event has a type, a timestamp, often a duration, the thread that produced it and, depending on the type, a stack trace.
They come in three families:
- events with a duration, like a wait on a lock or a file write. They have a threshold. If the event is shorter than the threshold, JFR does not keep it.
- periodic events, like CPU load or heap usage. JFR records them at a fixed interval.
- samples, like the stacks of active threads. JFR takes a limited number per second.
A recording contains close to 200 event types. The majority stay empty, because nothing triggered them.
The example application
Four worker threads process orders in a loop. For each order, the program reads a quantity, creates an invoice of about a hundred lines, writes it to an audit log, then keeps it in a history list.
public class Orders {
static final Object AUDIT = new Object();
static final List<String> HISTORY = new ArrayList<>();
static Writer auditLog; // a BufferedWriter on /tmp/audit.log
// Problem 1: an exception is used for control flow.
static int quantity(String raw) {
try {
return Integer.parseInt(raw);
} catch (NumberFormatException e) {
return 1;
}
}
// Problem 2: String.format in a loop.
static String invoice(int lines) {
StringBuilder sb = new StringBuilder();
for (int i = 0; i < lines; i++) {
sb.append(String.format("%05d;%-20s;%8.2f%n", i, "article-" + i, i * 1.5));
}
return sb.toString();
}
// Problem 3: the audit write happens under a global lock.
static void audit(String invoice) throws IOException {
synchronized (AUDIT) {
auditLog.write(invoice);
auditLog.flush();
try { Thread.sleep(5); } catch (InterruptedException e) {}
}
}
// Problem 4: a history list that never gets cleared.
static void remember(String invoice) {
synchronized (HISTORY) {
HISTORY.add(invoice.substring(0, 64) + System.nanoTime());
}
}
}
The Thread.sleep(5) simulates a slow disk. A third of the quantities are empty or invalid, and each one throws an exception. The main thread prints the throughput every five seconds:
162 orders/s
164 orders/s
165 orders/s
Recording from JVM startup
The -XX:StartFlightRecording flag starts a recording at the same time as the JVM:
java -XX:StartFlightRecording:duration=20s,filename=orders.jfr -cp classes Orders 4
The JVM prints this message at startup:
[0.597s][info][jfr,startup] Started recording 1. The result will be written to:
[0.597s][info][jfr,startup]
[0.597s][info][jfr,startup] /work/orders.jfr
Throughput does not change: about 164 orders per second, with and without JFR. After twenty seconds, the file is 394 kB.
Without the duration parameter, the recording continues until the JVM stops. The message then gives the command to get the data: Use jcmd <pid> JFR.dump name=1 to copy recording data to file.
Or without restarting the JVM, using jcmd
The flag is not required. On a JVM that is already running, jcmd starts a recording:
jcmd <pid> JFR.start name=prod settings=profile maxage=10m
Started recording 1.
Use jcmd 465 JFR.dump name=prod filename=FILEPATH to copy recording data to file.
Four commands cover the usual needs:
jcmd <pid> JFR.check # running recordings
jcmd <pid> JFR.view hot-methods # read without writing a file
jcmd <pid> JFR.dump name=prod filename=/tmp/dump.jfr # copy the data to disk
jcmd <pid> JFR.stop name=prod # stop
JFR.check prints Recording 1: name=prod maxage=10m (running). JFR.view shows the current data in the terminal. JFR.dump writes the data to disk, without stopping the recording.
Reading the recording without a GUI: jfr view
First, jfr summary gives the number of events per type. Here are the useful lines for our twenty seconds:
Event Type Count Size (bytes)
=============================================================
jdk.ObjectAllocationSample 2926 41404
jdk.JavaExceptionThrow 1137 17164
jdk.ExecutionSample 83 892
jdk.JavaMonitorEnter 24 562
jdk.GarbageCollection 23 484
jdk.FileWrite 2 65
To see the details, jfr view offers ready-to-use views. jfr help view lists close to 80 of them, in three groups: the JVM, the environment and the application. Here are the views that show our first three problems.
hot-methods: where the CPU works
jfr view hot-methods orders.jfr
Method Samples Percent
----------------------------------------------------------------- ------- -------
java.util.Formatter.parse(String) 7 8.43%
java.lang.AbstractStringBuilder.append(char) 7 8.43%
Orders.lambda$main$0(String[]) 5 6.02%
java.util.Formatter$FormatSpecifierParser.parse() 5 6.02%
java.util.Formatter$FormatSpecifier.print(Formatter, int, Locale) 4 4.82%
java.util.Formatter.format(Locale, String, Object[]) 4 4.82%
This view counts the method at the top of the stack when the sample was taken. This is why Orders.invoice does not appear. The list shows the java.util.Formatter methods it calls. This is problem 2: String.format parses the format string again for each line.
There are only 83 samples in twenty seconds. This is normal. jdk.ExecutionSample only takes threads that are running Java code. Our threads spend a large part of their time waiting for the lock.
contention-by-site: which code waits for a lock
jfr view contention-by-site orders.jfr
StackTrace Count Avg. Max.
--------------------------------------------------- ----- ------- -------
Orders.audit(String) 23 26.7 ms 39.2 ms
jdk.jfr.internal.PlatformRecorder.periodicTask() 1 22.2 ms 22.2 ms
This is problem 3. Threads wait 26.7 ms on average to enter audit. The second line comes from JFR itself, and you can ignore it.
Still, 23 waits in twenty seconds is low for four threads and a single lock. The section on thresholds explains why.
exception-by-site: who throws exceptions
jfr view exception-by-site orders.jfr
Method Count
--------------------------------------------------------------- -----
java.lang.NumberFormatException.forInputString(String, int) 1,131
java.lang.invoke.MethodHandleNatives.resolve(MemberName, ...) 9
This is problem 1: 1,131 exceptions in twenty seconds, or about 57 per second. On JDK 25, jdk.JavaExceptionThrow is on by default, with a limit of 100 events per second. Above this limit, JFR only keeps part of them.
allocation-by-class and gc: memory pressure
jfr view allocation-by-class orders.jfr
Object Type Allocation Pressure
------------------------------------------------- -------------------
byte[] 43.89%
java.util.Formatter$FormatSpecifier 10.87%
java.lang.Object[] 7.14%
java.lang.StringBuilder 7.14%
java.lang.String 5.86%
java.util.Formatter$FixedString 4.67%
Formatter appears again, because each call to String.format creates new objects. jfr view gc shows the result: about one young collection per second, with pauses of about 1 to 5 ms:
Start GC ID Type Heap Before GC Heap After GC Longest Pause
-------- ----- --------------------------- -------------- ------------- -------------
15:55:01 0 Young Garbage Collection 18.1 MB 3.6 MB 4.83 ms
15:55:01 1 Young Garbage Collection 16.6 MB 4.4 MB 5.23 ms
15:55:02 2 Young Garbage Collection 19.4 MB 4.2 MB 1.59 ms
This is not a problem here. To read GC logs in detail, see Reading GC logs without external tools.
Thresholds: what JFR does not see
The jfr summary output shows 2 jdk.FileWrite events. But the application writes to the audit log more than 160 times per second.
JFR does not record everything, to limit its cost. Each event with a duration has a threshold. A file write is kept only if it takes more than one millisecond. Our writes go to the system on each flush(), but they stay in the file system cache. So they are fast, and JFR kept only two of them.
It is the same for locks. The threshold for jdk.JavaMonitorEnter is 20 ms. JFR does not keep shorter waits.
The JDK provides two configurations, default and profile. Here is the same test with settings=profile, over about ten seconds:
StackTrace Count Avg. Max.
--------------------------------------------------- ----- ------- -------
Orders.audit(String) 1,665 18.3 ms 59.1 ms
The count goes from 23 waits to 1,665. With default, only about a dozen waits went over 20 ms every ten seconds. So nearly all waits last between 10 and 20 ms. default does not see them, profile does.
Here are the differences between the two, read from the JDK 25 default.jfc and profile.jfc files:
| Event | default | profile |
|---|---|---|
jdk.ExecutionSample | every 20 ms | every 10 ms |
jdk.JavaMonitorEnter | 20 ms threshold | 10 ms threshold |
jdk.ThreadPark, jdk.ThreadSleep | 20 ms threshold | 10 ms threshold |
jdk.JavaExceptionThrow | 100 per second | 300 per second |
jdk.ObjectAllocationSample | 150 per second | 300 per second |
jdk.OldObjectSample | no stack trace | with stack trace |
So read a recording with care. If an event is missing, it does not mean nothing happened. It means nothing went over the threshold.
Setting your own thresholds
jfr configure creates a configuration file from the default configuration:
jfr configure locking-threshold=5ms memory-leaks=stack-traces --output prod.jfc
java -XX:StartFlightRecording:settings=prod.jfc,filename=custom.jfr ...
locking-threshold changes the threshold for locks, park, sleep and wait at once. With --verbose, the command prints each setting it writes.
You can also put these options directly in the flag, without a file:
java -XX:StartFlightRecording:locking-threshold=5ms,filename=custom.jfr ...
jfr help configure lists the available options. For example: gc, compiler, allocation-profiling, method-profiling, exceptions, memory-leaks, thread-dump, class-loading, and two options new in JDK 25, method-timing and method-trace.
memory-leaks=stack-traces adds the stack trace to objects that stay in the heap for a long time. This option finds problem 4. Here is the result on a profile recording of about ten seconds:
jfr view memory-leaks-by-site dump1.jfr
Alloc. Time Application Method Object Age Heap Usage
----------- ------------------------------------ ---------- ----------
15:55:59 N/A 10.3 s 3.1 MB
15:56:05 Orders.main(String[]) 3.86 s 4.2 MB
15:56:06 Orders.remember(String) 2.57 s 4.2 MB
Orders.remember does appear. The view also shows other lines that are not useful. Over ten seconds, every object still in memory looks like a leak. This view is useful on a long recording, when the same method comes back often. To confirm a leak, use a heap dump: Diagnosing a memory leak with a heap dump.
What it costs
Because of the sleep, our application waits more than it computes. To measure the cost of JFR, I used another version, with no pause inside the lock. It writes to /dev/null and keeps no history. The measurements were made on a Mac with 10 cores.
I first measured throughput, with and without JFR. The result was not usable. Without JFR, throughput varied from 107,000 to 136,000 orders per second from one run to the next. This variation was larger than the effect of JFR.
I then measured the time for a fixed amount of work. Four threads process 200,000 orders each, then the JVM exits. I did six runs per mode, alternating the modes:
| Mode | Average time | Minimum | Maximum |
|---|---|---|---|
| no JFR | 7.17 s | 6.26 s | 8.00 s |
default | 7.21 s | 6.05 s | 8.58 s |
profile | 7.56 s | 6.24 s | 8.40 s |
With default, the average difference is 0.5%. This is less than the normal variation between two runs. With profile, the average difference is 5.4%. It is also smaller than the variation between two runs of the same mode.
On this application, the cost of default cannot be measured. The cost of profile is slightly visible. For your application, you need to measure with your own load.
Always on in production
JFR is more useful when it runs all the time. When an incident happens, the last minutes are already recorded.
java -XX:StartFlightRecording:maxage=1h,maxsize=200m,dumponexit=true,filename=/var/log/app/app.jfr ...
maxage and maxsize limit the data JFR keeps: one hour and 200 MB maximum. JFR stores its data in a working folder, the repository. You can choose this folder with -XX:FlightRecorderOptions:repository=/path. dumponexit=true writes the file when the JVM stops.
I tested this last point, because a production JVM does not always stop normally.
| Shutdown | Result |
|---|---|
SIGTERM | file written, 350 kB |
OutOfMemoryError kills the only thread, the JVM exits | file written, 388 kB |
OutOfMemoryError with -XX:+ExitOnOutOfMemoryError | empty file, 0 bytes |
kill -9 | empty file, 0 bytes |
The ExitOnOutOfMemoryError case is important, because this flag is common in containers. The JVM stops at once, without running dumponexit. The file exists, but it is empty, as after a kill -9.
After a kill -9 or an ExitOnOutOfMemoryError, the repository contains a file. But jfr cannot read it:
jfr summary: could not read recording at .../2026_09_23_18_28_12.jfr. Recording file is stuck in locked stream state.
Another point about the OutOfMemoryError: the error does not appear as an event. jfr view jdk.JavaErrorThrow prints No events found. You need to look at the GCs. Just before the error, an old collection frees almost nothing:
18:25:09 16 Old Garbage Collection 63.5 MB 60.8 MB 6.08 ms
The last two collections empty the heap. The main thread has stopped, so its list is no longer used and the GC can remove it.
In practice, you need to get the recording before the JVM stops. Run jcmd <pid> JFR.dump filename=... when latency goes up, when an alert fires, or before you restart a pod.
A note on JFR.dump. It accepts begin=-10s or maxage=10s to keep only the end. But JFR stores its data in blocks, the chunks, and a dump always contains whole chunks. On a JVM that had been running for 60 seconds, JFR.dump begin=-10s returned all 60 seconds. The cut is per chunk, not per second.
Timing a method without redeploying
JDK 25 adds two options, method-timing and method-trace. The first one counts the calls to a method and measures how long they take. You can turn it on at startup:
java "-XX:StartFlightRecording:method-timing=Orders::invoice;Orders::audit,duration=10s,filename=timing.jfr" ...
jfr view method-timing timing.jfr
The quotes are required. The semicolon separates the methods, and without quotes the shell reads it as the end of the command.
Timed Method Invocations Minimum Time Average Time Maximum Time
------------------------- ----------- ------------ ------------ ------------
Orders.audit(String) 1,644 5.100000 ms 23.900000 ms 37.900000 ms
Orders.invoice(int) 1,648 0.048300 ms 0.209000 ms 16.500000 ms
Over ten seconds, audit takes 23.9 ms on average, including the wait for the lock. invoice takes 0.2 ms. So the time is lost in the lock, not in creating the invoice.
The option also works on a JVM that is already running:
jcmd <pid> JFR.start method-timing=Orders::invoice duration=8s filename=/tmp/live-timing.jfr
Timed Method Invocations Minimum Time Average Time Maximum Time
------------------------- ----------- ------------ ------------ ------------
Orders.invoice(int) 1,318 0.043500 ms 0.181000 ms 4.290000 ms
There is no code to add and no redeploy. It is a simple way to check whether a method is slow in production.
CPU time, still experimental
JDK 25 also adds a new type of sample, based on CPU time. It exists only on Linux. It is experimental and off by default:
jcmd <pid> JFR.start "jdk.CPUTimeSample#enabled=true" duration=8s filename=/tmp/cputime.jfr
jfr view cpu-time-hot-methods /tmp/cputime.jfr
Java Methods that Execute the Most from CPU Time Sampler (Experimental)
Method Samples Percent
--------------------------------------------------------------- ------- -------
java.io.FileOutputStream.writeBytes(byte[], int, int, boolean) 5 11.63%
Orders.lambda$main$0(String[]) 5 11.63%
java.lang.AbstractStringBuilder.append(String) 5 11.63%
java.lang.AbstractStringBuilder.append(char) 4 9.30%
FileOutputStream.writeBytes is on the first line. It is a native method, and it was not in the 25 lines of hot-methods above. jdk.ExecutionSample only takes threads that run Java code. Threads in native code have another event, jdk.NativeMethodSample, which counts them even when they wait. The new sample only counts the CPU time used, native code included. The view’s title shows that it is still experimental.
Your own events
An application can also create its own events. All it takes is a class that extends jdk.jfr.Event:
@Name("shop.Order")
@Label("Order")
@Category("Shop")
@StackTrace(false)
public class OrderEvent extends Event {
@Label("Customer")
String customer;
@Label("Lines")
int lines;
}
You use it around the code you want to measure:
OrderEvent event = new OrderEvent();
event.begin();
process(order);
event.customer = customer;
event.lines = lines;
event.commit();
The event is on by default. You read it like the others:
jfr view shop.Order shop.jfr
Order
Start Time Duration Event Thread Stack Trace Customer Lines
---------- -------- -------------- ------------- -------- ------
18:22:15 45.1 ms main N/A alice 35
18:22:15 27.0 ms main N/A bob 18
18:22:15 63.0 ms main N/A carol 53
The @Label annotations give the column titles. The event contains business data: the customer and the number of lines. It is in the same file as the GCs and the locks, with the same time scale. So you can compare a slow order with the GC pauses at the same moment.
Watch out for one detail. To set a threshold on your own event, you need a + in front of its name:
java "-XX:StartFlightRecording:+shop.Order#threshold=40ms,filename=shop40.jfr" ...
With the +, JFR keeps the 22 orders out of 60 that go over 40 ms. Without the +, it keeps all 60. The setting is ignored, with no error message. The + adds a setting for an event that is not in the .jfc file.
Reading JFR from the application
RecordingStream reads events as they arrive, inside the JVM itself:
RecordingStream rs = new RecordingStream();
rs.enable("shop.Order").withThreshold(Duration.ofMillis(40));
rs.onEvent("shop.Order", e ->
System.out.println("slow order: " + e.getString("customer")
+ ", " + e.getInt("lines") + " lines, " + e.getDuration().toMillis() + " ms"));
rs.startAsync();
slow order: carol, 53 lines, 63 ms
slow order: carol, 30 lines, 40 ms
slow order: alice, 48 lines, 55 ms
Instead of the println, you can update a metric or write a log line. JFR applies the 40 ms threshold, so the code only receives the slow orders.
During the test, the JVM did not stop at the end of main. The cause: the JFR Event Stream 1 thread is not a daemon thread. An rs.close() at the end of the program fixes the problem. On a server that runs all the time, this problem does not exist.
Before you send a recording
A JFR file does not only contain stacks. It also contains the environment variables, the system properties and the JVM command line. Here is a test with a password in an environment variable and a key passed with -D:
jfr print --events InitialEnvironmentVariable app.jfr
key = "DB_PASSWORD"
value = "s3cr3t-db"
The app.api.key property also appears, and a second time in the jvmArguments field of jdk.JVMInformation.
Before you attach a file to a ticket, use jfr scrub to remove these events:
jfr scrub --exclude-events InitialEnvironmentVariable,InitialSystemProperty,JVMInformation app.jfr clean.jfr
Removed events:
jdk.InitialEnvironmentVariable 68/68
jdk.InitialSystemProperty 16/16
jdk.JVMInformation 1/1
Both secrets were readable in the original file. They are no longer in the cleaned file. The rest of the recording does not change.
In short
JFR is part of the JDK. -XX:StartFlightRecording starts it with the JVM, and jcmd <pid> JFR.start starts it on a JVM that is already running. jfr view reads the files with no GUI: hot-methods, contention-by-site, exception-by-site, allocation-by-class, gc.
JFR only keeps what goes over a threshold. With default, a 15 ms wait on a lock does not appear. Use profile, or lower the threshold with locking-threshold.
In our measurement, the cost of default is smaller than the variation between two runs. Leave JFR running with maxage and maxsize, and get the data with JFR.dump before the JVM stops. dumponexit does not work after a kill -9 or with ExitOnOutOfMemoryError.
JDK 25 adds method-timing, which measures the duration of a method on a JVM that is already running. It also adds a sample based on CPU time, which is still experimental.
Your own events are recorded with the JVM’s events. Do not forget the + for their settings. And use jfr scrub before you share a file.
What next?
JFR shows what happened inside the JVM. To see in detail where the CPU time goes, with native frames and a flamegraph, use async-profiler: Profiling the JVM with async-profiler.