关于perf report输出解读及单线程多栈迹问题的技术问询
Hey there! Let's break down your questions one by one— I’ve spent plenty of time digging into perf reports for Java applications, so I’ve got some practical insights to share.
1. Understanding the First Column in perf report
The first column represents the percentage of total samples that hit this function or call stack frame. When you run perf -g -p <pid>, perf captures full call stacks at regular intervals (default is every 1ms), and this percentage is calculated based on how often a particular frame or stack appears in those samples.
- For top-level functions (the first line of a stack trace), this number reflects how often perf caught the thread executing exactly that function.
- For nested stack frames (lines below the top), it’s the share of samples where this frame appeared anywhere in the call stack—meaning all execution paths that passed through this function contribute to this percentage.
Think of it as a "time weight": the higher the value, the more your process is spending cycles in that code path.
2. Analyzing Java_* Calls Taking >90% of Samples
First off, Java_* functions are almost always JNI bridge methods—they’re the glue between Java code and native (C/C++) logic, either from JVM internals or custom native libraries your app uses. Here’s how to dig into the bottleneck:
- Identify what the Java_ function does*: Run
perf annotate <Java_*_function_name>to inspect the assembly code. If you have JVM debug symbols installed, this will map directly to JVM internal operations (like GC work, native method dispatch, or JIT compiler tasks). - Link native stacks to Java code: Modern JVMs disable frame pointers by default (
-XX:-PreserveFramePointer), which blocks perf from resolving Java stack frames alongside native ones. Add-XX:+PreserveFramePointerto your JVM args, restart the process, and re-run perf—you’ll now see full stacks that include both Java methods and theJava_*JNI bridges. - Check if it’s expected or a problem:
- If the
Java_*calls tie to a custom native library, profile your Java code to see if you’re calling it excessively (use tools likejstack,jcmd, or async-profiler for Java-level context). - If it’s JVM-internal (e.g.,
Java_java_lang_System_arraycopyor GC-related functions), audit your JVM settings: Are you running frequent minor GCs? Is your heap sized appropriately?
- If the
- Use raw stack data for clarity: Export all samples with
perf script, then use a stack-collapsing tool (likestackcollapse-perf.pl) to aggregate similar stacks. This will highlight the exact call chains leading to those dominantJava_*calls.
3. Why Multiple Stack Traces for a Single-Threaded Process?
This is a common gotcha! Even a single-threaded process doesn’t stick to one call stack forever—remember, perf takes samples at intervals, not continuous tracking. Here’s why you see multiple entries:
- The thread switches execution paths: Your single thread might bounce between business logic, system calls (like
read/write), or JVM internal tasks (like class loading) between perf samples. Each of these states has a unique call stack. - Different parent calls to the same function: Suppose a function
Xis called by bothYandZ. Perf will capture two separate stack traces:X <- YandX <- Z, each with their own sample percentage. - Signal handling: Even single-threaded processes handle signals (like those sent by the JVM for GC or profiling). When a signal is delivered, the thread’s stack switches to the signal handler, which shows up as a distinct stack trace in perf.
In short: perf is showing you all the different execution states your single thread was in during the sampling period, not just a single snapshot.
内容的提问来源于stack exchange,提问作者J. Doe

