perf record的--call-graph选项:LBR、DWARF与fp的区别
Awesome question! When you're using perf record to dig into call graphs, the --call-graph flag's three options—fp, dwarf, and lbr—each have their own superpowers and tradeoffs, depending on what you need. Let's break them down clearly:
fp (Frame Pointer) This is the old-school, no-frills option. It relies on the CPU's frame pointer register (usually rbp on x86_64) to walk the call stack. When a function is called, the old frame pointer gets pushed onto the stack, and the new one points to the current stack frame. By following this chain of pointers, perf can reconstruct the call stack without much fuss.
- Pros: Blazingly fast—almost zero performance overhead, since it's just parsing existing register and stack data. No extra debug info needed either.
- Cons: It breaks if your code was compiled with frame pointer omission (like GCC's
-fomit-frame-pointer, which is default in many modern builds). When that happens, the call stack will be incomplete or totally wrong. - Best for: Code that hasn't been compiled with frame pointer optimizations, or when you need minimal overhead and don't mind the accuracy tradeoff if optimizations are enabled.
dwarf This is the workhorse for accurate call graphs, especially with optimized code. It uses DWARF debugging information (generated when you compile with -g) to reconstruct the stack. DWARF has detailed metadata that tells perf how to map the current stack state back to the calling functions—even without a frame pointer.
- Pros: Super accurate. It handles functions compiled with
-fomit-frame-pointer, tracks inline functions, and works with dynamic stack allocations. If you need a complete, correct call stack, this is your go-to. - Cons: Higher performance overhead.
perfhas to capture more stack data during recording, and parsing DWARF info during post-processing takes time. Also, you must have debug symbols available for your code (either in the binary or separate debug files). - Best for: When accuracy matters most, like debugging performance issues in optimized code, or tracking down where inline functions are contributing to overhead.
lbr (Last Branch Record) This one leverages CPU hardware magic. Modern Intel (Nehalem+) and AMD (Zen+) CPUs have a built-in feature that records the most recent branch jumps—like function calls and returns. perf reads this hardware buffer to build the call stack directly from the CPU's own records.
- Pros: Extremely fast, almost as low-overhead as
fp. No need for debug info or frame pointers, since it's using raw hardware data. - Cons: Stack depth is limited (usually 32-64 branches, depending on the CPU). If your call stack is deeper than that, you'll lose the upper layers. It also can't track inline functions (since there's no branch jump when a function is inlined). And of course, it only works on CPUs that support LBR.
- Best for: Short call stacks, when you need minimal overhead and have a compatible CPU. Great for quick profiling runs where you don't need deep stack traces.
Quick Cheat Sheet
| Option | Speed | Accuracy | Requires Debug Info? | CPU-Dependent? | Works with Frame Pointer Omission? |
|---|---|---|---|---|---|
fp | 🚀 Fast | 🟡 Inconsistent (breaks with optimizations) | ❌ No | ❌ No | ❌ No |
dwarf | 🐢 Moderate | 🟢 High | ✅ Yes | ❌ No | ✅ Yes |
lbr | 🚀 Fast | 🟡 Good (limited depth) | ❌ No | ✅ Yes | ✅ Yes |
内容的提问来源于stack exchange,提问作者The flash

