You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用perf收集可读堆栈追踪?Linux生产环境C++程序采样剖析

Got it, totally get why you’re avoiding gdb in a high-load production environment—nothing kills critical service performance faster than a debugger halting or dragging down execution. Perf is exactly the right tool here: it’s low-overhead, production-safe, and built to collect actionable stack trace data without disrupting your workload.

Using perf for Low-Overhead, Readable Stack Trace Profiling in Production

Let’s walk through the exact commands and tweaks you need to get clear, useful stack traces from your C++ program:

Step 1: Collect Stack Trace Data with Minimal Overhead

Run this command to sample your target process (replace <PID> with your C++ program’s process ID, and <duration> with how long you want to profile, e.g., 60 for 1 minute):

perf record -g --call-graph dwarf,16384 -p <PID> -o perf.prod.data --quiet sleep <duration>

Let’s break down the critical flags:

  • -g: Enables call graph (stack trace) collection.
  • --call-graph dwarf,16384: Uses DWARF debug info to capture accurate stack traces—this is far more reliable than the default frame pointer method for optimized C++ code. The 16384 sets the maximum stack depth to capture; adjust this if you need to trace deeper into nested calls.
  • -p <PID>: Targets only your specific process, avoiding unnecessary profiling of other system services.
  • -o perf.prod.data: Saves data to a named file (avoids overwriting the default perf.data if you run multiple profiles).
  • --quiet: Suppresses verbose output during collection to minimize overhead.

Step 2: Generate Readable, Demangled Stack Traces

By default, perf will show mangled C++ function names (like _ZN7MyClass8myMethodEv). Fix this and format the output for readability with:

perf report --demangle=smart --stdio -n -g

Key flags here:

  • --demangle=smart: Automatically translates mangled C++ names to human-readable ones (e.g., MyClass::myMethod()).
  • --stdio: Outputs to the terminal in clean text format, perfect for logging or remote analysis.
  • -n: Shows the number of samples per function, so you can instantly spot which code paths are consuming the most CPU.
  • -g: Displays the full call graph alongside each function, so you can trace execution flow from entry points down to low-level calls.

Production-Safe Best Practices

To keep profiling impact minimal in high-load environments:

  • Adjust sample rate: Perf defaults to 99 samples/second. If even that feels too aggressive, lower it with -F <rate> (e.g., -F 50 for 50 samples/second). Just balance rate with data quality—lower rates mean fewer samples, so don’t go too low.
  • Compile with minimal debug symbols: Even in production, compile your C++ code with -g1 (not full -g). This adds negligible overhead but provides enough debug info for perf to resolve stack traces correctly. If you can’t recompile, ship separate debug symbols and point perf to them with --symfs <path-to-debug-symbols>.
  • Filter to your binary: If you only care about stack traces from your program (not system libraries), filter during reporting:
    perf report --demangle=smart --stdio -g --dsos <path-to-your-c++-binary>
    

Perf’s overhead is typically well under 5% (often even lower with adjusted sample rates), so it’s safe to run in production without causing noticeable slowdowns.

内容的提问来源于stack exchange,提问作者ks1322

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:28:07