You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Linux perf合理采样已执行/退役指令?解决采样偏差问题

Great questions—this gets to the heart of how perf's hardware sampling works, especially around instruction-level events and the infamous "skid" problem. Let's break this down step by step.

1. Sampling Executed vs. Retired Instructions with perf

First, let's clarify the key difference between the two types of instruction events:

  • Executed instructions: Count all instructions that reach the CPU's execution units, including those speculatively executed but later discarded (e.g., due to branch mispredictions). Use events like inst_executed.any (check perf list for exact names on your CPU).
  • Retired instructions: Only count instructions that successfully complete and commit results to the architectural state (no speculative work included). This is the meaningful "instructions done" metric most developers care about. Common events are inst_retired.any_p (any retired instruction, kernel or user) or the generic alias instructions.

Basic Usage

  • To record samples for retired instructions:
    perf record -e inst_retired.any_p -c 10000 ./your_program
    
    The -c flag sets the sampling interval: perf takes a sample every 10,000 retired instructions.
  • To focus only on user-space instructions (avoid kernel noise), add the :u suffix:
    perf record -e inst_retired.any_p:u -c 10000 ./your_program
    
  • For executed (including speculative) instructions, swap the event:
    perf record -e inst_executed.any:u -c 10000 ./your_program
    

2. Fixing Biased Sampling (The Skid Problem)

Your issue—samples clustering on div or dec in the loop—is classic sampling skid. Here's why it happens:
Perf uses hardware counter overflow interrupts by default. When the counter hits your -c threshold, the CPU sends an interrupt—but by the time the interrupt is handled, the CPU has already advanced its instruction pointer (IP) past the instruction that actually triggered the overflow. This "skid" is far worse for long-latency instructions like div (which takes dozens of cycles): the CPU is still processing the div when the counter overflows, so the interrupt catches the IP at div or the following dec instead of earlier loop instructions.

Even though every instruction in your loop retires once per iteration, skid makes it look like div is responsible for nearly all samples.

Solutions for Accurate Instruction Sampling

a. Use Precision Event-Based Sampling (PEBS)

Intel CPUs (and AMD's equivalent IBS) support PEBS, which captures the exact IP of the instruction that caused the counter overflow at retirement time, eliminating skid entirely. To enable it, add the :p suffix to your event:

perf record -e inst_retired.any_p:p:u -c 1000 ./your_program
  • The :p flag enables PEBS precision sampling.
  • We use a smaller -c (1000 instead of 10000) here because PEBS has lower overhead for frequent samples—adjust based on your program's size and performance needs.

For your loop example, PEBS will distribute samples evenly across mov, xor, div, dec, and jne—matching the actual retirement rate of each instruction.

b. Reduce the Sampling Interval

If PEBS isn't available (e.g., on older CPUs), shrink the -c value to minimize skid. A smaller interval means the counter overflows more often, so the CPU can't advance as far past the triggering instruction before the interrupt. For example:

perf record -e inst_retired.any_p:u -c 1000 ./your_program

This won't eliminate skid entirely, but it will reduce the bias toward long-latency instructions. Just note that very small -c values can add overhead, so test to find a balance between precision and performance.

c. Verify Event Support

Check if your CPU supports PEBS for your target event with:

perf list | grep inst_retired

Look for entries with :p in the description—this confirms precision sampling is available. You can also run perf record -v to see if PEBS is being enabled (look for messages like "Using PEBS for sampling").

Final Notes

  • Always prefer retired instruction events (inst_retired.*) over executed ones unless you specifically need to measure speculative work—retired instructions reflect the actual, committed work your program does.
  • For user-space-only profiling, always add the :u suffix to avoid kernel instructions skewing your results.

内容的提问来源于stack exchange,提问作者BeeOnRope

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 15:18:09