You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何统计Skylake架构L2缓存预取总行数及无用预取占比?

Calculating Useless L2 Hardware Prefetch Ratio on Skylake

Great question—figuring out how to compute the ratio of unused L2 hardware prefetches to total prefetches on Skylake is tricky because the direct "total L2 hardware prefetches" event isn't immediately obvious. Here's how to tackle it:

Step 1: Identify the Right Events

You already know about l2_lines_out.useless_hwpf (counts lines hardware-prefetched into L2 that were evicted without being used). For the total number of hardware prefetches into L2, use the l2_lines_in.hwpf event. This event tracks all lines loaded into L2 via hardware prefetching, which is exactly the denominator you need for your ratio.

Step 2: Verify Event Availability

First, confirm these events exist on your system with:

perf list | grep -E "l2_lines_in.hwpf|l2_lines_out.useless_hwpf"

If l2_lines_in.hwpf doesn't show up, try looking for similar events with broader filters:

perf list | grep -i "l2.*prefetch"

On some Skylake systems, you might see l2_lines_in.all_hwpf instead—this is equivalent, as it counts all hardware prefetch lines entering L2.

Step 3: Measure and Calculate the Ratio

Use perf stat to collect both events during your program's execution:

perf stat -e l2_lines_in.hwpf,l2_lines_out.useless_hwpf ./your_target_program

Once you have the counts, compute the ratio with:

Useless Prefetch Ratio = (l2_lines_out.useless_hwpf / l2_lines_in.hwpf) * 100%

For example, if l2_lines_in.hwpf returns 15000 and l2_lines_out.useless_hwpf returns 3000, your ratio is 20%.

Notes for Edge Cases

  • Kernel Version: Ensure you're running a kernel 4.10 or newer—l2_lines_out.useless_hwpf was added around this release, so older kernels won't support it.
  • Software vs Hardware Prefetches: l2_lines_in.hwpf only counts hardware-initiated prefetches. If you want to include software prefetches (e.g., via prefetch intrinsics), add l2_lines_in.swpf to get the total prefetch count.
  • Event Aliases: Some distros or perf versions might use slightly different event names. If you're stuck, check the Intel Skylake PMU documentation for raw event codes (you can use raw codes directly with perf via rXX syntax, e.g., r04c6 for l2_lines_in.hwpf on some Skylake variants).

内容的提问来源于stack exchange,提问作者onlycparra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 17:12:43