Skylake处理器下,perf_event_open应优先选用哪种Intel PMU替代PERF_COUNT_HW_INSTRUCTIONS?
Great question—let’s dive into how to handle instruction counting on Skylake processors with perf_event_open, since Intel’s PMU has specific quirks you’ll want to leverage for reliable performance analysis.
Why You Should Be Cautious with PERF_COUNT_HW_INSTRUCTIONS
PERF_COUNT_HW_INSTRUCTIONS is a cross-architecture abstract hardware event, and on Skylake it maps to INST_RETIRED.ANY under the hood. The manual’s warning exists for two key reasons:
- Cross-architecture variability: This abstract event can map to slightly different underlying PMU events across Intel microarchitectures (even within Skylake variants), leading to less consistent or precise counts than using Intel’s specific events directly.
- Limited filtering: You can’t easily narrow counts to user-only or kernel-only instructions, or target specific instruction types (like vector ops) with this abstract event—features critical for targeted performance debugging.
Preferred Instruction-Counting Events on Skylake
For reliable instruction counting on Skylake, skip the abstract event and go straight to Intel-specific PMU events. Here’s what to pick based on your goal:
INST_RETIRED.ANY: Counts all retired instructions (user + kernel). This is the exact eventPERF_COUNT_HW_INSTRUCTIONSuses, but specifying it directly guarantees consistency on Skylake.INST_RETIRED.USER: Only counts user-space retired instructions. Perfect for focusing on your application’s execution without kernel noise.INST_RETIRED.KERNEL: Counts only kernel-space retired instructions. Use this if you’re debugging system call overhead or kernel paths affecting your app.INST_RETIRED.PREC_DISTINCT: Counts "unique" retired instruction addresses (great for analyzing instruction reuse or cache behavior if that’s part of your profile).
These events give you precise control and avoid the ambiguity of the cross-architecture abstract event.
Using Intel-Specific PMU Events with perf_event_open
Absolutely—Skylake fully supports Intel’s Performance Monitoring Unit (PMU), and using its specific events is straightforward with perf_event_open:
- Set the raw event type: Configure your
perf_event_attrstruct withtype = PERF_TYPE_RAWto tell the kernel you’re using a hardware-specific event code. - Specify the event code: Use the raw event encoding from Intel’s official docs (Skylake’s codes live in Volume 3B of the Intel Software Developer’s Manual). For example:
INST_RETIRED.ANY→0x00c0(Event 0xC0, Umask 0x00)
- Add filters: Use fields like
exclude_kernelorexclude_userto narrow the count to the privilege level you care about (this is cleaner than messing with extra bits in the event code). - Open and manage the event: Call
perf_event_open, then useioctlto enable/disable the event andreadto fetch the count.
Here’s a quick code snippet to illustrate:
#include <stdio.h> #include <stdlib.h> #include <unistd.h> #include <sys/ioctl.h> #include <linux/perf_event.h> #include <asm/unistd.h> static long perf_event_open(struct perf_event_attr *hw_event, pid_t pid, int cpu, int group_fd, unsigned long flags) { return syscall(__NR_perf_event_open, hw_event, pid, cpu, group_fd, flags); } int main(int argc, char **argv) { struct perf_event_attr attr = {0}; attr.type = PERF_TYPE_RAW; attr.size = sizeof(attr); attr.config = 0x00c0; // INST_RETIRED.ANY attr.exclude_kernel = 1; // Only count user-space instructions attr.disabled = 1; // Start with the event disabled int fd = perf_event_open(&attr, 0, -1, -1, 0); if (fd == -1) { perror("perf_event_open failed"); exit(EXIT_FAILURE); } // Enable the event ioctl(fd, PERF_EVENT_IOC_ENABLE, 0); // Run your target code here—example: loop a million times for (long i = 0; i < 1000000; i++); // Read the count uint64_t count; read(fd, &count, sizeof(count)); printf("Retired user instructions: %lu\n", count); // Disable and clean up ioctl(fd, PERF_EVENT_IOC_DISABLE, 0); close(fd); return 0; }
Quick Notes to Keep in Mind
- Permissions: You’ll need
CAP_PERFMON(Linux 5.8+) or root access to use PMU events—regular users can’t access hardware counters by default. - Hyper-Threading: Skylake’s logical cores each have their own PMU counters, so counts are isolated to the core your process is running on.
- Counter Limits: Skylake has 4 general-purpose PMU counters per core, so you can’t track more than 4 independent events at once without multiplexing (which can introduce minor overhead).
内容的提问来源于stack exchange,提问作者FloriUni

