如何用perf_event_open模拟perf record -g,实现全CPU定时调用栈记录?
Absolutely, let's walk through how to implement call stack sampling with perf_event_open—I've been there, and the man pages are definitely sparse on practical examples. Below is a working C example that mimics the core behavior of perf record -g, plus breakdowns of the key parts so you understand what's happening.
Key Prerequisites & Concepts
Before diving into code, keep these in mind:
- You'll need root or CAP_PERFMON/CAP_SYS_ADMIN privileges to run the program.
- For reliable user-space call stacks, compile your target programs with
-fno-omit-frame-pointerand consider disabling address space randomization temporarily (echo 0 > /proc/sys/kernel/randomize_va_space). - The
perf_event_openAPI uses a ring buffer (via mmap) to deliver samples efficiently—this is way faster than reading viaread(). - We need to set
PERF_SAMPLE_CALLCHAINalongsidePERF_SAMPLE_IP(and others) to get full call stack context.
Complete Example Code
#include <stdio.h> #include <stdlib.h> #include <unistd.h> #include <sys/mman.h> #include <linux/perf_event.h> #include <asm/unistd.h> #include <string.h> #include <sys/ioctl.h> #define PAGE_SIZE 4096 #define CALLCHAIN_MAX_DEPTH 64 // Wrapper for perf_event_open syscall (since it's not in standard libc) static long perf_event_open(struct perf_event_attr *hw_event, pid_t pid, int cpu, int group_fd, unsigned long flags) { return syscall(__NR_perf_event_open, hw_event, pid, cpu, group_fd, flags); } // Parse and print a call chain from a sample static void print_callchain(const struct perf_callchain_entry *chain) { printf("Call chain (depth: %u):\n", chain->nr); for (unsigned int i = 0; i < chain->nr; i++) { printf(" %p\n", (void *)chain->ips[i]); // Note: To resolve addresses to symbols, you'd use libelf/libdw here // to read the target's ELF files and look up symbols. } } int main(int argc, char **argv) { struct perf_event_attr attr; int fd; void *mmap_buf; size_t mmap_size; struct perf_event_mmap_page *header; struct perf_event_header *event_hdr; char *data_ptr; // Step 1: Initialize perf_event_attr memset(&attr, 0, sizeof(attr)); attr.type = PERF_TYPE_HARDWARE; attr.config = PERF_COUNT_HW_CPU_CYCLES; // Sample on CPU cycles attr.size = sizeof(attr); // Configure sampling: we want IP, call chain, and time attr.sample_type = PERF_SAMPLE_IP | PERF_SAMPLE_CALLCHAIN | PERF_SAMPLE_TIME; attr.sample_period = 1000000; // Sample every 1M cycles (adjust as needed) attr.freq = 0; // Use period-based sampling instead of frequency // Enable call chain sampling for both kernel and user space attr.callchain_kernel = CALLCHAIN_MAX_DEPTH; attr.callchain_user = CALLCHAIN_MAX_DEPTH; attr.disabled = 1; // Start disabled, enable later with ioctl attr.exclude_kernel = 0; // Include kernel stacks (set to 1 to exclude) attr.exclude_user = 0; // Include user stacks (set to 1 to exclude) // Step 2: Open the perf event for ALL CPUs (pid = -1, cpu = -1) fd = perf_event_open(&attr, -1, -1, -1, PERF_FLAG_FD_CLOEXEC); if (fd == -1) { perror("perf_event_open failed"); exit(EXIT_FAILURE); } // Step 3: Mmap the ring buffer mmap_size = PAGE_SIZE * 2; // 1 page for metadata, 1 page for samples mmap_buf = mmap(NULL, mmap_size, PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0); if (mmap_buf == MAP_FAILED) { perror("mmap failed"); close(fd); exit(EXIT_FAILURE); } header = (struct perf_event_mmap_page *)mmap_buf; // Step 4: Enable the event ioctl(fd, PERF_EVENT_IOC_ENABLE, 0); printf("Sampling call stacks... Press Ctrl+C to stop.\n"); // Step 5: Poll the ring buffer for samples while (1) { // Wait for new data (optional, but reduces busy waiting) while (header->data_head == header->data_tail) { usleep(1000); } // Get the current head and tail (memory barriers are important here!) uint64_t head = header->data_head; uint64_t tail = header->data_tail; __sync_synchronize(); // Iterate through all available events data_ptr = (char *)mmap_buf + PAGE_SIZE; // Skip metadata page while (tail != head) { event_hdr = (struct perf_event_header *)(data_ptr + tail % (mmap_size - PAGE_SIZE)); // Process only sample events if (event_hdr->type == PERF_RECORD_SAMPLE) { struct perf_sample_data sample; struct perf_callchain_entry chain; // Parse the sample data memset(&sample, 0, sizeof(sample)); perf_event__parse_sample(event_hdr, &attr, &sample, NULL, 0); // Extract the call chain if (sample.callchain) { chain = *(struct perf_callchain_entry *)sample.callchain; print_callchain(&chain); } } // Move tail forward tail += event_hdr->size; __sync_synchronize(); header->data_tail = tail; } } // Cleanup (never reached in this example, but good practice) ioctl(fd, PERF_EVENT_IOC_DISABLE, 0); munmap(mmap_buf, mmap_size); close(fd); return 0; }
Key Code Breakdown
Let's unpack the critical parts:
perf_event_attrInitialization:- We use
PERF_TYPE_HARDWAREwithPERF_COUNT_HW_CPU_CYCLESto sample on CPU cycles (likeperf record -e cycles). sample_typeincludesPERF_SAMPLE_IP(to get the instruction pointer) andPERF_SAMPLE_CALLCHAIN(to get the stack).callchain_kernelandcallchain_userset the maximum depth for kernel/user stacks.
- We use
Ring Buffer via mmap:
- The mmap buffer has two parts: a metadata page (
struct perf_event_mmap_page) and a data page for samples. - We use
data_headanddata_tailto track which samples have been processed—these are atomic variables, so we use memory barriers (__sync_synchronize()) to avoid race conditions.
- The mmap buffer has two parts: a metadata page (
Sample Parsing:
perf_event__parse_sample(from the perf internal headers, but you can implement your own parser if needed) extracts the call chain from the sample data.- The call chain is stored as an array of instruction pointers—you'd need to use
libelforlibdwto resolve these addresses to function names (likeperf reportdoes).
Next Steps to Expand This
- Symbol Resolution: Add code to read ELF files (using
libelf) and resolve IP addresses to function names. - Filter by PID: Modify the
perf_event_opencall to target a specific PID instead of all CPUs. - Adjust Sampling Parameters: Play with
sample_periodor switch to frequency-based sampling (attr.freq = 1,attr.sample_freq = 1000for 1000 samples per second). - Handle Different Event Types: Add support for software events (like
PERF_TYPE_SOFTWARE) or tracepoints.
内容的提问来源于stack exchange,提问作者Michael McLoughlin
相关产品推荐
相关产品推荐

