You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用perf_event_open模拟perf record -g,实现全CPU定时调用栈记录?

Absolutely, let's walk through how to implement call stack sampling with perf_event_open—I've been there, and the man pages are definitely sparse on practical examples. Below is a working C example that mimics the core behavior of perf record -g, plus breakdowns of the key parts so you understand what's happening.

Key Prerequisites & Concepts

Before diving into code, keep these in mind:

  • You'll need root or CAP_PERFMON/CAP_SYS_ADMIN privileges to run the program.
  • For reliable user-space call stacks, compile your target programs with -fno-omit-frame-pointer and consider disabling address space randomization temporarily (echo 0 > /proc/sys/kernel/randomize_va_space).
  • The perf_event_open API uses a ring buffer (via mmap) to deliver samples efficiently—this is way faster than reading via read().
  • We need to set PERF_SAMPLE_CALLCHAIN alongside PERF_SAMPLE_IP (and others) to get full call stack context.
Complete Example Code
#include <stdio.h>
#include <stdlib.h>
#include <unistd.h>
#include <sys/mman.h>
#include <linux/perf_event.h>
#include <asm/unistd.h>
#include <string.h>
#include <sys/ioctl.h>

#define PAGE_SIZE 4096
#define CALLCHAIN_MAX_DEPTH 64

// Wrapper for perf_event_open syscall (since it's not in standard libc)
static long perf_event_open(struct perf_event_attr *hw_event, pid_t pid,
                            int cpu, int group_fd, unsigned long flags) {
    return syscall(__NR_perf_event_open, hw_event, pid, cpu, group_fd, flags);
}

// Parse and print a call chain from a sample
static void print_callchain(const struct perf_callchain_entry *chain) {
    printf("Call chain (depth: %u):\n", chain->nr);
    for (unsigned int i = 0; i < chain->nr; i++) {
        printf("  %p\n", (void *)chain->ips[i]);
        // Note: To resolve addresses to symbols, you'd use libelf/libdw here
        // to read the target's ELF files and look up symbols.
    }
}

int main(int argc, char **argv) {
    struct perf_event_attr attr;
    int fd;
    void *mmap_buf;
    size_t mmap_size;
    struct perf_event_mmap_page *header;
    struct perf_event_header *event_hdr;
    char *data_ptr;

    // Step 1: Initialize perf_event_attr
    memset(&attr, 0, sizeof(attr));
    attr.type = PERF_TYPE_HARDWARE;
    attr.config = PERF_COUNT_HW_CPU_CYCLES; // Sample on CPU cycles
    attr.size = sizeof(attr);
    
    // Configure sampling: we want IP, call chain, and time
    attr.sample_type = PERF_SAMPLE_IP | PERF_SAMPLE_CALLCHAIN | PERF_SAMPLE_TIME;
    attr.sample_period = 1000000; // Sample every 1M cycles (adjust as needed)
    attr.freq = 0; // Use period-based sampling instead of frequency
    
    // Enable call chain sampling for both kernel and user space
    attr.callchain_kernel = CALLCHAIN_MAX_DEPTH;
    attr.callchain_user = CALLCHAIN_MAX_DEPTH;
    
    attr.disabled = 1; // Start disabled, enable later with ioctl
    attr.exclude_kernel = 0; // Include kernel stacks (set to 1 to exclude)
    attr.exclude_user = 0; // Include user stacks (set to 1 to exclude)

    // Step 2: Open the perf event for ALL CPUs (pid = -1, cpu = -1)
    fd = perf_event_open(&attr, -1, -1, -1, PERF_FLAG_FD_CLOEXEC);
    if (fd == -1) {
        perror("perf_event_open failed");
        exit(EXIT_FAILURE);
    }

    // Step 3: Mmap the ring buffer
    mmap_size = PAGE_SIZE * 2; // 1 page for metadata, 1 page for samples
    mmap_buf = mmap(NULL, mmap_size, PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
    if (mmap_buf == MAP_FAILED) {
        perror("mmap failed");
        close(fd);
        exit(EXIT_FAILURE);
    }
    header = (struct perf_event_mmap_page *)mmap_buf;

    // Step 4: Enable the event
    ioctl(fd, PERF_EVENT_IOC_ENABLE, 0);
    printf("Sampling call stacks... Press Ctrl+C to stop.\n");

    // Step 5: Poll the ring buffer for samples
    while (1) {
        // Wait for new data (optional, but reduces busy waiting)
        while (header->data_head == header->data_tail) {
            usleep(1000);
        }

        // Get the current head and tail (memory barriers are important here!)
        uint64_t head = header->data_head;
        uint64_t tail = header->data_tail;
        __sync_synchronize();

        // Iterate through all available events
        data_ptr = (char *)mmap_buf + PAGE_SIZE; // Skip metadata page
        while (tail != head) {
            event_hdr = (struct perf_event_header *)(data_ptr + tail % (mmap_size - PAGE_SIZE));
            
            // Process only sample events
            if (event_hdr->type == PERF_RECORD_SAMPLE) {
                struct perf_sample_data sample;
                struct perf_callchain_entry chain;

                // Parse the sample data
                memset(&sample, 0, sizeof(sample));
                perf_event__parse_sample(event_hdr, &attr, &sample, NULL, 0);
                
                // Extract the call chain
                if (sample.callchain) {
                    chain = *(struct perf_callchain_entry *)sample.callchain;
                    print_callchain(&chain);
                }
            }

            // Move tail forward
            tail += event_hdr->size;
            __sync_synchronize();
            header->data_tail = tail;
        }
    }

    // Cleanup (never reached in this example, but good practice)
    ioctl(fd, PERF_EVENT_IOC_DISABLE, 0);
    munmap(mmap_buf, mmap_size);
    close(fd);
    return 0;
}
Key Code Breakdown

Let's unpack the critical parts:

  1. perf_event_attr Initialization:

    • We use PERF_TYPE_HARDWARE with PERF_COUNT_HW_CPU_CYCLES to sample on CPU cycles (like perf record -e cycles).
    • sample_type includes PERF_SAMPLE_IP (to get the instruction pointer) and PERF_SAMPLE_CALLCHAIN (to get the stack).
    • callchain_kernel and callchain_user set the maximum depth for kernel/user stacks.
  2. Ring Buffer via mmap:

    • The mmap buffer has two parts: a metadata page (struct perf_event_mmap_page) and a data page for samples.
    • We use data_head and data_tail to track which samples have been processed—these are atomic variables, so we use memory barriers (__sync_synchronize()) to avoid race conditions.
  3. Sample Parsing:

    • perf_event__parse_sample (from the perf internal headers, but you can implement your own parser if needed) extracts the call chain from the sample data.
    • The call chain is stored as an array of instruction pointers—you'd need to use libelf or libdw to resolve these addresses to function names (like perf report does).
Next Steps to Expand This
  • Symbol Resolution: Add code to read ELF files (using libelf) and resolve IP addresses to function names.
  • Filter by PID: Modify the perf_event_open call to target a specific PID instead of all CPUs.
  • Adjust Sampling Parameters: Play with sample_period or switch to frequency-based sampling (attr.freq = 1, attr.sample_freq = 1000 for 1000 samples per second).
  • Handle Different Event Types: Add support for software events (like PERF_TYPE_SOFTWARE) or tracepoints.

内容的提问来源于stack exchange,提问作者Michael McLoughlin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:58:48