性能分析器如何采样运行程序?求多平台实用实现思路
Great question! Your core intuition about using timer interrupts to sample instruction pointers (IP) and find hot code is exactly how most production-grade sampling profilers work. Let’s dive into practical, actionable implementation approaches across Linux/Unix and Windows, with concrete details you can adapt.
1. Kernel-Level Sampling (Most Reliable & Precise)
The gold standard here is using the perf_event_open() system call, which lets you leverage the kernel's built-in sampling infrastructure without writing custom interrupt handlers:
- How it works: Configure a software timer event (e.g.,
PERF_COUNT_SW_CPU_CLOCK) to trigger at your desired interval (e.g., 1ms). The kernel automatically captures the current IP (and optionally stack traces) whenever the timer fires, writing data to a memory-mapped buffer your user-space program can read. - Concrete steps:
- Initialize a
perf_event_attrstruct to define your sampling parameters (event type, sample period, data to capture). - Call
perf_event_open()to bind the event to your target process/thread. - Map the kernel's sampling buffer into your process address space with
mmap(). - Read the buffer periodically to collect IP samples and tally their frequencies.
- Initialize a
- Simplified code snippet:
#include <linux/perf_event.h> #include <sys/ioctl.h> #include <sys/mman.h> struct perf_event_attr attr = {0}; attr.type = PERF_TYPE_SOFTWARE; attr.config = PERF_COUNT_SW_CPU_CLOCK; attr.sample_type = PERF_SAMPLE_IP; attr.sample_period = 1000000; // Sample every 1ms (nanoseconds) attr.size = sizeof(attr); // Replace target_pid with your process ID, or 0 for the current process int perf_fd = perf_event_open(&attr, target_pid, -1, -1, 0); if (perf_fd == -1) { /* Handle error */ } // Map the sampling buffer void* buf = mmap(NULL, 4096, PROT_READ | PROT_WRITE, MAP_SHARED, perf_fd, 0); // Now read samples from buf and count IP occurrences
2. User-Space Alternative (Simpler, Lower Precision)
If you want to avoid kernel-level code, use a timer-based signal handler:
- How it works: Use
setitimer()to set up anITIMER_PROFtimer, which sends aSIGPROFsignal to the process at your chosen interval. In the signal handler, usegetcontext()to grab the current CPU context and extract the IP. - Key notes:
- Only use async-signal-safe functions in the handler (e.g., avoid
printf()—use atomic counters or a lock-free queue to store IPs for later processing). - On x86_64, the IP is stored in
uc_mcontext.gregs[REG_RIP]; on ARM, it'suc_mcontext.arm_pc.
- Only use async-signal-safe functions in the handler (e.g., avoid
- Quick example snippet:
#include <signal.h> #include <ucontext.h> #include <sys/time.h> // Atomic counter map for IPs (use a thread-safe hash table in real code) _Atomic unsigned long ip_counts[0x10000] = {0}; void sigprof_handler(int sig, siginfo_t* info, void* ucontext) { ucontext_t* ctx = (ucontext_t*)ucontext; unsigned long ip = ctx->uc_mcontext.gregs[REG_RIP]; // Increment count for this IP (simplified—use a hash table for large ranges) if (ip < sizeof(ip_counts)/sizeof(ip_counts[0])) { __atomic_add_fetch(&ip_counts[ip], 1, __ATOMIC_RELAXED); } } // Set up the timer and handler struct sigaction sa = {0}; sa.sa_sigaction = sigprof_handler; sa.sa_flags = SA_SIGINFO; sigaction(SIGPROF, &sa, NULL); struct itimerval timer = {0}; timer.it_interval.tv_usec = 1000; // Trigger every 1ms timer.it_value.tv_usec = 1000; setitimer(ITIMER_PROF, &timer, NULL);
1. User-Space Profiling with ProfileAPI
Windows provides the ProfileAPI for easy sampling-based profiling, which handles the timer and context capture for you:
- How it works: Call
StartProfile()with thePROFILE_TIMERflag to start sampling your target process. The system sends your registered callback function the IP of the executing instruction at each interval. - Concrete steps:
- Define a callback function that receives
PROFILEINFOstructs containing the sampled IP. - Call
StartProfile()to bind the sampling to your target process ID. - After sampling completes, call
StopProfile()and tally the IP frequencies.
- Define a callback function that receives
- Simplified code snippet:
#include <windows.h> #include <profileapi.h> // Thread-safe map to count IP occurrences CRITICAL_SECTION ip_count_lock; unsigned long long ip_counts[0x10000] = {0}; BOOL CALLBACK ProfileCallback(LPVOID lpProfileData, DWORD dwSize, LPVOID lpUserContext) { PROFILEINFO* info = (PROFILEINFO*)lpProfileData; if (info->dwSize == sizeof(PROFILEINFO)) { EnterCriticalSection(&ip_count_lock); // Increment count for the sampled IP if ((unsigned long)info->lpIpAddress < sizeof(ip_counts)/sizeof(ip_counts[0])) { ip_counts[(unsigned long)info->lpIpAddress]++; } LeaveCriticalSection(&ip_count_lock); } return TRUE; } // Initialize and start sampling InitializeCriticalSection(&ip_count_lock); StartProfile(PROFILE_TIMER, target_pid, ProfileCallback, NULL); // Let the target run, then stop sampling Sleep(5000); // Sample for 5 seconds StopProfile(PROFILE_TIMER, target_pid); DeleteCriticalSection(&ip_count_lock);
2. Kernel-Mode Driver (High Precision, Advanced)
For maximum accuracy (e.g., sampling kernel-mode code or avoiding user-space signal delays), you can write a Windows kernel driver:
- How it works: Use
KeSetTimerEx()to set a periodic kernel timer. In the timer callback, usePsLookupProcessByProcessIdto find the target process, switch to its active thread's context, and read the IP register (e.g.,Ripon x86_64 viaCONTEXTstruct). - Note: This requires driver signing (for modern Windows versions) and familiarity with Windows kernel programming—only pursue this if you need ultra-precise sampling.
Once you have all your IP samples:
- Map IPs to code: Resolve IP addresses to function names and source lines using debug information. On Linux, use
libdworlibbfd; on Windows, use theDbgHelplibrary'sSymFromAddr()function. - Tally and sort: Use a hash table or sorted map to count how often each IP (or function) appears. Sort these counts in descending order to identify your hot code paths.
内容的提问来源于stack exchange,提问作者Zhou

