如何用ptrace统计x86_64程序运行时的CALL与RET指令数量?
Alright, let's walk through how to reliably count CALL and RET instructions using ptrace on x86_64 (Intel syntax) without enabling PTRACE_SYSCALL. I'll break this down into actionable steps, including a code example to tie it all together.
Since you're not leveraging PTRACE_SYSCALL, the core approach relies on single-step execution with ptrace. We'll attach to the target process, pause it after every single instruction, check what instruction is about to run, increment counters for CALL/RET, and repeat until the process exits. This avoids the complexity of manually parsing variable-length x86 instructions to jump between execution points.
1. Attach to the Target Process
First, attach ptrace to your target process using PTRACE_ATTACH. You'll need to wait for the process to pause after attaching using waitpid to ensure synchronization.
2. Single-Step Execution Loop
Set up a loop that repeats these actions:
- Fetch the current register state to get the value of
RIP(the instruction pointer, which points to the next instruction the process will execute) - Read the instruction at the address pointed to by
RIP - Check if the instruction is a CALL or RET, and increment the corresponding counter
- Tell the process to execute that one instruction with
PTRACE_SINGLESTEP - Wait for the process to pause again, then repeat until the process exits
3. Identify CALL and RET Instructions
x86_64 has multiple variants of CALL and RET—here's how to detect the most common ones:
CALL Instructions
- Direct relative CALL: Opcode
0xE8(the variant you mentioned, where the lowest byte is0xE8) - Indirect CALL: Opcode
0xFFfollowed by a ModR/M byte where bits 3-5 are010(binary), which signals the CALL opcode extension.
If you only care about the 0xE8 variant, you can skip checking the 0xFF case.
RET Instructions
- Near RET (no stack cleanup): Opcode
0xC3 - Near RET with immediate stack cleanup: Opcode
0xC2
Both count as valid RET operations, so we'll check for either.
4. Handle Process Exit
Once waitpid indicates the process has exited (via WIFEXITED or WIFSIGNALED), exit the loop, detach ptrace, and print your final counts.
Here's a complete C implementation that follows this logic:
#include <sys/ptrace.h> #include <sys/wait.h> #include <sys/user.h> #include <stdio.h> #include <stdlib.h> #include <unistd.h> #include <stdint.h> int main(int argc, char *argv[]) { if (argc != 2) { fprintf(stderr, "Usage: %s <target-pid>\n", argv[0]); return EXIT_FAILURE; } pid_t target_pid = strtoul(argv[1], NULL, 10); struct user_regs_struct regs; int status; uint64_t call_count = 0; uint64_t ret_count = 0; // Attach to the target process if (ptrace(PTRACE_ATTACH, target_pid, NULL, NULL) == -1) { perror("Failed to attach to process"); return EXIT_FAILURE; } waitpid(target_pid, &status, 0); while (!WIFEXITED(status) && !WIFSIGNALED(status)) { // Get current register state to read RIP if (ptrace(PTRACE_GETREGS, target_pid, NULL, ®s) == -1) { perror("Failed to get registers"); break; } // Read the first byte of the next instruction long instr_data = ptrace(PTRACE_PEEKDATA, target_pid, regs.rip, NULL); uint8_t op_code = (uint8_t)instr_data; // Check for CALL instructions if (op_code == 0xE8) { call_count++; } else if (op_code == 0xFF) { // Read ModR/M byte to confirm indirect CALL long modrm_data = ptrace(PTRACE_PEEKDATA, target_pid, regs.rip + 1, NULL); uint8_t modrm = (uint8_t)modrm_data; // Check if opcode extension is CALL (bits 3-5 = 010) if (((modrm >> 3) & 0x7) == 0x2) { call_count++; } } // Check for RET instructions if (op_code == 0xC3 || op_code == 0xC2) { ret_count++; } // Single-step to execute the current instruction if (ptrace(PTRACE_SINGLESTEP, target_pid, NULL, NULL) == -1) { perror("Failed to single-step process"); break; } waitpid(target_pid, &status, 0); } // Detach from the target process ptrace(PTRACE_DETACH, target_pid, NULL, NULL); printf("Total CALL instructions executed: %lu\n", call_count); printf("Total RET instructions executed: %lu\n", ret_count); return EXIT_SUCCESS; }
- Single-step reliability: x86_64 uses variable-length instructions, so single-step execution eliminates the need to manually calculate instruction lengths to find the next instruction. This is the most straightforward way to ensure we check every instruction.
- Permissions: You'll need sufficient permissions to attach to the target process—either run the tracer as root, or ensure the target process is owned by your user.
- Performance impact: Single-step execution will slow down the target process significantly, as it pauses after every instruction. This is suitable for short-lived programs or debugging, but not ideal for long-running production processes.
- Signal handling: The example exits if the target process receives a fatal signal. For more robustness, you could forward signals to the target process using
ptrace(PTRACE_CONT, target_pid, NULL, signal)if needed.
内容的提问来源于stack exchange,提问作者Paul-Marie

