OpenCL技术问询:能否利用DMA写入传入程序的主内存地址?
Short answer: Yes, this is technically feasible, though it comes with significant caveats tied to your operating system, hardware, and memory permissions. Since you’re expecting (and intending) crashes as part of your goal to overwrite the CPU program’s address space from the GPU process, let’s break down how this can work:
Key Background on OpenCL Host Memory Access
OpenCL is designed to support shared memory between the host (CPU) and compute devices (GPU/accelerator) via DMA, which is exactly what you’re looking to leverage. The critical flag here is CL_MEM_USE_HOST_PTR when creating a buffer: this tells the OpenCL runtime that you want the device to directly access the provided host memory pointer, rather than copying data to/from device-local memory.
Step-by-Step Approach
Here’s a rough outline of how to implement this:
- Get the target host memory address
- In your CPU program, retrieve the pointer to the memory region you want to overwrite (e.g., the address of a global variable, or the base address of your program’s code segment). If ASLR is enabled, you’ll need to resolve the actual runtime address of your target (e.g., using
dladdron Linux orGetModuleHandleon Windows to find your program’s base, then adding the static offset of your target).
- In your CPU program, retrieve the pointer to the memory region you want to overwrite (e.g., the address of a global variable, or the base address of your program’s code segment). If ASLR is enabled, you’ll need to resolve the actual runtime address of your target (e.g., using
- Adjust memory permissions (if needed)
- Modern OSes mark certain memory regions (like executable code segments) as read-only or non-writable. To let the GPU write to these areas, you’ll need to modify the permissions from your CPU program first:
- On Linux: Use
mprotect()to set the target page toPROT_READ | PROT_WRITE | PROT_EXEC(or at leastPROT_WRITE). - On Windows: Use
VirtualProtect()to set the region toPAGE_READWRITE.
- On Linux: Use
- Modern OSes mark certain memory regions (like executable code segments) as read-only or non-writable. To let the GPU write to these areas, you’ll need to modify the permissions from your CPU program first:
- Create an OpenCL buffer with direct host access
- Use
clCreateBuffer()with theCL_MEM_USE_HOST_PTRflag, passing your target host pointer as thehost_ptrargument. This tells the OpenCL runtime to map this host memory directly for device access via DMA.
- Use
- Write to the memory from an OpenCL kernel
- In your OpenCL kernel, declare a pointer to the buffer (using
__globalmemory space) and perform writes to it. For example:__kernel void overwrite_host_memory(__global char* target_addr) { // Write arbitrary data to the target host address target_addr[0] = 0x90; // NOP instruction, or any byte you want // Extend this to overwrite more bytes as needed }
- In your OpenCL kernel, declare a pointer to the buffer (using
- Execute the kernel
- Enqueue the kernel for execution on your GPU. Once the kernel runs, the GPU will use DMA to write directly to the host memory address you provided, overwriting the CPU program’s memory and triggering a crash (e.g., if you overwrite executable code with invalid instructions).
Critical Caveats
- Hardware Support: Not all OpenCL devices support direct host memory access. Most modern GPUs do, but you’ll want to verify your device supports
CL_MEM_USE_HOST_PTRviaclGetDeviceInfo(). - Memory Alignment: OpenCL devices often require host memory pointers to be aligned to a specific boundary (e.g., 64 bytes). Misaligned pointers may cause errors or undefined behavior, though some devices are more forgiving. You can use
posix_memalign()(Linux) or_aligned_malloc()(Windows) if you need to allocate aligned memory, but if you’re using existing program memory, you may need to adjust your target to an aligned address. - OS Security Restrictions: Features like Windows’ UMIP (User-Mode Instruction Prevention) or Linux’s
noexecmount flags may block writes to executable memory even after adjusting permissions. You may need to disable these (if possible) for your use case. - Undefined Behavior: Even if you get the write to work, overwriting a running program’s memory will lead to undefined behavior—exactly what you want, but be aware that the crash may not happen exactly as you expect (e.g., the CPU may have cached instructions, so the overwrite might not take effect immediately until the cache is invalidated).
内容的提问来源于stack exchange,提问作者novafacing

