OpenCL:如何无需多线程在不同设备上分配计算任务
Absolutely, you can pull off this multi-device task split without needing OpenMP or complex multi-threading—a single CPU core can handle coordinating your CPU and two GPUs perfectly. This approach avoids resource waste, keeps your code focused on OpenCL primitives, and skips any multi-threading compile conflicts. Let’s walk through how this works, step by step.
First: Feasibility Check
OpenCL is designed explicitly to manage multiple heterogeneous devices (CPUs, GPUs, etc.) from a single host thread. You don’t need parallelism on the host side to distribute work—you just need to:
- Enumerate all your target devices (CPU + GPU1 + GPU2)
- Create separate command queues for each device
- Split your array into chunks tailored to each device’s performance
- Dispatch work to each queue, wait for completion, and assemble results
Your core idea of splitting n = n1 + n2 + n3 such that f1(n1) = f2(n2) = f3(n3) is spot-on—this minimizes total runtime (since the slowest device dictates completion time).
Step-by-Step Implementation
1. Precompute Device Performance Functions
First, you need a way to get f1(n), f2(n), f3(n)—the time each device takes to process an array of length n. You can:
- Run a quick benchmark on startup for each device across a few array sizes (e.g., 2^10, 2^12, 2^14)
- Fit a simple linear model to these times (array addition is memory-bound, so runtime scales almost linearly with
nonce you’re past small array overhead)
2. Enumerate Devices & Setup OpenCL Objects
From a single CPU thread, you’ll:
- Get the platform(s) you’re using
- Fetch your target devices (filter for CPU + the two GPUs)
- Create a context that includes all three devices
- Create a separate command queue for each device (enable profiling with
CL_QUEUE_PROFILING_ENABLEif you want to validate runtime later) - Compile your array addition kernel once (it can be reused across all devices)
3. Calculate Task Splits
Solve for n1, n2, n3 where:
n1 + n2 + n3 = nf1(n1) = f2(n2) = f3(n3) = T(target runtime)
Since your functions are likely linear, this simplifies to solving a system of linear equations. For example, if f1(n) = a1*n + b1, f2(n)=a2*n +b2, f3(n)=a3*n +b3, you can solve for T and compute each ni. Don’t forget to round to integers and adjust the last chunk to ensure the total adds up to n.
4. Dispatch Work & Collect Results
Instead of splitting and copying arrays manually, you can map sections of your host arrays directly to device buffers. Here’s the key flow:
- Allocate host input arrays
A,Band output arrayC - For each device, create buffers pointing to the corresponding slice of
A,B,C(useCL_MEM_USE_HOST_PTRto avoid extra copies if your host memory is pinned) - For each device queue:
- Enqueue a write command to copy the device’s slice of
AandBto its local buffers (skip if usingCL_MEM_USE_HOST_PTR) - Enqueue the kernel execution with the device’s chunk size
ni - Enqueue a read command to copy the result back to the corresponding slice of
C
- Enqueue a write command to copy the device’s slice of
- Wait for all command queues to finish (use
clFinish()on each queue, or event objects to track completion)
Concrete Code Example
Here’s a trimmed-down C-OpenCL example implementing this flow:
#include <stdio.h> #include <stdlib.h> #include <CL/cl.h> // Simple array addition kernel const char* kernel_source = "__kernel void add_arrays(__global const float* A, __global const float* B, __global float* C, unsigned int n) { unsigned int i = get_global_id(0); if (i < n) { C[i] = A[i] + B[i]; } }"; int main() { cl_platform_id platform; cl_device_id devices[3]; // CPU + GPU1 + GPU2 cl_context context; cl_command_queue queues[3]; cl_program program; cl_kernel kernel; cl_int err; // 1. Enumerate platform and devices (adjust to pick your target devices) clGetPlatformIDs(1, &platform, NULL); clGetDeviceIDs(platform, CL_DEVICE_TYPE_CPU, 1, &devices[0], NULL); clGetDeviceIDs(platform, CL_DEVICE_TYPE_GPU, 1, &devices[1], NULL); clGetDeviceIDs(platform, CL_DEVICE_TYPE_GPU, 1, &devices[2], NULL); // 2. Create context and command queues context = clCreateContext(NULL, 3, devices, NULL, NULL, &err); queues[0] = clCreateCommandQueue(context, devices[0], 0, &err); queues[1] = clCreateCommandQueue(context, devices[1], 0, &err); queues[2] = clCreateCommandQueue(context, devices[2], 0, &err); // 3. Compile kernel program = clCreateProgramWithSource(context, 1, &kernel_source, NULL, &err); clBuildProgram(program, 3, devices, NULL, NULL, NULL); kernel = clCreateKernel(program, "add_arrays", &err); // 4. Define total array size and compute splits (replace with your f(n) calculation) unsigned int n = 1 << 16; // 65536 elements unsigned int n1 = 20000, n2 = 25000, n3 = n - n1 - n2; // 5. Allocate host arrays float* A = (float*)malloc(n * sizeof(float)); float* B = (float*)malloc(n * sizeof(float)); float* C = (float*)malloc(n * sizeof(float)); // Initialize A and B with test data... // 6. Create device buffers for each chunk cl_mem bufA[3], bufB[3], bufC[3]; bufA[0] = clCreateBuffer(context, CL_MEM_USE_HOST_PTR | CL_MEM_READ_ONLY, n1*sizeof(float), A, &err); bufB[0] = clCreateBuffer(context, CL_MEM_USE_HOST_PTR | CL_MEM_READ_ONLY, n1*sizeof(float), B, &err); bufC[0] = clCreateBuffer(context, CL_MEM_USE_HOST_PTR | CL_MEM_WRITE_ONLY, n1*sizeof(float), C, &err); bufA[1] = clCreateBuffer(context, CL_MEM_USE_HOST_PTR | CL_MEM_READ_ONLY, n2*sizeof(float), A + n1, &err); bufB[1] = clCreateBuffer(context, CL_MEM_USE_HOST_PTR | CL_MEM_READ_ONLY, n2*sizeof(float), B + n1, &err); bufC[1] = clCreateBuffer(context, CL_MEM_USE_HOST_PTR | CL_MEM_WRITE_ONLY, n2*sizeof(float), C + n1, &err); bufA[2] = clCreateBuffer(context, CL_MEM_USE_HOST_PTR | CL_MEM_READ_ONLY, n3*sizeof(float), A + n1 + n2, &err); bufB[2] = clCreateBuffer(context, CL_MEM_USE_HOST_PTR | CL_MEM_READ_ONLY, n3*sizeof(float), B + n1 + n2, &err); bufC[2] = clCreateBuffer(context, CL_MEM_USE_HOST_PTR | CL_MEM_WRITE_ONLY, n3*sizeof(float), C + n1 + n2, &err); // 7. Set kernel arguments and enqueue commands for each device for (int i = 0; i < 3; i++) { unsigned int ni = (i == 0) ? n1 : (i == 1) ? n2 : n3; clSetKernelArg(kernel, 0, sizeof(cl_mem), &bufA[i]); clSetKernelArg(kernel, 1, sizeof(cl_mem), &bufB[i]); clSetKernelArg(kernel, 2, sizeof(cl_mem), &bufC[i]); clSetKernelArg(kernel, 3, sizeof(unsigned int), &ni); // Enqueue kernel execution (tune local work size for better performance) size_t global_size = ni; clEnqueueNDRangeKernel(queues[i], kernel, 1, NULL, &global_size, NULL, 0, NULL, NULL); } // 8. Wait for all queues to finish and results to be written back for (int i = 0; i < 3; i++) { clFinish(queues[i]); } // 9. Cleanup resources clReleaseKernel(kernel); clReleaseProgram(program); for (int i = 0; i < 3; i++) { clReleaseMemObject(bufA[i]); clReleaseMemObject(bufB[i]); clReleaseMemObject(bufC[i]); clReleaseCommandQueue(queues[i]); } clReleaseContext(context); free(A); free(B); free(C); return 0; }
Key Notes
- Performance Tuning: Use pinned host memory (
CL_MEM_ALLOC_HOST_PTR+clEnqueueMapBuffer) to speed up data transfers between host and devices. - Task Split Adjustment: If your performance functions aren’t perfectly linear, iterate on split sizes slightly to balance runtimes. Use OpenCL profiling events to measure actual device runtime and adjust splits for future tasks.
- Typo Note: I noticed you mentioned a typo in your original graph—GPI1 should be GPU1, just wanted to confirm I understood that correctly!
内容的提问来源于stack exchange,提问作者Foad S. Farimani

