CUDA矩阵元素模p约简核函数异常:二维Stride循环线程索引排查
矩阵模p约简CUDA核函数的循环步长与网格配置错误修复
核心问题分析
你的代码存在两处致命错误,直接引发了无限循环:
- 循环步长逻辑完全错误:原代码中行/列的循环步长用
blockIdx.x*gridDim.x和blockIdx.y*gridDim.y,当blockIdx.x为0时步长会变成0,线程直接陷入死循环;正确步长应该是对应维度的总线程数(网格块数×每块线程数)。 - 网格维度计算逻辑错误:原代码用矩阵总元素数除以单块线程数来计算网格维度,完全不符合二维网格对应矩阵行列的逻辑,会导致网格配置混乱。
修正后的代码
核函数代码
__global__ void gpu_matrix_fma_reduction(double *matrix, int rows, int cols, double u, double p){ /* Reduce each coefficient of a matrix modulo p (with u = 1.0/p) */ for (int i = blockIdx.x * blockDim.x + threadIdx.x; i < rows; i += gridDim.x * blockDim.x) { for (int j = blockIdx.y * blockDim.y + threadIdx.y; j < cols; j += gridDim.y * blockDim.y) { gpu_fma_reduction(matrix + i*cols + j, u, p); } } }
核函数调用代码
dim3 threadsPerBlock(16, 16); // 对行列数分别向上取整,计算所需的网格块数 dim3 numBlocks( (n + threadsPerBlock.x - 1) / threadsPerBlock.x, (m + threadsPerBlock.y - 1) / threadsPerBlock.y); gpu_matrix_fma_reduction<<<numBlocks, threadsPerBlock>>>(partial_matrix, n, m, u, p);
额外建议
可以在核函数调用后添加CUDA错误检查,快速定位潜在问题:
gpu_matrix_fma_reduction<<<numBlocks, threadsPerBlock>>>(partial_matrix, n, m, u, p); cudaDeviceSynchronize(); cudaError_t err = cudaGetLastError(); if (err != cudaSuccess) { printf("CUDA error: %s\n", cudaGetErrorString(err)); }
内容的提问来源于stack exchange,提问作者Dimitri Lesnoff
相关产品推荐
相关产品推荐

