基于OpenACC并行化Fox矩阵乘法:数据依赖问题优化问询
Fox算法矩阵乘法并行化问题(OpenACC)
我正在尝试用Fox算法实现矩阵乘法的并行化,该算法将整个过程分为三个步骤。以下是我的实现代码:
void matmul_fox(int *matrixA, int *matrixB, int *zeroMatrix, int n){ #pragma acc kernels copyin(matrixA[0:n*n], matrixB[0:n*n]) copy(zeroMatrix[0:n*n]) { // Operation 1 #pragma acc loop independent for (int i = 0; i < n; i++){ // Extract the diagonal from arr1 #pragma acc loop independent for (int j = 0; j < n; j++){ // Add arr3 and multiplication of diag(broadcasted)*arr2 #pragma acc loop independent reduction (+: zeroMatrix[0:n*n]) for (int k = 0; k < n; k++){ zeroMatrix[j * n + k] += matrixA[j * n + j] * matrixB[j * n + k]; } } // Operation 2 // Shift-up on the rows of arr2 #pragma acc loop independent for (int a = 0; a < n; a++) { int temp1 = matrixB[0 * n + a]; #pragma acc loop independent for (int j = 0; j < n - 1; j++) { matrixB[j * n + a] = matrixB[(j + 1) * n + a]; } matrixB[(n - 1) * n + a] = temp1; } // Operation 3 // Shift-left on the rows of arr1 #pragma acc loop independent for (int b = 0; b < n; b++){ int temp2 = matrixA[b * n + 0]; // Store the first element of the row #pragma acc loop independent for (int j = 0; j < n - 1; j++){ matrixA[b * n + j] = matrixA[b * n + (j+1)]; // Shift elements to the left } matrixA[b * n + (n - 1)] = temp2; // Place the first element at the end } } } }
目前遇到的问题:Operation 2与Operation 3存在数据依赖,无法直接将三个操作全部并行化。我了解到可以通过OpenACC的async、wait和update等子句解决该问题,理论上可以让Operation 2和Operation 3等待Operation 1完成后再并行执行(二者之间无依赖)。
当前的运行状况:
- 未使用
wait子句时,因数据依赖导致计算结果错误; - 不使用OpenACC时代码可正常运行,但执行时间比传统串行算法更长。
希望能找到正确实现Fox算法并行化的方案,以获得更优的性能。
内容的提问来源于stack exchange,提问作者pranavs1997
相关产品推荐
相关产品推荐

