OpenMP并行Fortran代码性能与串行版本一致问题求助
Let's break down why your parallel version isn't speeding up compared to the serial one, and fix it step by step:
Key Issues Identified
1. Array Reduction Support Limitations
You're using REDUCTION(+:exp1,exp2,exp3) on arrays, but older OpenMP standards (pre-3.0) don't support array reductions. If your compiler defaults to an older standard, this reduction clause will either be ignored (falling back to serial execution) or cause unintended behavior like data races. Even if your compiler supports OpenMP 3.0+, array reductions can sometimes lead to poor performance due to cache inefficiencies.
2. Severe False Sharing
Your current loop structure has each thread iterate over all h values while processing different k values. This means multiple threads are writing to adjacent elements of exp1, exp2, and exp3 at the same time. These elements likely share the same CPU cache line, causing frequent cache invalidations (false sharing). This overhead completely cancels out any gains from parallelism.
3. Potential Insufficient Workload
If numRows is small, the overhead of creating/managing OpenMP threads will outweigh the time saved by parallel execution, resulting in runtime similar to the serial version.
Fixes to Try
Option 1: Manual Array Reduction (Compatible with All OpenMP Versions)
Instead of relying on the array reduction clause, create private arrays for each thread to accumulate values, then merge them into the global arrays after the parallel region:
exp1 = 0.0 exp2 = 0.0 exp3 = 0.0 !$OMP PARALLEL PRIVATE(h, xConfig, exp1_priv, exp2_priv, exp3_priv) SHARED(X, numRows, nPhi, s1, s2, s3, c1, c2, c3) ! Initialize private arrays to zero exp1_priv = 0.0 exp2_priv = 0.0 exp3_priv = 0.0 !$OMP DO do k = 0, numRows-1 xConfig(0:2) = X(k,0:2) do h = 0, nPhi-1 exp1_priv(h) = exp1_priv(h) + exp(-((xConfig(0)-c1(h))**2)*s1) exp2_priv(h) = exp2_priv(h) + exp(-((xConfig(1)-c2(h))**2)*s2) exp3_priv(h) = exp3_priv(h) + exp(-((xConfig(2)-c3(h))**2)*s3) end do end do !$OMP END DO ! Merge private arrays into global ones (critical section to avoid race conditions) !$OMP CRITICAL exp1 = exp1 + exp1_priv exp2 = exp2 + exp2_priv exp3 = exp3 + exp3_priv !$OMP END CRITICAL !$OMP END PARALLEL ALine = exp1 + exp2 + exp3
Option 2: Reorder Loops to Eliminate False Sharing
Swap the k and h loops so each thread handles a single h value, accumulating across all k values. This way, each thread writes to a single element of the global arrays, eliminating false sharing entirely:
exp1 = 0.0 exp2 = 0.0 exp3 = 0.0 !$OMP PARALLEL DO SHARED(X, numRows, s1, s2, s3, c1, c2, c3) PRIVATE(k, xConfig) REDUCTION(+:exp1, exp2, exp3) do h = 0, nPhi-1 do k = 0, numRows-1 xConfig(0:2) = X(k,0:2) exp1(h) = exp1(h) + exp(-((xConfig(0)-c1(h))**2)*s1) exp2(h) = exp2(h) + exp(-((xConfig(1)-c2(h))**2)*s2) exp3(h) = exp3(h) + exp(-((xConfig(2)-c3(h))**2)*s3) end do end do !$OMP END PARALLEL DO ALine = exp1 + exp2 + exp3
Additional Checks
- Verify OpenMP Version: Ensure your compiler supports OpenMP 3.0+. For GCC, you can add
-fopenmp-version=3.0to your--f90flagsto enforce this. - Workload Size: Make sure
numRowsis large enough (e.g., 10,000+ iterations) so thread overhead is negligible compared to computation time. - Compiler Optimization: Your existing
--opt='-O3'is good, but you can also try-march=nativeto enable CPU-specific optimizations.
内容的提问来源于stack exchange,提问作者user1443613

