You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

OpenMP并行Fortran代码性能与串行版本一致问题求助

Troubleshooting OpenMP Parallelization in Your Fortran Code

Let's break down why your parallel version isn't speeding up compared to the serial one, and fix it step by step:

Key Issues Identified

1. Array Reduction Support Limitations

You're using REDUCTION(+:exp1,exp2,exp3) on arrays, but older OpenMP standards (pre-3.0) don't support array reductions. If your compiler defaults to an older standard, this reduction clause will either be ignored (falling back to serial execution) or cause unintended behavior like data races. Even if your compiler supports OpenMP 3.0+, array reductions can sometimes lead to poor performance due to cache inefficiencies.

2. Severe False Sharing

Your current loop structure has each thread iterate over all h values while processing different k values. This means multiple threads are writing to adjacent elements of exp1, exp2, and exp3 at the same time. These elements likely share the same CPU cache line, causing frequent cache invalidations (false sharing). This overhead completely cancels out any gains from parallelism.

3. Potential Insufficient Workload

If numRows is small, the overhead of creating/managing OpenMP threads will outweigh the time saved by parallel execution, resulting in runtime similar to the serial version.

Fixes to Try

Option 1: Manual Array Reduction (Compatible with All OpenMP Versions)

Instead of relying on the array reduction clause, create private arrays for each thread to accumulate values, then merge them into the global arrays after the parallel region:

exp1 = 0.0
exp2 = 0.0
exp3 = 0.0

!$OMP PARALLEL PRIVATE(h, xConfig, exp1_priv, exp2_priv, exp3_priv) SHARED(X, numRows, nPhi, s1, s2, s3, c1, c2, c3)
    ! Initialize private arrays to zero
    exp1_priv = 0.0
    exp2_priv = 0.0
    exp3_priv = 0.0

    !$OMP DO
    do k = 0, numRows-1
        xConfig(0:2) = X(k,0:2)
        do h = 0, nPhi-1
            exp1_priv(h) = exp1_priv(h) + exp(-((xConfig(0)-c1(h))**2)*s1)
            exp2_priv(h) = exp2_priv(h) + exp(-((xConfig(1)-c2(h))**2)*s2)
            exp3_priv(h) = exp3_priv(h) + exp(-((xConfig(2)-c3(h))**2)*s3)
        end do
    end do
    !$OMP END DO

    ! Merge private arrays into global ones (critical section to avoid race conditions)
    !$OMP CRITICAL
        exp1 = exp1 + exp1_priv
        exp2 = exp2 + exp2_priv
        exp3 = exp3 + exp3_priv
    !$OMP END CRITICAL
!$OMP END PARALLEL

ALine = exp1 + exp2 + exp3

Option 2: Reorder Loops to Eliminate False Sharing

Swap the k and h loops so each thread handles a single h value, accumulating across all k values. This way, each thread writes to a single element of the global arrays, eliminating false sharing entirely:

exp1 = 0.0
exp2 = 0.0
exp3 = 0.0

!$OMP PARALLEL DO SHARED(X, numRows, s1, s2, s3, c1, c2, c3) PRIVATE(k, xConfig) REDUCTION(+:exp1, exp2, exp3)
do h = 0, nPhi-1
    do k = 0, numRows-1
        xConfig(0:2) = X(k,0:2)
        exp1(h) = exp1(h) + exp(-((xConfig(0)-c1(h))**2)*s1)
        exp2(h) = exp2(h) + exp(-((xConfig(1)-c2(h))**2)*s2)
        exp3(h) = exp3(h) + exp(-((xConfig(2)-c3(h))**2)*s3)
    end do
end do
!$OMP END PARALLEL DO

ALine = exp1 + exp2 + exp3

Additional Checks

  • Verify OpenMP Version: Ensure your compiler supports OpenMP 3.0+. For GCC, you can add -fopenmp-version=3.0 to your --f90flags to enforce this.
  • Workload Size: Make sure numRows is large enough (e.g., 10,000+ iterations) so thread overhead is negligible compared to computation time.
  • Compiler Optimization: Your existing --opt='-O3' is good, but you can also try -march=native to enable CPU-specific optimizations.

内容的提问来源于stack exchange,提问作者user1443613

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 06:38:10