使用OpenMP加速FFTW计算过慢及编译运行问题咨询
Hi Tom, let's tackle both your segmentation fault issue and the slow FFTW performance when using OpenMP. Here's a breakdown of the likely causes and actionable fixes:
1. Why You're Getting a Segmentation Fault with OpenMP
The stack overflow (and resulting segfault) is almost certainly tied to large local arrays combined with OpenMP's thread stack allocation. Here's the breakdown:
- Each OpenMP thread gets its own dedicated stack space. If you declare large arrays (like your 3D arrays sized
Nx=381,Ny=129,Nz=129) as local variables (stored on the stack), every thread will try to allocate that array on its own stack—quickly exceeding the default system stack limit. - Using
ulimit -s unlimitedworks around this temporarily, but it's not a robust solution (it can lead to poor memory locality or even system-wide memory pressure if threads consume too much stack space).
Fixes:
- Move large arrays from the stack to the heap using Fortran's dynamic allocation:
Real, Allocatable :: my_3d_array(:,:,:) Allocate(my_3d_array(Nx, Ny, Nz)) ! Use the array for computations Deallocate(my_3d_array) - Alternatively, declare large arrays in a module (module variables are stored in static memory, not the stack).
- Check for accidental large temporary arrays in loops or FFTW calls—these can also eat up stack space across threads.
2. Why FFTW is Slow with OpenMP
There are several common culprits for slow parallel FFTW performance. Let's go through the most impactful ones:
a. Incorrect OpenMP Compiler Flag
You're using gfortran with the -qopenmp flag—but -qopenmp is an Intel compiler (ifort) option. For gfortran, the correct flag to enable OpenMP is -fopenmp. Using the wrong flag means OpenMP isn't properly activated, so FFTW's parallel code won't run (or will run incorrectly).
Fix: Update your compile command to use the right flag:
gfortran -O -mcmodel=medium test.f90 -lfftw3_omp -lfftw3 -lm -fopenmp
b. Missing FFTW Thread Initialization
FFTW's OpenMP support requires explicit initialization to leverage multiple threads. If you skip this step, FFTW will default to single-threaded mode regardless of your OpenMP settings.
Fix: Add these lines early in your program (before creating any FFTW plans):
Call fftw_init_threads() Call fftw_plan_with_nthreads(OMP_GET_NUM_THREADS()) ! Uses your set OpenMP thread count
Make sure to set the number of OpenMP threads first—either via the environment variable export OMP_NUM_THREADS=4 (adjust to your core count) before running your program, or use call omp_set_num_threads(4) directly in your code.
c. Suboptimal FFTW Plan Selection
FFTW's performance depends heavily on the plan you create. If you're using FFTW_ESTIMATE (the default if you don't specify), it creates a quick plan but not necessarily the fastest for your data size.
Fix: For repeated FFT calls (standard in simulations), use FFTW_MEASURE or FFTW_PATIENT when creating your plan. These spend more time upfront finding the optimal algorithm but deliver far better performance for repeated runs:
! Example: Creating a 3D FFT plan with FFTW_MEASURE Call fftw_plan_dft_3d(plan, Nx, Ny, Nz, in_array, out_array, FFTW_FORWARD, FFTW_MEASURE)
Important: Create FFTW plans outside of parallel regions—plan creation is not thread-safe and should be done in a single-threaded context.
d. Problem Size is Too Small for Parallelism
Parallel computing has inherent overhead (thread creation, synchronization, data splitting). If your FFT size (381x129x129) is too small, the overhead of splitting work across threads can outweigh the speedup from parallelism.
Check: Run a quick test comparing single-threaded vs. multi-threaded FFTW. If single-threaded is faster, your problem size might be too small to benefit from parallelism. Try:
- Increasing the problem size if your simulation allows it.
- Using fewer threads (e.g., 2 instead of 8) to reduce overhead.
e. Non-Contiguous Memory Access
FFTW performs best with contiguous arrays. If your input/output arrays are non-contiguous (e.g., slices of larger arrays, or allocated with non-default strides), FFTW has to copy data, which slows things down.
Fix: Ensure your arrays are contiguous. In Fortran, arrays allocated with allocate(my_array(Nx, Ny, Nz)) are contiguous by default. Avoid using array sections like my_array(:,1:Ny-1,:) as FFT inputs/outputs—instead, copy the data to a contiguous temporary array first.
Final Notes
- After fixing the compile flag and FFTW thread initialization, recompile and test performance again.
- Use tools like
htopto monitor thread usage and confirm multiple threads are active during FFT computations.
内容的提问来源于stack exchange,提问作者Tom

