如何实现Intel MKL ddiagemv的多线程运行?
ddiagemv First off, great job outperforming MKL's serial ddiagemv—that's no small feat! Let's dig into why you can't get MKL's routine to run in parallel, even after tweaking environment variables. Here are the most likely fixes to try:
1. Verify if ddiagemv actually supports multi-threading
Not all MKL BLAS/LAPACK routines are parallelized. From Intel's documentation, ddiagemv (part of the sparse BLAS subset) is not inherently multi-threaded in many MKL versions. Its diagonal storage format's access pattern often doesn't lend itself to efficient parallelization without significant overhead, so it’s frequently implemented as a single-threaded routine.
Before wasting more time on thread settings, confirm this by checking your specific MKL version’s docs—if it’s listed as serial-only, that’s your root cause.
2. Double-check your environment variable setup
Environment variables can be finicky—make sure you’re setting them before launching your executable, and that they aren’t being overridden elsewhere:
- Set variables in the same shell session before running your program:
export MKL_NUM_THREADS=4 export OMP_NUM_THREADS=4 export MKL_DOMAIN_NUM_THREADS="MKL_BLAS=4" ./your_executable - Avoid conflicting settings (like setting
OMP_NUM_THREADS=1while trying to enable MKL threads—stick to consistent values for parallel domains). - Verify variables are being picked up by adding a quick check in your Fortran code:
use mkl_service integer :: num_threads num_threads = mkl_get_max_threads() print *, "MKL max threads detected: ", num_threads
3. Explicitly set MKL threads via API calls
Sometimes environment variables don’t stick (especially if your program launches via a script or IDE that overrides them). Try enforcing thread counts directly in your code:
use mkl_service call mkl_set_num_threads(4) ! Or target the BLAS domain specifically for granular control call mkl_domain_set_num_threads(4, MKL_DOMAIN_BLAS)
This bypasses external environment overrides and ensures your desired thread count is active.
4. Check your compilation and linking flags
Make sure you’re compiling with OpenMP support and linking against MKL’s parallel libraries:
- For Intel compilers:
ifort -qopenmp -mkl=parallel your_code.f90 -o your_executable - For GCC:
gfortran -fopenmp -lmkl_rt -lmkl_intel_lp64 -lmkl_gnu_thread -lmkl_core your_code.f90 -o your_executable
Using serial MKL libraries (-mkl=sequential) will disable all parallelism, so double-check your linker flags.
5. Ensure your problem size is large enough to trigger parallelism
MKL’s runtime often skips parallelization for small problem sizes because thread-spawning overhead outweighs performance gains. Try scaling up your matrix dimensions (more rows/columns, more diagonals) and see if multi-threading kicks in.
内容的提问来源于stack exchange,提问作者Rey Reddington

