关于将类变量XY、V移出OpenMP并行循环的技术咨询(scrypt并行版)
XY and V Out of OpenMP Parallel Loop for Scrypt Parallelization Hey there! Let's work through how to move your XY and V SecByteBlock variables outside the OpenMP parallel loop while keeping your scrypt implementation correct.
First, let's clarify the core issue: if you just move these variables directly outside the loop without handling thread isolation, multiple threads will try to access and modify the same memory blocks, leading to data races and corrupted results. The key here is to give each thread its own private copy of XY and V—just like when you declared them inside the loop, but now we'll manage that ownership outside the loop body.
Approach 1: Use OpenMP private Clause
This is the most straightforward way to replicate the behavior of declaring variables inside the loop, but with declarations moved outside:
// Declare variables outside the parallel loop (default-constructed) SecByteBlock XY; SecByteBlock V; // Mark XY and V as private to each thread #pragma omp parallel for private(XY, V) for (unsigned int i = 0; i < parallel; ++i) { // Initialize each thread's private copy with the required size XY.resize(static_cast<size_t>(blockSize * 256U)); V.resize(static_cast<size_t>(blockSize * cost * 128U)); // 3: B_i <-- MF(B_i, N) const ptrdiff_t offset = ...; // Your existing logic here // Rest of your scrypt iteration code... }
Why this works:
- The
private(XY, V)clause tells OpenMP to create a separate instance of each variable for every thread in the parallel region. - Each thread initializes its own
XYandVviaresize(), just like you did when declaring them inside the loop. No two threads will touch the same memory, so data races are eliminated. - This maintains the exact correctness of your working scenarios (no OpenMP, or variables inside loop) while moving the declarations outside the loop.
Approach 2: Preallocate Thread-Private Storage (For Reuse)
If you're running multiple parallel loops or want to avoid reallocating memory on each iteration (though scrypt's per-iteration work likely dominates), you can preallocate storage for each thread once:
// Get the maximum number of threads OpenMP will use int maxThreads; #pragma omp parallel { maxThreads = omp_get_num_threads(); } // Create a vector of SecByteBlock instances, one per thread std::vector<SecByteBlock> XY_per_thread(maxThreads); std::vector<SecByteBlock> V_per_thread(maxThreads); // Preallocate each thread's storage once for (int t = 0; t < maxThreads; ++t) { XY_per_thread[t].resize(static_cast<size_t>(blockSize * 256U)); V_per_thread[t].resize(static_cast<size_t>(blockSize * cost * 128U)); } // In the parallel loop, each thread uses its preallocated block #pragma omp parallel for for (unsigned int i = 0; i < parallel; ++i) { int threadId = omp_get_thread_num(); SecByteBlock& XY = XY_per_thread[threadId]; SecByteBlock& V = V_per_thread[threadId]; // 3: B_i <-- MF(B_i, N) const ptrdiff_t offset = ...; // Your existing logic here // Rest of your scrypt iteration code... }
Why this works:
- We preallocate a
SecByteBlockfor every possible thread once, before the parallel loop starts. - Each thread uses its own preallocated block via
omp_get_thread_num(), avoiding repeated memory allocation/deallocation in the loop. - This is a good option if you want to optimize memory overhead, especially for large
costvalues whereVcan be very large.
Key Notes to Verify
- Make sure your
SecByteBlockimplementation is thread-safe for individual instance access (which it should be, since each thread only touches its own copy). - Test both approaches against your working baseline scenarios to confirm the output is identical—this ensures you haven't introduced any bugs.
内容的提问来源于stack exchange,提问作者jww

