MPI_Allreduce自定义操作中len参数异常及段错误问题排查
len in MPI custom reduction function not equal the total array size? Let's break down why you're seeing unexpected values for *len in your custom stitchResultArrays function, and how to fix the segmentation fault.
The Core Issue: MPI Chunks Data for Reduction Operations
MPI doesn't pass the entire array to your custom reduction function in one go. Instead, it splits the large buffer into smaller, contiguous chunks (the size depends on your MPI implementation, number of processes, and data size) and calls your function once per chunk. That's why you're seeing *len values like 1251 and 1252 instead of 2503—those are the sizes of the individual chunks being processed in each invocation of your function.
When MPI calls stitchResultArrays, the in and inout pointers point to the start of the current chunk, not the start of the full array. So accessing in[2502] or inout[2502] is trying to read/write way beyond the bounds of the current chunk, which causes the segmentation fault.
Why Your Assumption About len Was Wrong
The len parameter in the custom MPI reduction function represents the number of elements in the current chunk, not the total size of the original array. MPI does this for performance reasons—processing smaller chunks can be more efficient, especially with large datasets or distributed memory systems.
Fixing Your Code
Since your goal is only to sum the value at index 2502 across all processes, you don't need a custom reduction function for the entire array. A simpler approach is to split your data into two parts:
- The first 2502 elements (which you don't need to modify)
- The single element at index 2502 (which needs to be summed)
Here's how you can adjust your code:
#include <mpi.h> #include <math.h> #include <stdio.h> #include <stddef.h> #include <stdlib.h> int main(int argc, char** argv) { MPI_Init(&argc, &argv); int world_size; MPI_Comm_size(MPI_COMM_WORLD, &world_size); int world_rank; MPI_Comm_rank(MPI_COMM_WORLD, &world_rank); const int total_size = 2503; double sendMpiMessage[total_size]; double test[total_size]; // Initialize your data sendMpiMessage[2502] = 4.1; sendMpiMessage[2501] = 4.4; sendMpiMessage[2500] = 4.2; sendMpiMessage[0] = 4.5; sendMpiMessage[1] = (double)(world_rank + 5); // Copy the unchanged elements directly (no reduction needed) for (int i = 0; i < total_size - 1; i++) { test[i] = sendMpiMessage[i]; } // Only reduce the last element using MPI_SUM MPI_Allreduce(&sendMpiMessage[2502], &test[2502], 1, MPI_DOUBLE, MPI_SUM, MPI_COMM_WORLD); printf("result = %f\n", test[2502]); MPI_Finalize(); return 0; }
If you really need to use a custom reduction function (for more complex logic), you must adjust it to only operate on the len elements provided in each chunk. You can't access indices outside the current chunk. For example, if your logic needed to modify elements based on their position in the full array, you'd need to track global indices—which isn't straightforward with MPI's reduction model. In most cases, splitting the data into relevant parts (like we did above) is cleaner and safer.
Additional Notes
- The
commuteparameter (set to 1 inMPI_Op_create) indicates your operation is commutative (order of operations doesn't matter), which is correct for summing values, but it doesn't affect how MPI chunks the data. - Always avoid hardcoding indices like 2502 in your custom reduction functions—use the
lenparameter to stay within the bounds of the current chunk.
内容的提问来源于stack exchange,提问作者Danny Sortino

