如何理解Valgrind报错及MPI相关运行错误?
Let's break down your Valgrind and MPI issues step by step—most of these aren't actually problems with your code, but rather quirks or bugs in the underlying libraries (hwloc, MPI/PSM) you're using. Here's a detailed breakdown:
1. Understanding "Still Reachable" Valgrind Entries
Valgrind’s "still reachable" label doesn’t mean you have a memory leak in your code—it refers to memory that was allocated but never freed, but remains accessible via valid program pointers when the program exits. This is often safe (and common in library code that doesn’t clean up global state on exit).
- 0-byte malloc entries (from hwloc): These come from the
hwloclibrary, which Open MPI uses to discover your system’s hardware topology. The 0-byte allocations are harmless—they’re likely placeholder allocations the library uses for internal bookkeeping. This has nothing to do with yournew/deletecalls. - 48-byte entry linked to MPI_File_open: This traces back to Open MPI’s attribute system (
ompi_attr_set_c). The library stores internal metadata for MPI file handles, and doesn’t clean this up on exit. Since you’ve confirmed you’re properly closing your MPI files, this is another library-level residual allocation—not a leak in your code.
2. Segmentation Fault (Signal 11) During MPI_Init
The segfault at rank 0 points to an issue when Open MPI initializes, specifically when it uses hwloc to build the system topology. Possible causes:
- A bug in the version of hwloc or Open MPI you’re running (older versions had issues with certain hardware topologies).
- Corrupted command-line arguments passed to
MPI_Init—double-check thatargcandargvare intact before callingMPI_Init(e.g., don’t modify them in a way that invalidates pointers prior to initialization). - Stack overflow or corrupted stack state before MPI starts up.
I’d suggest debugging this with gdb alongside MPI:
mpirun -n 1 gdb ./your_executable
Once in gdb, type run to start the program. When it segfaults, use bt (backtrace) to get a detailed stack trace—this will show exactly where the crash happens in the library code.
3. Uninitialized Stack Bytes in MPI_Finalize
The "Syscall param write(buf) points to uninitialised byte(s)" errors during MPI_Finalize come from the PSM (Portals-Style Messaging) library, which handles InfiniPath interconnects. The error means the library is accessing uninitialized stack memory during cleanup.
This is almost certainly a bug in the PSM/InfiniPath library, not your code:
- The uninitialized memory is allocated in a stack frame from
libpsm_infinipath.so, not your code’s stack. - If this noise is distracting you in Valgrind, you can suppress it by creating a suppression file that matches the error pattern. Alternatively, check if an updated version of the PSM library is available for your system.
4. Actionable Next Steps
- Update your libraries: Look for newer versions of Open MPI, hwloc, and PSM—many of these issues are fixed in recent releases.
- Test a minimal MPI program: Write a simple test with just
MPI_Init,MPI_Finalize, and anMPI_File_opencall. If the errors persist, the problem is definitely with the libraries, not your code. - Suppress harmless Valgrind errors: Create a Valgrind suppression file to filter out library-specific "still reachable" and uninitialized stack errors, so you can focus on real issues in your code.
- Debug the segfault thoroughly: Use the gdb method mentioned earlier to get a precise stack trace—this will help you confirm if it’s a library bug or a subtle issue in how you’re calling MPI functions.
内容的提问来源于stack exchange,提问作者Antonella Longo

