AIX-6环境下第三方无限循环应用Core Dump获取问题
Let's break down this problem and walk through the likely causes and fixes, since you've already ruled out basic system configuration issues (great job validating with sleep and checking the official guidelines!).
Problem Recap
You're working on AIX 6, trying to capture a core dump from a third-party executable stuck in an infinite loop. Running kill -11 $thepid terminates the process, but no core dump is generated. However, testing with a sleep 60 process works perfectly—so your system's core dump setup is generally functional.
Likely Causes & Fixes
1. The Third-Party Program Handles SIGSEGV (Signal 11) Customly
Many applications register custom handlers for SIGSEGV to run cleanup logic or error reporting instead of letting the kernel generate a core dump. If this is the case, kill -11 will trigger the app's custom code rather than core creation.
How to check & fix:
- Use
procflags $thepidto inspect the process's signal handling state. Look forSIGSEGVin the output—if it showscaughtinstead ofdefault, the app is intercepting the signal. - Attach to the process with
gdbto force a core dump:gdb -p $thepid # Once in GDB, run either of these commands: generate-core-file # Directly generates a core dump file call abort() # Triggers an abort signal, which should produce a core
2. The Process Has Core Dump Limits Set to Zero
It's possible the third-party program (or its startup script) explicitly sets a core dump limit of 0 via setrlimit, overriding your system-wide settings.
How to check & fix:
- Check the process's resource limits with
prlimit -c $thepid. If thesoftorhardlimit for core files is 0, that's the blockage. - If you can modify the startup process, add
ulimit -c unlimitedbefore launching the executable. If you can't edit the startup script, use thegdbmethod above to force a core dump anyway.
3. The Process's Working Directory Lacks Write Permissions
Even if your system allows core dumps, the process needs write access to its current working directory to save the core file. The sleep test likely runs in a directory you have access to, but the third-party app might be running elsewhere.
How to check & fix:
- Find the process's working directory with
pwdx $thepid. - Verify the user running the process has write permissions to that directory. If not, either adjust the directory permissions or configure a global core dump path with
coreadm:# Create the directory first if it doesn't exist mkdir -p /var/core chmod 777 /var/core # Set a global core path and enable global core dumps coreadm -g /var/core/core.%n.%p coreadm -e global
4. The Program is a Setuid/Setgid Executable
AIX has security restrictions that prevent setuid/setgid programs from generating core dumps by default (to avoid leaking sensitive data). If the third-party app has the s permission bit set, this could be blocking core creation.
How to check & fix:
- Check the executable's permissions with
ls -l /path/to/third-party-executable. Look forsin the user or group permissions (e.g.,-rwsr-xr-x). - Enable core dumps for setuid/setgid processes with:
Note: This has security implications, so only do this if you trust the executable and understand the risks of exposing sensitive data in core dumps.coreadm -e setid
5. The Process is Stuck in a Kernel-Mode Loop (Rare)
In rare cases, an infinite loop running entirely in kernel space might not respond to SIGSEGV in a way that triggers a core dump. This is less likely since sleep works, but it's worth checking.
How to check & fix:
- Use
ps -k $thepidto see if the process is in kernel mode (look forKin theSTATEcolumn). - If it's stuck in kernel space, you might need to use the AIX
snapcommand or reach out to IBM support for deeper kernel-level debugging—this is a last resort.
内容的提问来源于stack exchange,提问作者Be Kind To New Users

