循环启动多管道子进程的程序运行一段时间后崩溃问题排查
Let’s break down what’s happening here and how to fix it. That error almost always points to file descriptor exhaustion—your monitoring program is opening more files/sockets/pipes than it’s closing, until it hits the system or process-level limit. Here’s how to diagnose and fix the issue:
1. Confirm File Descriptor Exhaustion First
First, verify this is indeed the problem:
- Check your process’s file descriptor limit with
ulimit -n(run this as the user running the monitor). The default is often 1024. - Count how many file descriptors your monitor is currently holding open:
# Replace <monitor_pid> with your monitor program's PID ls /proc/<monitor_pid>/fd | wc -l - If the count is close to or exceeds the
ulimit -nvalue, you’ve confirmed the leak.
2. Your Pipeline Command Is Likely the Culprit
The pipeline ps aux | grep <proc_name> | grep -v grep | grep -v <program_name> creates multiple subprocesses and pipe file descriptors every time it runs. If this is in a loop (like a while true loop checking every few seconds), your shell might not clean up these FDs properly over time, leading to a slow leak.
Fix: Simplify the Monitoring Command
Replace that multi-pipe mess with a more efficient, FD-friendly alternative:
- Use
pgrep(cleanest option for exact process name matches):# Checks if <proc_name> is running; starts it if not if ! pgrep -x "<proc_name>" > /dev/null; then /path/to/your/target/process & fi - Or use
pswith built-in filtering to avoid extragrepcalls:# Lists only processes named <proc_name>, excluding your monitor script if ! ps aux --no-heading -C "<proc_name>" | grep -v "<program_name>" > /dev/null; then /path/to/your/target/process & fi
Both options cut down on the number of subprocesses and pipes created per check, reducing FD usage significantly.
3. Diagnose Leaks in the Monitor Program Itself
If simplifying the command doesn’t fix the issue, check if your monitor (whether shell script or custom binary) is leaking FDs elsewhere:
- For shell scripts: Ensure you’re redirecting unused output/errors properly (e.g.,
> /dev/null 2>&1for commands that don’t need to print anything) and that any temporary files/sockets are closed or deleted after use. - For custom binaries: Use
straceto track FD operations in real time:
Look for patterns wherestrace -e open,close,dup2 -p <monitor_pid>opencalls aren’t followed by correspondingclosecalls—those are your leak points.
4. Temporary Mitigation (Prioritize Long-Term Fix)
If you need to keep the monitor running while troubleshooting, you can temporarily increase the file descriptor limit:
- For the current session:
ulimit -n 4096 - For permanent changes (edit
/etc/security/limits.conf):your_username soft nofile 4096 your_username hard nofile 8192
Note: This is a band-aid, not a solution. You still need to find and plug the root leak.
内容的提问来源于stack exchange,提问作者rusty doe

