JBoss服务器‘Too many open files’问题排查咨询——systemd启动场景下的异常情况
Let's break down your problem step by step—first clearing up the confusing discrepancy between lsof and /proc/fd counts, then diving into actionable checks to fix the root cause:
1. Why lsof and /proc/fd counts don't match
Your lsof output shows nearly 500k entries, but /proc/42777/fd only lists ~2.7k files. This is not a bug—here's the key difference:
lsofcounts every instance of a file descriptor being used by a thread. JBoss (a Java application) runs hundreds of threads, and a single shared file descriptor (like a socket or connection pool handle) will be listed once per thread that uses it. This leads to massive duplicate counting./proc/<pid>/fdis the authoritative count of unique file descriptors held by the entire process—this is the number that directly corresponds to theulimit -n(open files) limit.
So your actual open file count when running is ~2.7k, which is well below 30k. The "too many open files" error likely happens during a transient spike in file descriptors during startup—or the limit applied to the JBoss process isn't actually 30k at all.
2. Verify the actual open files limit applied to JBoss
The biggest pitfall when switching from shell scripts to systemd is assuming user-level ulimit settings are inherited. Systemd manages process limits independently, even when switching to a non-root user.
Check the effective limit for your JBoss process:
cat /proc/42777/limits | grep "Max open files"If this shows anything lower than 30000 (like the systemd default of 1024), that's the immediate problem.
Fix your systemd service file (e.g.,
/etc/systemd/system/jboss.service):
Add these lines to the[Service]section:LimitNOFILE=30000 User=appuser(You can increase this to 65536 if you need extra headroom.) Then reload systemd and restart JBoss:
systemctl daemon-reload systemctl restart jboss.service
3. Align user-level limits for consistency
Even with systemd's LimitNOFILE, ensure appuser's limits are set properly in case systemd falls back to them:
- Edit
/etc/security/limits.conf(or a file in/etc/security/limits.d/):
Note: Systemd'sappuser soft nofile 30000 appuser hard nofile 30000LimitNOFILEtakes precedence over these settings, but aligning them avoids future confusion.
4. Track transient file descriptor spikes during startup
Since the error happens during startup (but the final count is low), you need to catch the moment when descriptors exceed the limit:
- Use
straceto log all file operations during startup:
(Or attach to the process immediately after starting withstrace -e open,openat,close -f -o jboss_fd_trace.log /path/to/jboss/startup/commandstrace -e open,openat,close -p <pid>.) - Alternatively, run
lsofin a loop to track counts over time:
Look for sudden jumps that cross your limit.while true; do echo "$(date) $(lsof -p 42777 | wc -l)"; sleep 1; done
5. Investigate application-level leaks in the new deployment
The problem started after a new deployment, so focus on changes that could increase file descriptor usage:
- Check components highlighted in your
lsofoutput: ActiveMQ, Elasticsearch connections, database pools. Did the new version increase connection pool sizes? - Look for code that opens files/sockets without closing them—e.g., unclosed
FileInputStreaminstances, missingclose()calls on network connections. - Review JBoss configuration changes: Did you enable new services, listeners, or logging handlers that open more files?
6. Rule out other systemd restrictions
While less likely for this specific error, double-check other limits that could interfere:
- Check max processes/threads limit:
If your JBoss process spawns more threads than this limit allows, it can trigger related errors. Adjust withcat /proc/42777/limits | grep "Max processes"LimitNPROC=4096(or higher) in the systemd service file if needed.
内容的提问来源于stack exchange,提问作者Vlastimil Unucka

