You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Ehcache+JGroups引发Tomcat内存耗尽问题排查求助

Ehcache+JGroups引发Tomcat内存耗尽问题排查求助

大家好,我最近碰到一个棘手的内存泄漏问题,折腾了好一阵没找到根因,想请各位大佬帮忙分析下:

我的部署环境是这样的:

  • 两台Tomcat服务器,运行完全相同的JSF Web应用
  • Nginx做负载均衡,把请求分发到两台Tomcat节点
  • 应用用Ehcache结合JGroups做缓存跨节点同步,一开始缓存变更通知都能正常工作,但过段时间后,两台Tomcat的内存会慢慢被耗尽

我抓了堆dump分析后发现,单个字节数组实例就占用了1.6GB内存,对应的调用栈信息如下:

Connection.Receiver [11.130.1.175:40014 - 11.130.1.175:40014],EH_CACHE,TOMCAT_1-26339
at java.net.SocketInputStream.socketRead0(Ljava/io/FileDescriptor;[BIII)I (SocketInputStream.java(Native Method))
at java.net.SocketInputStream.socketRead(Ljava/io/FileDescriptor;[BIII)I (SocketInputStream.java:116)
at java.net.SocketInputStream.read([BIII)I (SocketInputStream.java:171)
at java.net.SocketInputStream.read([BII)I (SocketInputStream.java:141)
at java.io.BufferedInputStream.read1([BII)I (BufferedInputStream.java:284)
at java.io.BufferedInputStream.read([BII)I (BufferedInputStream.java:345)
at java.io.DataInputStream.readFully([BII)V (DataInputStream.java:195)
at org.jgroups.blocks.TCPConnectionMap$TCPConnection$ConnectionPeerReceiver.run()V (TCPConnectionMap.java:595)
at java.lang.Thread.run()V (Thread.java:750)

下面是我的相关配置,供大家参考:

Nginx 配置(截断版)

http {
 server_tokens off ;
 include /etc/nginx/mime.types;
 default_type application/octet-stream;
 set_real_ip_from 192.168.0.0/16;
 set_real_ip_from 11.0.0.0/8;
 set_real_ip_from 174.16.0.0/16;
 real_ip_header X-Forwarded-For;
 real_ip_recursive on;
 log_format main '$scheme - $remote_addr '
 '$http_x_forwarded_for - $remote_user [$time_local] '
 '$request $status $bytes_sent '
 '$http_referer $http_user_agent '
 '$gzip_ratio' ;
 client_header_timeout 11m;
 client_body_timeout 11m;
 send_timeout 11m;
 client_max_body_size 50m;
 client_body_buffer_size 11m;
 sendfile on;
 tcp_nopush on;
 tcp_nodelay on;
 keepalive_timeout 75 20;
 ignore_invalid_headers on;
 map $http_upgrade $connection_upgrade {
 default upgrade;
 '' close;
 }
 map $upstream_addr $group {
 default "";
 ~TOMCAT_1:80$ common;
 ~TOMCAT_2:80$ common;
 }
 upstream default_upstream{
 server TOMCAT_1;
 server TOMCAT_2;
 sticky path=/;
 keepalive 64;}
 include /etc/nginx/conf.d/*.conf;
}

Ehcache 配置(两台Tomcat完全一致)

<ehcache xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
 xsi:noNamespaceSchemaLocation="http://ehcache.org/ehcache.xsd"
 monitoring="autodetect"
 dynamicConfig="true"
 updateCheck="false">
 <!-- Disk Store configuration -->
 <diskStore path="java.io.tmpdir/ehcache" />
 <cacheManagerPeerProviderFactory
 class="net.sf.ehcache.distribution.jgroups.JGroupsCacheManagerPeerProviderFactory"
 properties="connect=TCP( bind_port=40001; port_range=30; enable_diagnostics=false): TCPPING(initial_hosts=TOMCAT_1[40001],TOMCAT_2[40001]; port_range=30; num_initial_members=1; ergonomics=false): FD_ALL: VERIFY_SUSPECT: MERGE3: pbcast.NAKACK2(use_mcast_xmit=false): UNICAST2: pbcast.STABLE: pbcast.GMS(view_bundling=true; print_local_addr=true; log_view_warnings=true): UFC(max_credits=4M): MFC(max_credits=4M): FRAG2: pbcast.STATE_TRANSFER"
 propertySeparator="::" />
 <cache name="org.hibernate.cache.UpdateTimestampsCache"
 maxElementsInMemory="100000"
 eternal="true"
 overflowToDisk="false"
 diskPersistent="false"
 timeToIdleSeconds="0"
 timeToLiveSeconds="0"
 statistics="false">
 <cacheEventListenerFactory
 class="net.sf.ehcache.distribution.jgroups.JGroupsCacheReplicatorFactory"
 properties="replicateAsynchronously=true, replicatePuts=true, replicateUpdates=true, replicateUpdatesViaCopy=true, replicatePutsViaCopy=true, replicateRemovals=false" />
 </cache>
 <cache name="org.hibernate.cache.StandardQueryCache"
 maxElementsInMemory="100000"
 eternal="false"
 overflowToDisk="false"
 diskPersistent="false"
 timeToIdleSeconds="600"
 timeToLiveSeconds="600"
 statistics="false">
 </cache>
 <defaultCache
 maxElementsInMemory="100000"
 eternal="false"
 overflowToDisk="false"
 diskPersistent="false"
 timeToIdleSeconds="600"
 timeToLiveSeconds="600"
 statistics="false">
 <cacheEventListenerFactory
 class="net.sf.ehcache.distribution.jgroups.JGroupsCacheReplicatorFactory"
 properties="replicateAsynchronously=true, replicatePuts=true, replicateUpdates=true, replicateUpdatesViaCopy=true, replicatePutsViaCopy=true, replicateRemovals=false" />
 </defaultCache>
</ehcache>

目前我猜测是JGroups的TCP连接处理逻辑有问题,导致数据堆积在输入流的字节数组中无法释放,但不确定是配置参数不合理还是JGroups本身的潜在bug。有没有大佬遇到过类似的情况?或者有什么具体的调试方向可以推荐的?

内容来源于stack exchange

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.08 09:59:51