Ehcache+JGroups引发Tomcat内存耗尽问题排查求助
Ehcache+JGroups引发Tomcat内存耗尽问题排查求助
大家好,我最近碰到一个棘手的内存泄漏问题,折腾了好一阵没找到根因,想请各位大佬帮忙分析下:
我的部署环境是这样的:
- 两台Tomcat服务器,运行完全相同的JSF Web应用
- Nginx做负载均衡,把请求分发到两台Tomcat节点
- 应用用Ehcache结合JGroups做缓存跨节点同步,一开始缓存变更通知都能正常工作,但过段时间后,两台Tomcat的内存会慢慢被耗尽
我抓了堆dump分析后发现,单个字节数组实例就占用了1.6GB内存,对应的调用栈信息如下:
Connection.Receiver [11.130.1.175:40014 - 11.130.1.175:40014],EH_CACHE,TOMCAT_1-26339 at java.net.SocketInputStream.socketRead0(Ljava/io/FileDescriptor;[BIII)I (SocketInputStream.java(Native Method)) at java.net.SocketInputStream.socketRead(Ljava/io/FileDescriptor;[BIII)I (SocketInputStream.java:116) at java.net.SocketInputStream.read([BIII)I (SocketInputStream.java:171) at java.net.SocketInputStream.read([BII)I (SocketInputStream.java:141) at java.io.BufferedInputStream.read1([BII)I (BufferedInputStream.java:284) at java.io.BufferedInputStream.read([BII)I (BufferedInputStream.java:345) at java.io.DataInputStream.readFully([BII)V (DataInputStream.java:195) at org.jgroups.blocks.TCPConnectionMap$TCPConnection$ConnectionPeerReceiver.run()V (TCPConnectionMap.java:595) at java.lang.Thread.run()V (Thread.java:750)
下面是我的相关配置,供大家参考:
Nginx 配置(截断版)
http { server_tokens off ; include /etc/nginx/mime.types; default_type application/octet-stream; set_real_ip_from 192.168.0.0/16; set_real_ip_from 11.0.0.0/8; set_real_ip_from 174.16.0.0/16; real_ip_header X-Forwarded-For; real_ip_recursive on; log_format main '$scheme - $remote_addr ' '$http_x_forwarded_for - $remote_user [$time_local] ' '$request $status $bytes_sent ' '$http_referer $http_user_agent ' '$gzip_ratio' ; client_header_timeout 11m; client_body_timeout 11m; send_timeout 11m; client_max_body_size 50m; client_body_buffer_size 11m; sendfile on; tcp_nopush on; tcp_nodelay on; keepalive_timeout 75 20; ignore_invalid_headers on; map $http_upgrade $connection_upgrade { default upgrade; '' close; } map $upstream_addr $group { default ""; ~TOMCAT_1:80$ common; ~TOMCAT_2:80$ common; } upstream default_upstream{ server TOMCAT_1; server TOMCAT_2; sticky path=/; keepalive 64;} include /etc/nginx/conf.d/*.conf; }
Ehcache 配置(两台Tomcat完全一致)
<ehcache xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:noNamespaceSchemaLocation="http://ehcache.org/ehcache.xsd" monitoring="autodetect" dynamicConfig="true" updateCheck="false"> <!-- Disk Store configuration --> <diskStore path="java.io.tmpdir/ehcache" /> <cacheManagerPeerProviderFactory class="net.sf.ehcache.distribution.jgroups.JGroupsCacheManagerPeerProviderFactory" properties="connect=TCP( bind_port=40001; port_range=30; enable_diagnostics=false): TCPPING(initial_hosts=TOMCAT_1[40001],TOMCAT_2[40001]; port_range=30; num_initial_members=1; ergonomics=false): FD_ALL: VERIFY_SUSPECT: MERGE3: pbcast.NAKACK2(use_mcast_xmit=false): UNICAST2: pbcast.STABLE: pbcast.GMS(view_bundling=true; print_local_addr=true; log_view_warnings=true): UFC(max_credits=4M): MFC(max_credits=4M): FRAG2: pbcast.STATE_TRANSFER" propertySeparator="::" /> <cache name="org.hibernate.cache.UpdateTimestampsCache" maxElementsInMemory="100000" eternal="true" overflowToDisk="false" diskPersistent="false" timeToIdleSeconds="0" timeToLiveSeconds="0" statistics="false"> <cacheEventListenerFactory class="net.sf.ehcache.distribution.jgroups.JGroupsCacheReplicatorFactory" properties="replicateAsynchronously=true, replicatePuts=true, replicateUpdates=true, replicateUpdatesViaCopy=true, replicatePutsViaCopy=true, replicateRemovals=false" /> </cache> <cache name="org.hibernate.cache.StandardQueryCache" maxElementsInMemory="100000" eternal="false" overflowToDisk="false" diskPersistent="false" timeToIdleSeconds="600" timeToLiveSeconds="600" statistics="false"> </cache> <defaultCache maxElementsInMemory="100000" eternal="false" overflowToDisk="false" diskPersistent="false" timeToIdleSeconds="600" timeToLiveSeconds="600" statistics="false"> <cacheEventListenerFactory class="net.sf.ehcache.distribution.jgroups.JGroupsCacheReplicatorFactory" properties="replicateAsynchronously=true, replicatePuts=true, replicateUpdates=true, replicateUpdatesViaCopy=true, replicatePutsViaCopy=true, replicateRemovals=false" /> </defaultCache> </ehcache>
目前我猜测是JGroups的TCP连接处理逻辑有问题,导致数据堆积在输入流的字节数组中无法释放,但不确定是配置参数不合理还是JGroups本身的潜在bug。有没有大佬遇到过类似的情况?或者有什么具体的调试方向可以推荐的?
内容来源于stack exchange
相关产品推荐
相关产品推荐

