SpringBoot+Infinispan应用周期性无响应:线程死锁排查求助
Spring Boot + Infinispan 13.0.12 周期性无响应问题分析与优化建议
问题现象
运行基于Spring Boot的REST应用,集成Infinispan 13.0.12缓存后,出现周期性随机无响应。线程转储显示超过200个请求线程(http-nio-8080-exec-*)处于TIMED_WAITING状态,阻塞点集中在JGroups的流控逻辑。
线程转储关键信息
"http-nio-8080-exec-379" #11999 daemon prio=5 os_prio=0 tid=0x00007f28900f9800 nid=0x2c68 waiting on condition [0x00007f28485c2000] java.lang.Thread.State: TIMED_WAITING (parking) at sun.misc.Unsafe.park(Native Method) - parking to wait for <0x00000006c09af3e8> (a java.util.concurrent.locks.AbstractQueuedSynchronizer$ConditionObject) at java.util.concurrent.locks.LockSupport.parkNanos(LockSupport.java:215) at java.util.concurrent.locks.AbstractQueuedSynchronizer$ConditionObject.await(AbstractQueuedSynchronizer.java:2163) at org.jgroups.util.Credit.decrementIfEnoughCredits(Credit.java:65) at org.jgroups.protocols.UFC.handleDownMessage(UFC.java:119) at org.jgroups.protocols.FlowControl.down(FlowControl.java:323) at org.jgroups.protocols.FlowControl.down(FlowControl.java:317) at org.jgroups.protocols.FRAG3.down(FRAG3.java:139) at org.jgroups.stack.ProtocolStack.down(ProtocolStack.java:927) at org.jgroups.JChannel.down(JChannel.java:645) at org.jgroups.JChannel.send(JChannel.java:484) at org.infinispan.remoting.transport.jgroups.JGroupsTransport.send(JGroupsTransport.java:1161)
当前配置信息
Java配置
@Autowired @Bean public SpringEmbeddedCacheManagerFactoryBean springEmbeddedCacheManagerFactoryBean(GlobalConfigurationBuilder gcb, ConfigurationBuilder configurationBuilder) { SpringEmbeddedCacheManagerFactoryBean springEmbeddedCacheManagerFactoryBean = new SpringEmbeddedCacheManagerFactoryBean(); springEmbeddedCacheManagerFactoryBean.addCustomGlobalConfiguration(gcb); springEmbeddedCacheManagerFactoryBean.addCustomCacheConfiguration(configurationBuilder); return springEmbeddedCacheManagerFactoryBean; } @Autowired @Bean public EmbeddedCacheManager defaultCacheManager(SpringEmbeddedCacheManager springEmbeddedCacheManager) throws Exception { return springEmbeddedCacheManager.getNativeCacheManager(); } @Bean public GlobalConfigurationBuilder globalConfigurationBuilder() { GlobalConfigurationBuilder result = GlobalConfigurationBuilder.defaultClusteredBuilder(); result.transport().addProperty("configurationFile", jgroupsConfigFile); result.cacheManagerName(IDENTITY_CACHE); result.defaultCacheName(IDENTITY_CACHE + "-default"); result.serialization() .marshaller(new JavaSerializationMarshaller()) .allowList() .addClasses( LinkedMultiValueMap.class, String.class ); result.globalState().enable().persistentLocation(DATA_DIR); return result; } @Bean public ConfigurationBuilder configurationBuilder() { ConfigurationBuilder result = new ConfigurationBuilder(); result.clustering().cacheMode(CacheMode.REPL_SYNC); return result; } @Bean public org.infinispan.configuration.cache.Configuration cacheConfiguration() { ConfigurationBuilder builder = new ConfigurationBuilder(); return builder .clustering() .cacheMode(CacheMode.REPL_SYNC) .remoteTimeout(replicationTimeoutSeconds, TimeUnit.SECONDS) .stateTransfer().timeout(stateTransferTimeoutMinutes, TimeUnit.MINUTES) .persistence() .addSoftIndexFileStore() .shared(false) .fetchPersistentState(true) .expiration().lifespan(expirationHours, TimeUnit.HOURS) .build(); } @Autowired @Bean public Cache<String, MultiValueMap<String, String>> identityCache(EmbeddedCacheManager manager, org.infinispan.configuration.cache.Configuration cacheConfiguration) throws IOException { Cache<String, MultiValueMap<String, String>> result = manager .administration().withFlags(CacheContainerAdmin.AdminFlag.VOLATILE) .getOrCreateCache(IDENTITY_CACHE, cacheConfiguration); result.getAdvancedCache().getStats().setStatisticsEnabled(true); return result; }
JGroups UDP配置
<config xmlns="urn:org:jgroups" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="urn:org:jgroups http://www.jgroups.org/schema/jgroups-4.0.xsd"> <UDP mcast_addr="${jgroups.udp.mcast_addr:228.6.7.9}" mcast_port="${jgroups.udp.mcast_port:46655}" ucast_send_buf_size="1m" mcast_send_buf_size="1m" ucast_recv_buf_size="20m" mcast_recv_buf_size="25m" ip_ttl="${jgroups.ip_ttl:2}" thread_naming_pattern="pl" enable_diagnostics="false" bundler_type="no-bundler" max_bundle_size="8500" thread_pool.min_threads="${jgroups.thread_pool.min_threads:0}" thread_pool.max_threads="${jgroups.thread_pool.max_threads:200}" thread_pool.keep_alive_time="60000" /> <PING /> <MERGE3 min_interval="10000" max_interval="30000" /> <FD_SOCK /> <FD_ALL timeout="60000" interval="15000" timeout_check_interval="5000" /> <VERIFY_SUSPECT timeout="5000" /> <pbcast.NAKACK2 xmit_interval="100" xmit_table_num_rows="50" xmit_table_msgs_per_row="1024" xmit_table_max_compaction_time="30000" resend_last_seqno="true" /> <UNICAST3 xmit_interval="100" xmit_table_num_rows="50" xmit_table_msgs_per_row="1024" xmit_table_max_compaction_time="30000" conn_expiry_timeout="0" /> <pbcast.STABLE stability_delay="500" desired_avg_gossip="5000" max_bytes="1M" /> <pbcast.GMS print_local_addr="false" install_view_locally_first="true" join_timeout="${jgroups.join_timeout:5000}" /> <UFC max_credits="2m" min_threshold="0.40" /> <MFC max_credits="2m" min_threshold="0.40" /> <FRAG3 frag_size="8000"/> </config>
原因分析
线程阻塞在Credit.decrementIfEnoughCredits,这是JGroups UFC(Upper Flow Control) 机制的核心逻辑:发送端向集群节点发送消息时会消耗接收端分配的credits,当credits耗尽,发送端线程会阻塞,直到接收端返回新的credits。结合配置和现象,核心原因包括:
- 同步复制放大阻塞风险:使用
REPL_SYNC同步复制,每个缓存写操作都需要等待3个节点全部响应。一旦某个节点处理消息缓慢(如序列化耗时、资源瓶颈),所有请求线程都会阻塞在等待credits或节点响应上。 - Java序列化效率低下:
JavaSerializationMarshaller序列化速度慢、消息体积大,增加网络传输时间和节点处理负载,导致credits恢复延迟。 - 流控配置不合理:UFC的
max_credits=2m可能过小,高并发场景下很快耗尽;min_threshold=0.40设置较高,接收端需要积累更多未处理消息才会返回credits,延长发送端阻塞时间。 - 线程池配置失衡:JGroups线程池
max_threads=200过大,导致上下文切换开销增加;Tomcat请求线程池与缓存操作线程耦合,一旦缓存阻塞,所有请求线程都会被占用,引发应用无响应。
优化建议
1. 调整JGroups流控与传输配置
- 增大UFC和MFC的
max_credits,例如从2m改为8m,减少因credits耗尽导致的阻塞:<UFC max_credits="8m" min_threshold="0.20"/> <MFC max_credits="8m" min_threshold="0.20"/> - 降低
min_threshold至0.20,让接收端更早返还credits,缩短发送端等待时间。 - 启用JGroups诊断(
enable_diagnostics="true"),收集流控统计数据,定位具体是哪个节点或消息类型导致的credits耗尽。 - 调整UDP的
bundler_type为fixed-size,开启消息批量发送,减少网络IO次数:<UDP bundler_type="fixed-size" max_bundle_size="64k"/>
2. 优化Infinispan缓存模式与序列化
- 若业务允许最终一致性,将缓存模式从
REPL_SYNC改为REPL_ASYNC,写操作无需等待所有节点响应,彻底避免同步阻塞:builder.clustering().cacheMode(CacheMode.REPL_ASYNC); - 替换
JavaSerializationMarshaller为ProtoStreamMarshaller,大幅提升序列化效率并减小消息体积:
(注:ProtoStream需提前定义序列化协议或使用注解自动生成)result.serialization().marshaller(new ProtoStreamMarshaller()) .allowList().addClasses(LinkedMultiValueMap.class, String.class);
3. 线程池优化
- 调整JGroups线程池大小,建议设置为节点CPU核心数的2倍,避免线程过多导致上下文切换:
<UDP thread_pool.max_threads="16"/> <!-- 假设8核CPU --> - 限制Tomcat请求线程池大小(在
application.properties中配置),避免请求线程全部阻塞在缓存操作上:server.tomcat.max-threads=200 server.tomcat.accept-count=100
4. 缓存持久化与状态传输优化
- 若无需持久化状态,关闭
fetchPersistentState,减少节点启动时的状态传输负载:builder.persistence().addSoftIndexFileStore().fetchPersistentState(false); - 调整状态传输的
chunk_size,避免状态传输时占用过多资源:builder.clustering().stateTransfer().chunkSize(1000);
5. 集群监控与排查
- 监控各节点的CPU、内存、磁盘IO,排查是否有节点存在资源瓶颈(如磁盘IO过高导致持久化操作缓慢)。
- 利用已启用的Infinispan统计,监控缓存命中率、复制延迟等指标,定位性能瓶颈。
内容的提问来源于stack exchange,提问作者Adrian Smith
相关产品推荐
相关产品推荐

