Spring Integration LoggingHandler抛出StackOverflowError问题求助
我遇到一个奇怪的StackOverflowError问题,由org.springframework.integration.handler.LoggingHandler抛出,且无太多详细信息。该问题发生在流终止之后,具体是在savingLionProcessingRequestToDB流终止后出现。
已进行的排查
- 确认输入通道与输出通道未重复,不存在循环问题;
- 调试发现,在数据持久化的终止步骤(即DSL的
.to(savingLionProcessingRequestToDB())方法)中,线程进入日志所述的饥饿阶段并抛出StackOverflowError。
相关日志信息
2022-11-14T21:56:22,176 WARN [HikariPool-1 housekeeper] com.zaxxer.hikari.pool.HikariPool:HikariPool-1 - Thread starvation or clock leap detected (housekeeper delta=1h40m40s250ms340µs500ns). ERROR [pool-19-thread-1] org.springframework.integration.handler.LoggingHandler: java.lang.StackOverflowError at org.springframework.integration.support.channel.BeanFactoryChannelResolver.resolveDestination(BeanFactoryChannelResolver.java:86) at org.springframework.integration.support.channel.BeanFactoryChannelResolver.resolveDestination(BeanFactoryChannelResolver.java:45) at org.springframework.messaging.core.AbstractDestinationResolvingMessagingTemplate.resolveDestination(AbstractDestinationResolvingMessagingTemplate.java:78) at org.springframework.messaging.core.AbstractDestinationResolvingMessagingTemplate.send(AbstractDestinationResolvingMessagingTemplate.java:71) at org.springframework.integration.handler.AbstractMessageProducingHandler.sendOutput(AbstractMessageProducingHandler.java:465) at org.springframework.integration.handler.AbstractMessageProducingHandler.doProduceOutput(AbstractMessageProducingHandler.java:325) at org.springframework.integration.handler.AbstractMessageProducingHandler.produceOutput(AbstractMessageProducingHandler.java:268)
相关分散收集流代码
MessagingGateway定义
@MessagingGateway public interface LionGateway { @Gateway(requestChannel = "flow.input", replyChannel = "lionReplyChannel") Object processlionRequest( @Payload Message lionRequest, @Header("dbID") Long dbID, @Header(value = "reqNumber", required = false) String requestNumber); }
主流程定义
// Main flow @Bean public IntegrationFlow flow() { return flow -> flow.handle(validatorService, "validateRequest") .split() .channel(c -> c.executor(Executors.newCachedThreadPool())) .scatterGather( scatterer -> scatterer .applySequence(true) .recipientFlow(saveLionDB()) .recipientFlow(preparingRequestFromLionRequest()), gatherer -> gatherer.outputProcessor(prepareLionRequest())) .to(sampleFlow1()); } @Bean public IntegrationFlow sampleFlow1() { return flow -> flow.enrichHeaders(h -> h.replyChannel("responseChannel", true)) .handle( Http.outboundGateway(serviceURL, restTemplateConfig.restTemplate()) .mappedRequestHeaders("type", "name") .httpMethod(HttpMethod.POST) .expectedResponseType(response.class)) .handle(lionsService, "lionService1") .logAndReply("Response"); }
后续处理流程
// Another flow being leveraged further in processing :- @Bean @BridgeFrom("responseChannel") public MessageChannel lionChannel() { return MessageChannels.direct().get(); } @Bean public IntegrationFlow lionSampleFlow() { return IntegrationFlows.from(lionChannel()) .split() .channel(c -> c.executor(Executors.newCachedThreadPool())) .log("payload") .scatterGather( scatterer -> scatterer .applySequence(true) .recipientFlow(getLionResponse()) .recipientFlow(prepareResponse()), gatherer -> gatherer.outputProcessor(sendLionResponse())) .to(savingLionProcessingRequestToDB()); } @Bean public IntegrationFlow savingLionProcessingRequestToDB() { return flow -> flow.filter( conditionToPersistDBData, f -> f.discardFlow(BaseIntegrationFlowDefinition::bridge)) .channel(c -> c.executor(Executors.newCachedThreadPool())) .enrichHeaders(h -> h.errorChannel("lionErrorChannel1", true)) .enrichHeaders(h -> h.replyChannel(lionReplyChannel(), true)) // defined at gateway layer as the replyChannel .handle( (payload, header) -> lionsService.saveLionResponse( header.get("dbID", Long.class), header.get("reqNumber", String.class), header.get("createdAt", Timestamp.class))) .logAndReply("persist_response"); } public LionModel saveLionResponse( long dbId, String requestNumber, Timestamp createdAt) { // return lionProcessingRepository.save(lionModel); }
原因分析
线程饥饿引发连锁反应:日志先出现Hikari连接池的线程饥饿警告,说明连接池管家线程长时间无法正常运行,大概率是应用中大量线程被阻塞(比如数据库操作耗时过长、线程池耗尽),导致JVM线程调度异常,进而引发后续调用栈溢出。
logAndReply与手动设置回复通道的冲突:savingLionProcessingRequestToDB流中,既通过.enrichHeaders设置了网关层的lionReplyChannel作为回复通道,又使用了.logAndReply。logAndReply会自动将结果发送到当前消息的回复通道,若回复通道的下游逻辑触发了递归调用(或者线程饥饿导致消息发送逻辑异常循环),就会引发StackOverflowError,从堆栈信息看,确实卡在了通道解析和消息发送的循环调用中。无界线程池的隐患:多个流使用
Executors.newCachedThreadPool(),这种线程池会无限制创建线程,请求量大时会导致JVM线程数量过多,上下文切换开销剧增,既容易引发线程饥饿,也会因每个线程的栈空间占用,增加栈溢出风险。
解决方案
修复
logAndReply的使用逻辑:- 如果
savingLionProcessingRequestToDB是终止流,不需要返回结果到网关,移除.enrichHeaders(h -> h.replyChannel(lionReplyChannel(), true)),或者将.logAndReply改为.log,避免自动发送回复。 - 如果确实需要发送结果到
lionReplyChannel,改用.channel(lionReplyChannel())显式发送,避免logAndReply的默认逻辑与手动配置冲突。
- 如果
优化线程池配置:
- 替换无界线程池为有界线程池,比如根据业务场景调整线程数:
Executors.newFixedThreadPool(10) - 调整Hikari连接池参数(如
maximumPoolSize、connectionTimeout),确保连接池管家线程正常运行,避免数据库连接耗尽导致线程阻塞。
- 替换无界线程池为有界线程池,比如根据业务场景调整线程数:
排查数据库操作性能:
- 检查
saveLionResponse中的数据库操作是否存在性能瓶颈(如缺少索引、数据量过大),优化读写逻辑,减少线程阻塞时间。
- 检查
检查回复通道下游逻辑:
- 确认
lionReplyChannel的下游处理是否存在循环调用,比如是否有流监听该通道后又将消息发送回上游流,导致消息循环。
- 确认
内容的提问来源于stack exchange,提问作者Aurora

