使用Ribbon时Connection refused无法触发重试的问题排查求助
问题分析与解决方案
环境信息
spring-cloud.version=Dalston.SR5 spring-boot-starter-parent=1.5.7.RELEASE zuul-core=1.3.0
问题场景
当客户端实例下线(杀死进程)后,Zuul仍缓存该实例信息,请求被路由至已下线实例抛出Connection refused异常。明明配置了Ribbon重试和Zuul重试参数,也引入了spring-retry依赖,但重试逻辑完全没触发——调试发现异常在RetryableRibbonLoadBalancingHttpClient.delegate.execute()处抛出,重试逻辑根本没机会执行。
已配置参数
ribbon: MaxAutoRetries: 1 MaxAutoRetriesNextServer: 2 OkToRetryOnAllOperations: true ReadTimeout: 1000 ConnectTimeout: 250 ServerListRefreshInterval: 1000 zuul: retryable: true
已引入依赖
<dependency> <groupId>org.springframework.retry</groupId> <artifactId>spring-retry</artifactId> </dependency>
问题根源
在Dalston版本的RetryableRibbonLoadBalancingHttpClient中,默认的重试逻辑没有把HttpHostConnectException(也就是底层的Connection refused异常)纳入可重试异常范围。重试逻辑写在delegate.execute()之后,但这个方法抛出的连接异常没被RetryTemplate识别,直接跳过了重试流程。
解决办法
1. 自定义Ribbon重试策略,扩展可重试异常
创建自定义的RetryHandler和RetryTemplate,把连接类异常加入重试列表:
import com.netflix.client.RetryHandler; import com.netflix.client.config.IClientConfig; import org.apache.http.conn.HttpHostConnectException; import org.springframework.cloud.netflix.ribbon.RibbonClient; import org.springframework.context.annotation.Bean; import org.springframework.context.annotation.Configuration; import org.springframework.retry.policy.SimpleRetryPolicy; import org.springframework.retry.support.RetryTemplate; import java.util.HashMap; import java.util.Map; @Configuration @RibbonClient(name = "你的服务ID", configuration = CustomRibbonConfig.class) public class CustomRibbonConfig { @Bean public RetryHandler retryHandler() { return new RetryHandler() { @Override public boolean isRetriableException(Throwable e, boolean sameServer) { // 让连接异常、IO异常都能触发重试 return e instanceof HttpHostConnectException || e instanceof java.io.IOException; } @Override public boolean isCircuitTrippingException(Throwable e) { return false; } @Override public int getMaxRetriesOnSameServer() { return 1; // 对应ribbon.MaxAutoRetries配置 } @Override public int getMaxRetriesOnNextServer() { return 2; // 对应ribbon.MaxAutoRetriesNextServer配置 } }; } @Bean public RetryTemplate retryTemplate(IClientConfig config) { RetryTemplate retryTemplate = new RetryTemplate(); SimpleRetryPolicy retryPolicy = new SimpleRetryPolicy(); // 设置总重试次数(初始请求+重试次数) retryPolicy.setMaxAttempts(config.getIntegerProperty("ribbon.MaxAutoRetries", 1) + 1); // 配置需要重试的异常类型 Map<Class<? extends Throwable>, Boolean> retryableExceptions = new HashMap<>(); retryableExceptions.put(HttpHostConnectException.class, true); retryableExceptions.put(java.io.IOException.class, true); retryPolicy.setRetryableExceptions(retryableExceptions); retryTemplate.setRetryPolicy(retryPolicy); return retryTemplate; } }
2. 优化Eureka实例感知速度,减少缓存影响
在Eureka客户端配置里添加健康检查和实例过期配置,让Eureka更快剔除下线实例,从根源减少路由到无效节点的概率:
eureka: client: serviceUrl: defaultZone: http://你的Eureka地址:8761/eureka/ instance: lease-renewal-interval-in-seconds: 5 # 缩短心跳间隔 lease-expiration-duration-in-seconds: 15 # 缩短实例过期时间 health-check-url-path: /actuator/health # 开启健康检查
3. 添加重试监控,验证逻辑是否生效
可以加个RetryListener来打印重试日志,确认重试是否真的触发:
@Bean public org.springframework.retry.RetryListener retryListener() { return new org.springframework.retry.RetryListener() { @Override public <T, E extends Throwable> boolean open(org.springframework.retry.RetryContext context, org.springframework.retry.RetryCallback<T, E> callback) { return true; } @Override public <T, E extends Throwable> void close(org.springframework.retry.RetryContext context, org.springframework.retry.RetryCallback<T, E> callback, Throwable throwable) { if (throwable != null) { System.out.println("重试流程结束,最终异常:" + throwable.getMessage()); } } @Override public <T, E extends Throwable> void onError(org.springframework.retry.RetryContext context, org.springframework.retry.RetryCallback<T, E> callback, Throwable throwable) { System.out.println("触发重试,异常类型:" + throwable.getClass().getSimpleName()); } }; }
补充说明
Dalston是比较旧的Spring Cloud版本,上述配置是针对该版本的适配方案。如果后续有升级计划,建议考虑升级到Greenwich及以上版本——新版本的Zuul/Ribbon重试逻辑更完善,对异常的覆盖更全面,很多这类问题都已经被修复。
内容的提问来源于stack exchange,提问作者Hope Dc
相关产品推荐
相关产品推荐

