You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spring Boot负载均衡:实现故障服务器自动忽略与切换

Spring Boot负载均衡故障节点自动切换实现方案

一、添加服务实例健康校验逻辑

你目前的实现是静态返回所有ServiceInstance,没有判断实例是否存活,导致故障节点仍会被选中调用。需要在返回实例前完成健康检查,过滤不可用节点:

@Autowired
private WebClient webClient;

@Override
public Flux<List<ServiceInstance>> get() {
    List<ServiceInstance> allInstances = Arrays.asList(
        new DefaultServiceInstance(serviceId + "1", serviceId, "localhost", 8080, false),
        new DefaultServiceInstance(serviceId + "2", serviceId, "localhost", 9092, false),
        new DefaultServiceInstance(serviceId + "3", serviceId, "localhost", 9999, false)
    );
    
    // 异步校验每个实例健康状态,过滤故障节点
    return Flux.fromIterable(allInstances)
        .flatMap(instance -> {
            String healthCheckUrl = String.format("http://%s:%s/actuator/health", instance.getHost(), instance.getPort());
            return webClient.get()
                .uri(healthCheckUrl)
                .retrieve()
                .toBodilessEntity()
                .map(res -> instance)
                .onErrorResume(ex -> Mono.empty()); // 校验失败则丢弃实例
        })
        .collectList();
}

注意:目标服务需开启Spring Boot Actuator的/actuator/health端点;若无Actuator,可替换为服务的核心业务接口作为健康校验地址。

二、配置负载均衡重试机制

仅过滤健康节点还不够,需在调用失败时自动重试其他可用实例。基于Spring Cloud LoadBalancer的配置如下:

  1. 引入依赖:
<dependency>
    <groupId>org.springframework.cloud</groupId>
    <artifactId>spring-cloud-starter-loadbalancer</artifactId>
</dependency>
<dependency>
    <groupId>org.springframework.retry</groupId>
    <artifactId>spring-retry</artifactId>
</dependency>
  1. 配置文件(application.yml)中开启重试规则:
spring:
  cloud:
    loadbalancer:
      retry:
        enabled: true
      client:
        config:
          default:
            retry:
              max-retries-on-next-server: 2 # 切换至其他实例的重试次数
              max-retries-on-current-server: 0 # 当前实例不重试
              retryable-status-codes: 500,502,503,504 # 触发重试的HTTP状态码

若使用OpenFeign调用服务,可额外配置Feign重试:

feign:
  client:
    config:
      default:
        retryer:
          max-attempts: 3
          period: 100
          max-period: 500

三、优化:缓存健康实例+定时更新

每次请求都做健康校验会影响性能,可通过定时任务维护健康实例缓存:

private List<ServiceInstance> healthyInstances = new ArrayList<>();
private final ScheduledExecutorService scheduler = Executors.newSingleThreadScheduledExecutor();

@PostConstruct
public void initHealthCheckTask() {
    // 每30秒刷新一次健康实例列表
    scheduler.scheduleAtFixedRate(this::refreshHealthyInstances, 0, 30, TimeUnit.SECONDS);
}

private void refreshHealthyInstances() {
    List<ServiceInstance> allInstances = Arrays.asList(
        new DefaultServiceInstance(serviceId + "1", serviceId, "localhost", 8080, false),
        new DefaultServiceInstance(serviceId + "2", serviceId, "localhost", 9092, false),
        new DefaultServiceInstance(serviceId + "3", serviceId, "localhost", 9999, false)
    );
    
    List<ServiceInstance> tempHealthy = new ArrayList<>();
    for (ServiceInstance instance : allInstances) {
        try {
            String healthUrl = String.format("http://%s:%s/actuator/health", instance.getHost(), instance.getPort());
            webClient.get().uri(healthUrl).retrieve().toBodilessEntity().block();
            tempHealthy.add(instance);
        } catch (Exception ignored) {
            // 实例不可用,跳过
        }
    }
    this.healthyInstances = tempHealthy;
}

@Override
public Flux<List<ServiceInstance>> get() {
    return Flux.just(healthyInstances);
}

四、兜底处理:避免全节点故障导致的500

当所有实例都故障时,可添加降级逻辑,返回兜底实例或友好提示:

@Override
public Flux<List<ServiceInstance>> get() {
    if (healthyInstances.isEmpty()) {
        // 返回降级实例(需提前部署降级服务)
        return Flux.just(Collections.singletonList(
            new DefaultServiceInstance(serviceId + "-fallback", serviceId, "localhost", 8081, false)
        ));
    }
    return Flux.just(healthyInstances);
}

内容的提问来源于stack exchange,提问作者Mario Alejandro Ortiz Vargas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 12:50:23