SpringBoot3.2.5 RestClient偶发3秒请求延迟问题排查求助
偶发3秒延迟:SpringBoot3.2.5 RestClient调用K8s集群内API时的异常
问题概述
使用SpringBoot3.2.5的RestClient调用Kubernetes集群内其他Namespace的网关服务API时,偶发出现请求发出到接收方收到请求之间存在约3秒延迟的情况,该问题无法稳定复现。
相关配置与代码
RestClient配置类
import org.yu.feign.GatewayRestService; import io.micrometer.observation.ObservationRegistry; import jakarta.annotation.Resource; import org.apache.hc.client5.http.classic.HttpClient; import org.apache.hc.client5.http.config.RequestConfig; import org.apache.hc.client5.http.impl.classic.HttpClientBuilder; import org.apache.hc.client5.http.impl.io.PoolingHttpClientConnectionManager; import org.apache.hc.client5.http.socket.ConnectionSocketFactory; import org.apache.hc.client5.http.socket.PlainConnectionSocketFactory; import org.apache.hc.client5.http.ssl.SSLConnectionSocketFactory; import org.apache.hc.core5.http.URIScheme; import org.apache.hc.core5.http.config.RegistryBuilder; import org.apache.hc.core5.util.Timeout; import org.springframework.context.annotation.Bean; import org.springframework.context.annotation.Configuration; import org.springframework.http.client.HttpComponentsClientHttpRequestFactory; import org.springframework.web.client.RestClient; import org.springframework.web.client.support.RestClientAdapter; import org.springframework.web.service.invoker.HttpServiceProxyFactory; @Configuration public class RestClientConfig { @Resource private ObservationRegistry observationRegistry; public HttpClient myHttpClient() { PoolingHttpClientConnectionManager connectionManager = new PoolingHttpClientConnectionManager( RegistryBuilder.<ConnectionSocketFactory>create() .register(URIScheme.HTTP.id, PlainConnectionSocketFactory.getSocketFactory()) .register(URIScheme.HTTPS.id, SSLConnectionSocketFactory.getSystemSocketFactory()) .build()); connectionManager.setDefaultMaxPerRoute(50); connectionManager.setMaxTotal(100); RequestConfig requestConfig = RequestConfig.custom() .setResponseTimeout(Timeout.ofSeconds(300)) .build(); return HttpClientBuilder.create() .useSystemProperties() .setDefaultRequestConfig(requestConfig) .setConnectionManager(connectionManager) .build(); } @Bean("gatewayRestService") public GatewayRestService gatewayRestService() { RestClient restClient = RestClient.builder() .observationRegistry(observationRegistry) .requestFactory(new HttpComponentsClientHttpRequestFactory(this.myHttpClient())) .baseUrl("http://gateway.namespace.svc:8080/rest/") .build(); final RestClientAdapter adapter = RestClientAdapter.create(restClient); final HttpServiceProxyFactory factory = HttpServiceProxyFactory.builderFor(adapter).build(); return factory.createClient(GatewayRestService.class); } }
GatewayRestService接口
import org.springframework.web.bind.annotation.RequestBody; import org.springframework.web.service.annotation.GetExchange; public interface GatewayRestService { @GetExchange(value = "getData") String getData(@RequestBody RequestOfGetData request); }
API调用代码
@Resource private GatewayRestService gatewayRestService; final RequestOfGetData request = new RequestOfGetData(); String response = gatewayRestService.getData(request);
已排查情况
- 已配置HTTP连接池,
maxTotal设为100,defaultMaxPerRoute设为50,理论上并发请求不超过50时不会出现排队等待; - 延迟发生时,在途请求数不足10,排除连接池排队导致延迟的可能;
- 服务部署在Kubernetes集群,采用集群内DNS域名(
http://service.namespace.svc:port)通信,暂不确定是否与K8s相关; - 延迟发生时的Zipkin链路追踪显示:请求从调用方发出到接收方接收到请求的时间间隔约为3秒,接收方处理请求本身耗时正常。
排查方向与解决方案建议
1. 补全HttpComponents超时配置
当前仅配置了响应超时,缺少连接超时和连接请求超时设置,偶发延迟可能来自连接建立阶段:
RequestConfig requestConfig = RequestConfig.custom() .setResponseTimeout(Timeout.ofSeconds(300)) .setConnectTimeout(Timeout.ofMilliseconds(500)) // 连接超时 .setConnectionRequestTimeout(Timeout.ofMilliseconds(500)) // 连接池获取连接超时 .build();
2. 优化连接池空闲连接复用
未配置空闲连接验证与存活时间,可能存在空闲连接被K8s网络组件回收的情况,导致需要重建连接:
// 设置空闲连接验证时间,每次获取连接前验证是否可用 connectionManager.setValidateAfterInactivity(Timeout.ofSeconds(5)); // 设置连接最大存活时间,避免长期空闲连接 connectionManager.setConnectionTimeToLive(Timeout.ofMinutes(5));
3. 排查Kubernetes DNS解析问题
集群内服务调用依赖CoreDNS,偶发3秒延迟可能与DNS解析超时有关:
- 在Pod内执行
nslookup gateway.namespace.svc多次,观察是否存在解析超时; - 检查CoreDNS Pod的日志与监控,确认是否存在请求堆积或性能瓶颈;
- 考虑在Pod内配置本地DNS缓存(如dnsmasq),减少DNS解析依赖。
4. 排查Kubernetes网络层面问题
- 查看延迟发生时Pod所在Node的网卡监控,检查是否存在丢包、带宽瓶颈;
- 检查网关服务的Service Endpoints状态,确认是否存在Endpoint频繁切换;
- 临时改用Pod IP直接调用,排除Service转发层面的问题。
5. 细化日志定位耗时阶段
- 启用HttpComponents的DEBUG日志:配置
logging.level.org.apache.hc.client5.http=DEBUG,观察连接建立、DNS解析阶段的耗时; - 自定义
ClientHttpRequestInterceptor,在请求前后记录时间戳,精准定位延迟发生的阶段。
内容的提问来源于stack exchange,提问作者Renjun Yu
相关产品推荐
相关产品推荐

