You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Quarkus响应式REST端点@Bulkhead指标返回NaN问题咨询

响应式@Bulkhead端点的ft_bulkhead_executionsRunning/executionsWaiting指标返回NaN

我有一个标注了@Bulkhead注解的响应式REST端点,在通过curl获取Prometheus指标时,ft_bulkhead_executionsRunning和ft_bulkhead_executionsWaiting两个指标返回NaN值,不确定这是框架BUG还是自身配置问题。

curl请求指标的输出(过滤ft_bulkhead相关结果):

curl -v http://localhost:8080/ncudr-gud-dr/v1/directory-data/q/metrics | grep ft_bulkhead_
 % Total    % Received % Xferd  Average Speed   Time    Time     Time  Current
                                 Dload  Upload   Total   Spent    Left  Speed
  0     0    0     0    0     0      0      0 --:--:-- --:--:-- --:--:--     0*   Trying 127.0.0.1:8080...
* Connected to localhost (127.0.0.1) port 8080 (#0)
> GET /ncudr-gud-dr/v1/directory-data/q/metrics HTTP/1.1
> Host: localhost:8080
> User-Agent: curl/8.0.1
> Accept: */*
> 
< HTTP/1.1 200 OK
< Content-Type: application/openmetrics-text; version=1.0.0; charset=utf-8
< content-length: 65979
< 
{ [65979 bytes data]
# TYPE ft_bulkhead_calls counter
# HELP ft_bulkhead_calls  
ft_bulkhead_calls_total{bulkheadResult="rejected",method="org.cheva.ngud.SearchResource.query"} 120696.0
ft_bulkhead_calls_total{bulkheadResult="accepted",method="org.cheva.ngud.SearchResource.query"} 210840.0
# TYPE ft_bulkhead_waitingDuration_seconds summary
# HELP ft_bulkhead_waitingDuration_seconds  
ft_bulkhead_waitingDuration_seconds_count{method="org.cheva.ngud.SearchResource.query"} 210840.0
ft_bulkhead_waitingDuration_seconds_sum{method="org.cheva.ngud.SearchResource.query"} 7047.583918116
# TYPE ft_bulkhead_waitingDuration_seconds_max gauge
# HELP ft_bulkhead_waitingDuration_seconds_max  
ft_bulkhead_waitingDuration_seconds_max{method="org.cheva.ngud.SearchResource.query"} 0.363759962
# TYPE ft_bulkhead_executionsRunning gauge
# HELP ft_bulkhead_executionsRunning  
ft_bulkhead_executionsRunning{method="org.cheva.ngud.SearchResource.query"} NaN
# TYPE ft_bulkhead_executionsWaiting gauge
# HELP ft_bulkhead_executionsWaiting  
ft_bulkhead_executionsWaiting{method="org.cheva.ngud.SearchResource.query"} NaN
# TYPE ft_bulkhead_runningDuration_seconds summary
# HELP ft_bulkhead_runningDuration_seconds  
ft_bulkhead_runningDuration_seconds_count{method="org.cheva.ngud.SearchResource.query"} 210840.0
ft_bulkhead_runningDuration_seconds_sum{method="org.cheva.ngud.SearchResource.query"} 21331.330171846
# TYPE ft_bulkhead_runningDuration_seconds_max gauge
# HELP ft_bulkhead_runningDuration_seconds_max  
ft_bulkhead_runningDuration_seconds_max{method="org.cheva.ngud.SearchResource.query"} 0.578839078

可能的原因及解决步骤

1. 响应式Bulkhead的指标实现限制

多数熔断框架(如Resilience4j)中,响应式@Bulkhead默认使用信号量模式,而运行中/等待中的执行数指标在该模式下未被正确支持——信号量仅通过许可数控制并发,不追踪实时活跃执行数,导致指标采集时返回NaN。

解决:

  • 切换到固定线程池模式的响应式Bulkhead。以Resilience4j为例,修改配置文件:
    resilience4j.bulkhead:
      instances:
        yourBulkheadName:
          type: THREAD_POOL
          maxThreadPoolSize: 10
          coreThreadPoolSize: 5
          queueCapacity: 20
    

2. 指标初始化时机问题

即使已有大量请求,若指标采集逻辑在Bulkhead实例完全初始化前触发,也会返回NaN。

解决:

  • 触发一批新请求到端点,强制Bulkhead完成初始化后重新采集指标。
  • 确认框架的指标自动配置已开启,保证Bulkhead指标正确绑定到MeterRegistry。

3. 版本兼容性问题

熔断框架(如Resilience4j)与Spring Boot/Actuator版本不兼容,导致响应式Bulkhead的指标逻辑存在BUG。

解决:

  • 升级到框架最新稳定版,或匹配Spring Boot版本的兼容分支(如Spring Boot 3.x对应Resilience4j 2.x,Spring Boot 2.x对应Resilience4j 1.x)。
  • 查看框架官方Issue,确认是否有同类NaN指标问题已被修复。

4. 自定义指标采集逻辑错误

若自定义了指标采集器,可能在获取Bulkhead状态时处理不当导致NaN。

解决:

  • 检查自定义MeterBinder或采集代码,确保正确调用响应式Bulkhead的getNumberOfRunningCalls()和getNumberOfWaitingCalls()方法。
  • 注意:仅线程池模式的Bulkhead能返回有效运行/等待数,信号量模式下该指标未被支持。

内容的提问来源于stack exchange,提问作者Cheva

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 16:53:17