如何为ThreadPoolTaskExecutor在Grafana中添加连接与队列监控指标
我发现某仅从数据库查询实体的接口请求响应耗时3-5秒,每小时处理1.5k次请求,当前状况不佳。
接口定义如下:
@GetMapping @Async("operationThreadExecutor") public CompletableFuture<OperationDto> read()
我的线程池配置如下:
@Bean public AsyncTaskExecutor operationThreadExecutor() { ThreadPoolTaskExecutor executor = new ThreadPoolTaskExecutor(); executor.setCorePoolSize(5); executor.setMaxPoolSize(10); executor.setQueueCapacity(50); executor.setThreadNamePrefix("OperationSelfExecutor-"); executor.initialize(); }
请问如何在Grafana中添加该接口的当前连接数及线程池队列的监控指标?
要在Grafana中监控该接口的当前连接数和线程池队列状态,需先完成指标采集,再在Grafana中配置可视化面板,具体步骤如下:
一、采集所需指标(基于Spring Boot + Micrometer)
你的项目属于Spring Boot环境,推荐用Micrometer暴露线程池和HTTP连接的监控指标,结合Prometheus作为中间存储(Grafana对Prometheus支持最优):
1. 添加依赖
确保项目引入Micrometer核心包和Prometheus导出器:
<!-- Micrometer核心依赖 --> <dependency> <groupId>io.micrometer</groupId> <artifactId>micrometer-core</artifactId> </dependency> <!-- Prometheus指标导出器 --> <dependency> <groupId>io.micrometer</groupId> <artifactId>micrometer-registry-prometheus</artifactId> </dependency>
2. 暴露线程池指标
ThreadPoolTaskExecutor默认不会自动暴露全量监控指标,需要手动绑定Micrometer的ThreadPoolMetrics:
修改线程池配置Bean,添加指标绑定逻辑:
import io.micrometer.core.instrument.MeterRegistry; import io.micrometer.core.instrument.binder.jvm.ThreadPoolMetrics; @Bean public AsyncTaskExecutor operationThreadExecutor(MeterRegistry meterRegistry) { ThreadPoolTaskExecutor executor = new ThreadPoolTaskExecutor(); executor.setCorePoolSize(5); executor.setMaxPoolSize(10); executor.setQueueCapacity(50); executor.setThreadNamePrefix("OperationSelfExecutor-"); executor.initialize(); // 绑定线程池到Micrometer,指定指标标识为operation.self.executor ThreadPoolMetrics.monitor(meterRegistry, executor.getThreadPoolExecutor(), "operation.self.executor"); return executor; }
绑定后会生成以下关键线程池指标(Prometheus格式):
executor_threads_active{name="operation.self.executor"}:当前活跃线程数executor_queue_size{name="operation.self.executor"}:队列中等待的任务数executor_threads_max{name="operation.self.executor"}:线程池最大线程数
3. 暴露HTTP连接数指标
通过Spring Boot Actuator暴露容器(以Tomcat为例)的连接相关指标,在application.yml中配置:
management: endpoints: web: exposure: include: prometheus,health metrics: tags: application: your-app-name export: prometheus: enabled: true
配置后可获取的核心连接指标:
tomcat_sessions_active_current:当前活跃HTTP会话数(对应连接数)tomcat_threads_current:Tomcat工作线程当前数量http_server_requests_seconds_active{uri="/read"}:该接口的实时并发请求数
二、在Grafana中配置监控面板
假设已将Prometheus配置为Grafana的数据源,且Prometheus已添加Spring Boot的/actuator/prometheus端点作为采集目标:
1. 线程池队列监控面板
- 新建Stat或Graph面板,选择Prometheus数据源
- 输入查询语句:
- 队列等待任务数:
executor_queue_size{name="operation.self.executor"} - 活跃线程数:
executor_threads_active{name="operation.self.executor"}
- 队列等待任务数:
- 设置面板标题为「Operation线程池队列状态」,配置图例和显示样式(比如用折线图展示趋势)
2. 接口当前连接数监控面板
- 新建面板,输入查询语句:
- 接口实时并发请求数:
sum(http_server_requests_seconds_active{uri="/read"}) - Tomcat当前活跃连接数:
tomcat_sessions_active_current
- 接口实时并发请求数:
- 设置面板标题为「Read接口当前连接数」,选择数字或折线图展示
3. 额外优化:接口耗时监控
针对接口3-5秒的耗时问题,建议同时添加耗时分布监控,查询语句:avg(http_server_requests_seconds_bucket{uri="/read"}) by (le)
可直观看到请求耗时的区间分布,辅助定位瓶颈(比如数据库查询慢导致线程池队列堆积)
三、性能问题排查提示
从你的线程池配置(核心5、最大10、队列50)和请求量(每秒约0.4次)来看,理论上不应该出现高耗时,大概率是数据库查询本身效率低(比如缺少索引、查询数据量过大)。建议先排查SQL执行计划,再结合线程池监控确认是否存在队列堆积的情况。
内容的提问来源于stack exchange,提问作者Khilarian

