Prometheus默认缺失http_request_duration_seconds_count指标的配置求助
解决Prometheus无法采集
http_request_duration_seconds_*指标的问题 这些指标属于HTTP请求时长的Histogram类型指标,通常由应用的Prometheus客户端生成,找不到的核心原因要么是应用未生成/暴露指标,要么是Prometheus未正确配置抓取规则,以下是分步排查和配置方案:
1. 确认应用已生成并暴露指标
http_request_duration_seconds_count/sum/bucket这类指标需要应用通过Prometheus客户端库生成并暴露,不同技术栈的配置方式如下:
- Java/Spring Boot:
引入依赖(Maven为例):
在<dependency> <groupId>org.springframework.boot</groupId> <artifactId>spring-boot-starter-actuator</artifactId> </dependency> <dependency> <groupId>io.micrometer</groupId> <artifactId>micrometer-registry-prometheus</artifactId> </dependency>application.properties中开启Prometheus端点:management.endpoints.web.exposure.include=prometheus management.metrics.tags.application=your-app-name - Go:
使用promhttp库暴露指标,确保默认HTTP采集器被注册:import ( "net/http" "github.com/prometheus/client_golang/prometheus/promhttp" ) func main() { // 注册默认的HTTP请求时长指标 http.Handle("/metrics", promhttp.Handler()) http.ListenAndServe(":8080", nil) } - Python:
使用prometheus-client库,若框架未自动集成,需手动定义Histogram指标:from prometheus_client import start_http_server, Histogram import time REQUEST_DURATION = Histogram('http_request_duration_seconds', 'HTTP request duration in seconds') @REQUEST_DURATION.time() def handle_request(): time.sleep(0.1) if __name__ == '__main__': start_http_server(8000) while True: handle_request()
验证:直接访问应用的/metrics端点(如http://your-app-ip:port/metrics),搜索是否存在目标指标。若没有,优先解决应用端的指标生成问题。
2. 配置Prometheus抓取目标
确保Prometheus配置文件(通常为prometheus.yml)中添加了目标命名空间下的应用作为抓取源:
Kubernetes环境(自动发现Pod)
scrape_configs: - job_name: 'namespace-apps-monitor' kubernetes_sd_configs: - role: pod namespaces: names: ['your-target-namespace'] # 替换为你的目标命名空间 relabel_configs: # 仅抓取带有Prometheus注解的Pod - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape] action: keep regex: true # 从注解中获取指标端口 - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_port] action: replace target_label: __address__ regex: ([^:]+)(?::\d+)?;(\d+) replacement: $1:$2 # 添加命名空间标签,方便后续过滤 - source_labels: [__meta_kubernetes_namespace] action: replace target_label: kubernetes_namespace
静态目标(非Kubernetes环境)
scrape_configs: - job_name: 'namespace-apps-monitor' static_configs: - targets: ['app-1:8080', 'app-2:8080'] # 替换为你的应用地址和端口 labels: kubernetes_namespace: 'your-target-namespace' # 标记目标命名空间
配置完成后重启Prometheus,在http://prometheus-ip:9090/targets页面检查该job的抓取状态是否为UP。
3. 验证Prometheus采集状态
在Prometheus表达式浏览器(http://prometheus-ip:9090/graph)中执行查询:
http_request_duration_seconds_count{kubernetes_namespace="your-target-namespace"}
若返回结果,说明采集成功;若仍无结果,排查以下点:
- 应用
/metrics端点是否能被Prometheus所在网络访问(如Kubernetes网络策略是否限制) - 检查
scrape_interval配置,等待一个采集周期后重试 - 确认指标名称是否被自定义(如Spring Boot中修改了
management.metrics.web.server.request.metric-name配置,指标名会变更)
4. 在Grafana中使用指标
采集成功后,可在Grafana中创建面板实现API监控:
- QPS统计:
rate(http_request_duration_seconds_count[1m]) - 平均响应时间:
rate(http_request_duration_seconds_sum[1m]) / rate(http_request_duration_seconds_count[1m]) - 95分位响应时间:
histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket[1m])) by (le, kubernetes_namespace, endpoint))
内容的提问来源于stack exchange,提问作者ziv ziv
相关产品推荐
相关产品推荐

