You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

配置Prometheus告警监控特定端点4xx/5xx错误及本地部署排查

问题排查与解决方案

一、先解决Prometheus无法访问本地API的问题

1. 修复Docker网络访问问题

Prometheus跑在Docker容器里时,配置里的127.0.0.1指的是容器自己的localhost,不是宿主机的API地址。两种解决办法:

  • 用宿主机专属地址替换:Windows/macOS用host.docker.internal,Linux用Docker网桥IP(一般是172.17.0.1),把scrape目标改成:
    static_configs:
      - targets: ['host.docker.internal:8081'] # Windows/macOS环境
      # 或者 Linux环境用:
      # - targets: ['172.17.0.1:8081']
    
  • 让Prometheus用宿主机网络启动:启动容器时加--network host参数,这样容器内的127.0.0.1就指向宿主机。

2. 修正Scrape配置的错误

你把metrics_path设成了/api/users,但Prometheus默认是拉取目标的/metrics端点获取监控指标的——除非你的API专门在/api/users暴露了Prometheus格式的metrics,否则这个配置完全错误。先改回默认配置,确保能拿到up指标:

scrape_configs:
  - job_name: 'my-api'
    scrape_interval: 5s
    # 删掉错误的metrics_path和无用的params
    static_configs:
      - targets: ['host.docker.internal:8081'] # 替换成正确的访问地址

3. 验证访问是否正常

进入Prometheus容器测试:

# 进入容器
docker exec -it <你的Prometheus容器名> sh
# 测试访问API的metrics端点(如果有的话)
curl http://host.docker.internal:8081/metrics
# 或者直接测/api/users
curl http://host.docker.internal:8081/api/users

改完后打开Prometheus UI(默认http://localhost:9090),查询up{job="my-api"},如果返回1,那up == 1的告警10秒后就会触发。

二、配置特定端点4xx/5xx的告警

要监控API端点的HTTP状态码,得用Blackbox Exporter主动探测,步骤如下:

1. 用Docker部署Blackbox Exporter

docker run -d --name blackbox_exporter -p 9115:9115 prom/blackbox-exporter:latest

2. 更新Prometheus配置,添加Blackbox探测任务

修改prometheus.yml的scrape_configs部分:

scrape_configs:
  # 保留原有监控API自身metrics的任务(如果需要)
  - job_name: 'my-api'
    scrape_interval: 5s
    static_configs:
      - targets: ['host.docker.internal:8081']

  # 新增Blackbox探测任务,专门监控/api/users端点
  - job_name: 'blackbox-api-users'
    scrape_interval: 5s
    metrics_path: /probe
    params:
      module: [http_2xx] # 用http_2xx模块探测,会返回http_status_code指标
    static_configs:
      - targets: ['http://host.docker.internal:8081/api/users']
    relabel_configs:
      - source_labels: [__address__]
        target_label: __param_target
      - source_labels: [__param_target]
        target_label: instance
      - target_label: __address__
        replacement: blackbox_exporter:9115 # Blackbox容器的地址(同一Docker网络下直接用容器名)

3. 编写4xx/5xx的告警规则

更新alert.rules:

groups:
- name: example
  rules:
  # 保留原有测试告警
  - alert: test
    expr: up == 1
    for: 10s
    labels:
      severity: warning
    annotations:
      summary: "站点正常: {{$labels.instance}}"
      description: "站点 {{$labels.instance}} 已正常运行超过10秒"

  # 新增端点错误告警
  - alert: APIEndpointError
    expr: probe_http_status_code{job="blackbox-api-users"} >= 400
    for: 10s
    labels:
      severity: critical
    annotations:
      summary: "API端点异常: {{$labels.instance}}"
      description: "API端点 {{$labels.instance}} 返回错误状态码 {{$value}},持续时间超过10秒"

4. 重启Prometheus容器让配置生效

docker restart <你的Prometheus容器名>

三、验证告警

  1. 打开Prometheus UI的Alerts页面,当API端点返回4xx/5xx时,APIEndpointError会先进入Pending状态,10秒后变为Firing。
  2. 确保Alertmanager和Prometheus在同一Docker网络,或者把alerting配置里的target改成host.docker.internal:9093,保证告警能发到Alertmanager。

内容的提问来源于stack exchange,提问作者Gonçalo C

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 04:55:35