You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

升级容器版本后Prometheus告警一直处于pending状态无法触发firing

问题:Prometheus告警始终处于pending状态无法触发firing

更新内容

  1. 使用另一个exporter(fly-exporter)时,告警可以正常触发。在另一台主机上用近似配置运行Prometheus、Alertmanager和gcp-exporter,告警也能正常触发。推测gcp-exporter的抓取会周期性失败,数据缺失导致告警被重置。
  2. 发现ALERTS合成时间序列的图表存在间隙,通过API查询min_over_time(ALERTS{alertname="gcp_cloud_run_services_running"}[12h])结果为1。
  3. activeAt(Active Since)会周期性重置,暂未找到原因。告警对应的时间序列中未发现0值。

我为Google Cloud服务编写了gcp-exporter,用于在资源消耗超出预期时发送告警(邮件/Pushover)。该exporter已正常运行数月,此前告警功能正常。

近期仅进行了以下版本升级操作:

  • 容器镜像版本:
    • docker.io/prom/prometheus:v2.37.0
    • docker.io/prom/alertmanager:v0.24.0
  • google.golang.org/api依赖库

但近几周不再收到预期告警,原因是告警进入“pending”状态后,永远不会变为“firing”状态。问题出在Prometheus告警无法触发“firing”,而非AlertManager的问题。

已确认两个受监控服务(Cloud Functions、Cloud Run)的资源已运行超过1天:

date --rfc-3339=seconds --utc
2022-07-22 22:03:37+00:00

gcloud functions list \
--project=${PROJECT} \
--format="value(updateTime)"
2020-12-03T20:23:30.940Z # 2020年12月3日

gcloud run services list --project=${PROJECT} \
--format="value(status.conditions['lastTransitionTime'].slice(0))"
2022-07-21T16:21:24.293561Z # 2022年7月21日
2022-07-21T16:20:48.439194Z # 2022年7月21日

调用/api/v1/alerts接口返回结果:

{
  "status": "success",
  "data": {
    "alerts": [
      {
        "labels": {
          "alertname": "gcp_cloud_functions_running",
          "severity": "page"
        },
        "annotations": {
          "summary": "GCP Cloud Functions running"
        },
        "state": "pending",
        "activeAt": "2022-07-22T21:56:02.984175081Z",
        "value": "1e+00"
      },
      {
        "labels": {
          "alertname": "gcp_cloud_run_services_running",
          "severity": "page"
        },
        "annotations": {
          "summary": "GCP Cloud Run services running"
        },
        "state": "pending",
        "activeAt": "2022-07-22T21:56:02.984175081Z",
        "value": "2e+00"
      }
    ]
  }
}

注意 撰写此问题时,activeAt值为21:56:02,处于6小时告警等待窗口内。但相关资源已存在更久,且Prometheus自2022年7月21日16:14重启后一直运行(当时认为它陷入僵死状态),为何activeAt仅从该时间点开始?

更新 告警的activeAt属性似乎每30分钟重置一次:

2022-07-23T01:26:02.984175081Z
2022-07-23T01:26:02.984175081Z

2022-07-23T00:56:02.984175081Z
2022-07-23T00:56:02.984175081Z

查询两个时间序列后,未发现可能重置6小时计时器的0值,且两个告警的activeAt值每次都相同,这一点令人困惑。

此外,查询时间序列的结果如下:

QUERY="..." # 见下文

# 当前时间为22:03,因此该时间范围涵盖22小时的数据
START="2022-07-01T00:00:00.000Z"
END="2022-07-22T23:59:59.999Z
STEP="5m"

# 结果见下文对应的QUERY
curl \
--silent \
--data-urlencode "query=${QUERY}" \
--data "start=${START}" \
--data "end=${END}" \
--data "step=${STEP}" \
"http://${HOST}:${PORT}/api/v1/query_range" \
| jq -r '.data.result[].values[][1]' \
| sort \
| uniq -c

# Cloud Functions
QUERY="min_over_time(gcp_cloud_functions_functions[15m]>0"
    348 1
     60 2

# Cloud Run
QUERY="min_over_time(gcp_cloud_run_services[15m])>0"
     60 10
    609 2

对上述结果的解读:两个告警的时间序列(按5分钟步长采样)从未包含0值,因此查询结果始终>0,但告警仍未触发。

问题

  1. 这是什么原因导致的?
  2. 我存在哪些理解误区?
  3. 有没有更好的调试方法?

配置文件

prometheus.yml

global:
  scrape_interval: 1m
  scrape_timeout: 10s
  evaluation_interval: 1m
alerting:
  alertmanagers:
  - follow_redirects: true
    enable_http2: true
    scheme: http
    timeout: 10s
    api_version: v2
    static_configs:
    - targets:
      - localhost:9093
rule_files:
- /etc/alertmanager/rules.yml
scrape_configs:
- job_name: gcp-exporter
  honor_timestamps: true
  scrape_interval: 15m
  scrape_timeout: 30s
  metrics_path: /metrics
  scheme: http
  follow_redirects: true
  enable_http2: true
  static_configs:
  - targets:
    - localhost:9402

rules.yml

groups:
- name: gcp_exporter
  rules:
  - alert: gcp_cloud_functions_running
    expr: min_over_time(gcp_cloud_functions_functions{}[15m]) > 0
    for: 6h
    labels:
      severity: page
    annotations:
      summary: GCP Cloud Functions running
  - alert: gcp_cloud_run_services_running
    expr: min_over_time(gcp_cloud_run_services{}[15m]) > 0
    for: 6h
    labels:
      severity: page
    annotations:
      summary: GCP Cloud Run services running

运行信息

运行时信息

{
  "status": "success",
  "data": {
    "startTime": "2022-07-21T16:14:23.941571056Z",
    "CWD": "/prometheus",
    "reloadConfigSuccess": true,
    "lastConfigTime": "2022-07-21T16:14:23Z",
    "corruptionCount": 0,
    "goroutineCount": 43,
    "GOMAXPROCS": 4,
    "GOGC": "",
    "GODEBUG": "",
    "storageRetention": "15d"
  }
}

构建信息

{
  "status": "success",
  "data": {
    "version": "2.37.0",
    "revision": "b41e0750abf5cc18d8233161560731de05199330",
    "branch": "HEAD",
    "buildUser": "root@0ebb6827e27f",
    "buildDate": "20220714-15:19:21",
    "goVersion": "go1.18.4"
  }
}

Alertmanagers信息

{
  "status": "success",
  "data": {
    "activeAlertmanagers": [
      {
        "url": "http://localhost:9093/api/v2/alerts"
      }
    ],
    "droppedAlertmanagers": []
  }
}

内容的提问来源于stack exchange,提问作者DazWilkin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 18:31:25