升级容器版本后Prometheus告警一直处于pending状态无法触发firing
问题:Prometheus告警始终处于pending状态无法触发firing
更新内容
- 使用另一个exporter(
fly-exporter)时,告警可以正常触发。在另一台主机上用近似配置运行Prometheus、Alertmanager和gcp-exporter,告警也能正常触发。推测gcp-exporter的抓取会周期性失败,数据缺失导致告警被重置。- 发现
ALERTS合成时间序列的图表存在间隙,通过API查询min_over_time(ALERTS{alertname="gcp_cloud_run_services_running"}[12h])结果为1。activeAt(Active Since)会周期性重置,暂未找到原因。告警对应的时间序列中未发现0值。
我为Google Cloud服务编写了gcp-exporter,用于在资源消耗超出预期时发送告警(邮件/Pushover)。该exporter已正常运行数月,此前告警功能正常。
近期仅进行了以下版本升级操作:
- 容器镜像版本:
docker.io/prom/prometheus:v2.37.0docker.io/prom/alertmanager:v0.24.0
google.golang.org/api依赖库
但近几周不再收到预期告警,原因是告警进入“pending”状态后,永远不会变为“firing”状态。问题出在Prometheus告警无法触发“firing”,而非AlertManager的问题。
已确认两个受监控服务(Cloud Functions、Cloud Run)的资源已运行超过1天:
date --rfc-3339=seconds --utc 2022-07-22 22:03:37+00:00 gcloud functions list \ --project=${PROJECT} \ --format="value(updateTime)" 2020-12-03T20:23:30.940Z # 2020年12月3日 gcloud run services list --project=${PROJECT} \ --format="value(status.conditions['lastTransitionTime'].slice(0))" 2022-07-21T16:21:24.293561Z # 2022年7月21日 2022-07-21T16:20:48.439194Z # 2022年7月21日
调用/api/v1/alerts接口返回结果:
{ "status": "success", "data": { "alerts": [ { "labels": { "alertname": "gcp_cloud_functions_running", "severity": "page" }, "annotations": { "summary": "GCP Cloud Functions running" }, "state": "pending", "activeAt": "2022-07-22T21:56:02.984175081Z", "value": "1e+00" }, { "labels": { "alertname": "gcp_cloud_run_services_running", "severity": "page" }, "annotations": { "summary": "GCP Cloud Run services running" }, "state": "pending", "activeAt": "2022-07-22T21:56:02.984175081Z", "value": "2e+00" } ] } }
注意 撰写此问题时,
activeAt值为21:56:02,处于6小时告警等待窗口内。但相关资源已存在更久,且Prometheus自2022年7月21日16:14重启后一直运行(当时认为它陷入僵死状态),为何activeAt仅从该时间点开始?
更新 告警的
activeAt属性似乎每30分钟重置一次:2022-07-23T01:26:02.984175081Z 2022-07-23T01:26:02.984175081Z 2022-07-23T00:56:02.984175081Z 2022-07-23T00:56:02.984175081Z查询两个时间序列后,未发现可能重置6小时计时器的0值,且两个告警的
activeAt值每次都相同,这一点令人困惑。
此外,查询时间序列的结果如下:
QUERY="..." # 见下文 # 当前时间为22:03,因此该时间范围涵盖22小时的数据 START="2022-07-01T00:00:00.000Z" END="2022-07-22T23:59:59.999Z STEP="5m" # 结果见下文对应的QUERY curl \ --silent \ --data-urlencode "query=${QUERY}" \ --data "start=${START}" \ --data "end=${END}" \ --data "step=${STEP}" \ "http://${HOST}:${PORT}/api/v1/query_range" \ | jq -r '.data.result[].values[][1]' \ | sort \ | uniq -c # Cloud Functions QUERY="min_over_time(gcp_cloud_functions_functions[15m]>0" 348 1 60 2 # Cloud Run QUERY="min_over_time(gcp_cloud_run_services[15m])>0" 60 10 609 2
对上述结果的解读:两个告警的时间序列(按5分钟步长采样)从未包含0值,因此查询结果始终>0,但告警仍未触发。
问题
- 这是什么原因导致的?
- 我存在哪些理解误区?
- 有没有更好的调试方法?
配置文件
prometheus.yml
global: scrape_interval: 1m scrape_timeout: 10s evaluation_interval: 1m alerting: alertmanagers: - follow_redirects: true enable_http2: true scheme: http timeout: 10s api_version: v2 static_configs: - targets: - localhost:9093 rule_files: - /etc/alertmanager/rules.yml scrape_configs: - job_name: gcp-exporter honor_timestamps: true scrape_interval: 15m scrape_timeout: 30s metrics_path: /metrics scheme: http follow_redirects: true enable_http2: true static_configs: - targets: - localhost:9402
rules.yml
groups: - name: gcp_exporter rules: - alert: gcp_cloud_functions_running expr: min_over_time(gcp_cloud_functions_functions{}[15m]) > 0 for: 6h labels: severity: page annotations: summary: GCP Cloud Functions running - alert: gcp_cloud_run_services_running expr: min_over_time(gcp_cloud_run_services{}[15m]) > 0 for: 6h labels: severity: page annotations: summary: GCP Cloud Run services running
运行信息
运行时信息
{ "status": "success", "data": { "startTime": "2022-07-21T16:14:23.941571056Z", "CWD": "/prometheus", "reloadConfigSuccess": true, "lastConfigTime": "2022-07-21T16:14:23Z", "corruptionCount": 0, "goroutineCount": 43, "GOMAXPROCS": 4, "GOGC": "", "GODEBUG": "", "storageRetention": "15d" } }
构建信息
{ "status": "success", "data": { "version": "2.37.0", "revision": "b41e0750abf5cc18d8233161560731de05199330", "branch": "HEAD", "buildUser": "root@0ebb6827e27f", "buildDate": "20220714-15:19:21", "goVersion": "go1.18.4" } }
Alertmanagers信息
{ "status": "success", "data": { "activeAlertmanagers": [ { "url": "http://localhost:9093/api/v2/alerts" } ], "droppedAlertmanagers": [] } }
内容的提问来源于stack exchange,提问作者DazWilkin
相关产品推荐
相关产品推荐

