You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Prometheus未触发告警求助:Scraper已上报但无告警通知

Prometheus/Alertmanager告警无法触发排查与解决

问题背景

部署WebAPI应用后,通过Selenium执行多场景测试并向Prometheus输出监控指标,其中操作耗时类警告可正常触发,但关键操作超时/失败类告警始终无法触发。今日应用出现不可用情况,Scraper持续上报login_failed_alert:1、ShellTube_scenario_failed_alert:1达20分钟,却未收到任何告警通知。

Scraper日志片段

[2025-04-11 08:07:53.897 +02:00  INF]  ===== Metrics Snapshot =====
[2025-04-11 08:07:53.897 +02:00  INF]  tmLogin: 2.3422876
[2025-04-11 08:07:53.897 +02:00  INF]  tmSTPrTree: -1
[2025-04-11 08:07:53.897 +02:00  INF]  tmSTCalculate: -1
[2025-04-11 08:07:53.897 +02:00  INF]  tmStGenPDF: 0
[2025-04-11 08:07:53.897 +02:00  INF]  login_failed_alert: 0
[2025-04-11 08:07:53.897 +02:00  INF]  ShellTube_scenario_failed_alert: 1
[2025-04-11 08:07:53.897 +02:00  INF]  PDFCreation_failed_alert: 0
[2025-04-11 08:07:53.897 +02:00  INF]  Application_unavailable_alert: 0


[2025-04-11 08:11:40.618 +02:00  INF]  ===== Metrics Snapshot =====
[2025-04-11 08:11:40.618 +02:00  INF]  tmLogin: -1
[2025-04-11 08:11:40.618 +02:00  INF]  tmSTPrTree: -1
[2025-04-11 08:11:40.618 +02:00  INF]  tmSTCalculate: -1
[2025-04-11 08:11:40.618 +02:00  INF]  tmStGenPDF: 0
[2025-04-11 08:11:40.618 +02:00  INF]  login_failed_alert: 1
[2025-04-11 08:11:40.618 +02:00  INF]  ShellTube_scenario_failed_alert: 1
[2025-04-11 08:11:40.618 +02:00  INF]  PDFCreation_failed_alert: 0
[2025-04-11 08:11:40.618 +02:00  INF]  Application_unavailable_alert: 0

上述日志持续出现20分钟,期间无成功运行记录

告警规则配置(alert.rules.yml片段)

- alert: ShellTubeScenarioFailed
  expr: ShellTube_scenario_failed_alert == 1
  for: 60s
  labels:
   severity: critical
  annotations:
   summary: "Shell Tube Scenario execution failed"
   description: "ShellTube_scenario_failed_alert is 1 – indicating scenario failure."

- alert: LoginFailed
  expr: login_failed_alert == 1
  for: 30s
  labels:
   severity: critical
  annotations:
   summary: "Login failed"
   description: "Login attempt failed."

- alert: ApplicationUnavailable
  expr: Application_unavailable_alert == 1
  for: 30s
  labels:
   severity: critical
  annotations:
   summary: "Application is unavailable"
   description: "The application did not respond"

Prometheus采集配置(prometheus.yml片段)

job_name: "cairo_monitoring"
scrape_interval: 60s
scrape_timeout: 30s
static_configs:
  - targets: ["localhost:5001"]

排查步骤

  • 确认指标采集状态:登录Prometheus UI,在表达式浏览器输入ShellTube_scenario_failed_alert、login_failed_alert,检查是否存在数值为1的时间序列,确认Prometheus是否正确采集到指标。
  • 检查告警规则状态:进入Prometheus的Alerts页面,查看对应告警规则的状态(Inactive/Pending/Firing)。若为Pending,需确认for参数与采集间隔的匹配逻辑。
  • 验证Alertmanager配置:
    1. 确认Prometheus配置中已正确指定Alertmanager地址(alerting节点)。
    2. 检查Alertmanager的路由规则,确保severity: critical的告警被路由到正确接收方。
    3. 查看Alertmanager日志,排查通知发送失败的错误(如邮件服务器配置、API密钥问题)。
  • 检查指标逻辑正确性:日志中Application_unavailable_alert始终为0,但实际应用不可用,说明Scraper未在应用不可用时将该指标置1;需确认所有失败场景下对应告警指标的映射逻辑是否正确。
  • 适配采集间隔与for参数:当前采集间隔为60s,LoginFailed的for:30s,Prometheus需要至少两次连续采集满足条件才会触发告警,实际触发时间约为60s,需根据需求调整参数。

解决方案

  1. 修复指标采集逻辑:

    • 当应用不可用时,将Application_unavailable_alert设为1,确保直接反映应用状态。
    • 验证所有失败场景(如登录超时、场景执行失败)下,对应告警指标能正确置1。
  2. 调整告警规则参数:

    • 由于采集间隔为60s,将LoginFailed、ApplicationUnavailable的for参数调整为60s或更长,确保Prometheus能收集到连续的异常样本触发告警。
  3. 完善Alertmanager配置:

    • 在Prometheus.yml中添加Alertmanager地址:
      alerting:
        alertmanagers:
        - static_configs:
          - targets: ["localhost:9093"] # 替换为实际Alertmanager地址
      
    • 配置Alertmanager的接收规则(以邮件为例):
      route:
        group_by: ['alertname']
        group_wait: 10s
        group_interval: 10s
        repeat_interval: 1h
        receiver: 'critical-alerts'
      receivers:
      - name: 'critical-alerts'
        email_configs:
        - to: 'your-alert-email@example.com'
          from: 'prometheus-alerts@example.com'
          smarthost: 'smtp.example.com:587'
          auth_username: 'smtp-username'
          auth_password: 'smtp-password'
      
  4. 排查日志错误:

    • 查看Prometheus日志:journalctl -u prometheus(Systemd环境)
    • 查看Alertmanager日志:journalctl -u alertmanager,定位配置或运行错误。

内容的提问来源于stack exchange,提问作者Kal800

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 10:25:54