Prometheus未触发告警求助:Scraper已上报但无告警通知
Prometheus/Alertmanager告警无法触发排查与解决
问题背景
部署WebAPI应用后,通过Selenium执行多场景测试并向Prometheus输出监控指标,其中操作耗时类警告可正常触发,但关键操作超时/失败类告警始终无法触发。今日应用出现不可用情况,Scraper持续上报login_failed_alert:1、ShellTube_scenario_failed_alert:1达20分钟,却未收到任何告警通知。
Scraper日志片段
[2025-04-11 08:07:53.897 +02:00 INF] ===== Metrics Snapshot ===== [2025-04-11 08:07:53.897 +02:00 INF] tmLogin: 2.3422876 [2025-04-11 08:07:53.897 +02:00 INF] tmSTPrTree: -1 [2025-04-11 08:07:53.897 +02:00 INF] tmSTCalculate: -1 [2025-04-11 08:07:53.897 +02:00 INF] tmStGenPDF: 0 [2025-04-11 08:07:53.897 +02:00 INF] login_failed_alert: 0 [2025-04-11 08:07:53.897 +02:00 INF] ShellTube_scenario_failed_alert: 1 [2025-04-11 08:07:53.897 +02:00 INF] PDFCreation_failed_alert: 0 [2025-04-11 08:07:53.897 +02:00 INF] Application_unavailable_alert: 0 [2025-04-11 08:11:40.618 +02:00 INF] ===== Metrics Snapshot ===== [2025-04-11 08:11:40.618 +02:00 INF] tmLogin: -1 [2025-04-11 08:11:40.618 +02:00 INF] tmSTPrTree: -1 [2025-04-11 08:11:40.618 +02:00 INF] tmSTCalculate: -1 [2025-04-11 08:11:40.618 +02:00 INF] tmStGenPDF: 0 [2025-04-11 08:11:40.618 +02:00 INF] login_failed_alert: 1 [2025-04-11 08:11:40.618 +02:00 INF] ShellTube_scenario_failed_alert: 1 [2025-04-11 08:11:40.618 +02:00 INF] PDFCreation_failed_alert: 0 [2025-04-11 08:11:40.618 +02:00 INF] Application_unavailable_alert: 0
上述日志持续出现20分钟,期间无成功运行记录
告警规则配置(alert.rules.yml片段)
- alert: ShellTubeScenarioFailed expr: ShellTube_scenario_failed_alert == 1 for: 60s labels: severity: critical annotations: summary: "Shell Tube Scenario execution failed" description: "ShellTube_scenario_failed_alert is 1 – indicating scenario failure." - alert: LoginFailed expr: login_failed_alert == 1 for: 30s labels: severity: critical annotations: summary: "Login failed" description: "Login attempt failed." - alert: ApplicationUnavailable expr: Application_unavailable_alert == 1 for: 30s labels: severity: critical annotations: summary: "Application is unavailable" description: "The application did not respond"
Prometheus采集配置(prometheus.yml片段)
job_name: "cairo_monitoring" scrape_interval: 60s scrape_timeout: 30s static_configs: - targets: ["localhost:5001"]
排查步骤
- 确认指标采集状态:登录Prometheus UI,在表达式浏览器输入
ShellTube_scenario_failed_alert、login_failed_alert,检查是否存在数值为1的时间序列,确认Prometheus是否正确采集到指标。 - 检查告警规则状态:进入Prometheus的Alerts页面,查看对应告警规则的状态(Inactive/Pending/Firing)。若为Pending,需确认
for参数与采集间隔的匹配逻辑。 - 验证Alertmanager配置:
- 确认Prometheus配置中已正确指定Alertmanager地址(
alerting节点)。 - 检查Alertmanager的路由规则,确保
severity: critical的告警被路由到正确接收方。 - 查看Alertmanager日志,排查通知发送失败的错误(如邮件服务器配置、API密钥问题)。
- 确认Prometheus配置中已正确指定Alertmanager地址(
- 检查指标逻辑正确性:日志中
Application_unavailable_alert始终为0,但实际应用不可用,说明Scraper未在应用不可用时将该指标置1;需确认所有失败场景下对应告警指标的映射逻辑是否正确。 - 适配采集间隔与
for参数:当前采集间隔为60s,LoginFailed的for:30s,Prometheus需要至少两次连续采集满足条件才会触发告警,实际触发时间约为60s,需根据需求调整参数。
解决方案
修复指标采集逻辑:
- 当应用不可用时,将
Application_unavailable_alert设为1,确保直接反映应用状态。 - 验证所有失败场景(如登录超时、场景执行失败)下,对应告警指标能正确置1。
- 当应用不可用时,将
调整告警规则参数:
- 由于采集间隔为60s,将
LoginFailed、ApplicationUnavailable的for参数调整为60s或更长,确保Prometheus能收集到连续的异常样本触发告警。
- 由于采集间隔为60s,将
完善Alertmanager配置:
- 在Prometheus.yml中添加Alertmanager地址:
alerting: alertmanagers: - static_configs: - targets: ["localhost:9093"] # 替换为实际Alertmanager地址 - 配置Alertmanager的接收规则(以邮件为例):
route: group_by: ['alertname'] group_wait: 10s group_interval: 10s repeat_interval: 1h receiver: 'critical-alerts' receivers: - name: 'critical-alerts' email_configs: - to: 'your-alert-email@example.com' from: 'prometheus-alerts@example.com' smarthost: 'smtp.example.com:587' auth_username: 'smtp-username' auth_password: 'smtp-password'
- 在Prometheus.yml中添加Alertmanager地址:
排查日志错误:
- 查看Prometheus日志:
journalctl -u prometheus(Systemd环境) - 查看Alertmanager日志:
journalctl -u alertmanager,定位配置或运行错误。
- 查看Prometheus日志:
内容的提问来源于stack exchange,提问作者Kal800
相关产品推荐
相关产品推荐

