AlertmanagerConfig未向邮件接收器发送告警问题排查
我部署的一个Pod持续处于CrashLoopBackoff状态,已配置对应告警,但告警仅触发AlertManager自带的默认接收器,未发送到自定义邮件接收器。该AlertManager属于bitnami/kube-prometheus stack。
自定义邮件接收器配置(AlertmanagerConfig)
apiVersion: monitoring.coreos.com/v1alpha1 kind: AlertmanagerConfig metadata: name: pod-restarts-receiver namespace: monitoring labels: alertmanagerConfig: email release: prometheus spec: route: receiver: 'email-receiver' groupBy: ['alertname'] groupWait: 30s groupInterval: 5m repeatInterval: 5m matchers: - name: job value: pod-restarts receivers: - name: 'email-receiver' emailConfigs: - to: 'etshuma@mycompany.com' sendResolved: true from: 'ops@mycompany.com' smarthost: 'mail2.mycompany.com:25'
告警规则配置(PrometheusRule)
apiVersion: monitoring.coreos.com/v1 kind: PrometheusRule metadata: name: pod-restarts-alert namespace: monitoring labels: app: kube-prometheus-stack release: prometheus spec: groups: - name: api rules: - alert: PodRestartsAlert expr: sum by (namespace, pod) (kube_pod_container_status_restarts_total{namespace="labs", pod="crash-loop-pod"}) > 5 for: 1m labels: severity: critical job: pod-restarts annotations: summary: "Pod {{ $labels.pod }} in namespace {{ $labels.namespace }} has more than 5 restarts" description: "The pod {{ $labels.pod }} in namespace {{ $labels.namespace }} has experienced more than 5 restarts."
AlertManager Pod内默认配置提取
提取命令:
kubectl -n monitoring exec -it alertmanager-prometheus-kube-prometheus-alertmanager-0 -- sh cd conf cat config.yaml
默认config.yaml内容:
route: group_by: ['alertname'] group_wait: 30s group_interval: 5m repeat_interval: 1h receiver: 'web.hook' receivers: - name: 'web.hook' webhook_configs: - url: 'http://127.0.0.1:5001/' inhibit_rules: - source_match: severity: 'critical' target_match: severity: 'warning' equal: ['alertname', 'dev', 'instance']
AlertManager UI显示的全局配置
global: resolve_timeout: 5m http_config: follow_redirects: true enable_http2: true smtp_hello: localhost smtp_require_tls: true pagerduty_url: https://events.pagerduty.com/v2/enqueue opsgenie_api_url: https://api.opsgenie.com/ wechat_api_url: https://qyapi.weixin.qq.com/cgi-bin/ victorops_api_url: https://alert.victorops.com/integrations/generic/20131114/alert/ telegram_api_url: https://api.telegram.org webex_api_url: https://webexapis.com/v1/messages route: receiver: "null" group_by: - job continue: false routes: - receiver: monitoring/pod-restarts-receiver/email-receiver group_by: - alertname match: job: pod-restarts matchers: - namespace="monitoring" continue: true group_wait: 30s group_interval: 5m repeat_interval: 5m - receiver: "null" match: alertname: Watchdog continue: false group_wait: 30s group_interval: 5m repeat_interval: 12h receivers: - name: "null" - name: monitoring/pod-restarts-receiver/email-receiver email_configs: - send_resolved: true to: etshuma@mycompany.com from: ops@mycompany.com hello: localhost smarthost: mail2.mycompany.com:25 headers: From: ops@mycompany.com Subject: '{{ template "email.default.subject" . }}' To: etshuma@mycompany.com html: '{{ template "email.default.html" . }}' require_tls: true templates: []
疑问解答
1. 全局配置中接收器显示为"null",原因是什么?
这是kube-prometheus-stack的默认行为:当未配置全局默认接收器时,系统会自动生成一个名为null的空接收器作为路由的兜底项。你通过AlertmanagerConfig配置的是局部路由规则,未覆盖全局默认接收器,所以全局路由默认指向null。
2. 全局配置顶部无邮件设置,是否会导致问题?
不会。你在AlertmanagerConfig的emailConfigs中已经指定了smarthost、from等必要参数,这些局部配置会被合并到AlertManager的最终配置中,且优先级高于全局SMTP设置,全局SMTP配置属于可选项。
3. 不确定AlertManagerConfig层级定义的邮件设置是否生效,也不清楚如何更新仅能从Pod访问的全局配置;部署所用values.yaml无smarthost或邮件设置选项
- 从UI显示的配置来看,你的邮件设置已经生效:
monitoring/pod-restarts-receiver/email-receiver接收器包含完整的邮件配置项。 - 若要更新全局配置,bitnami的kube-prometheus-stack可通过
alertmanager.config字段在values.yaml中添加全局SMTP设置,示例:
alertmanager: config: global: smtp_smarthost: 'mail2.mycompany.com:25' smtp_from: 'ops@mycompany.com'
添加后执行helm upgrade即可生效。
4. 全局配置中存在额外匹配器- namespace="monitoring",是否需要在PrometheusRule中添加类似命名空间标签?
这是告警无法触发邮件接收器的核心原因!你的告警规则监控的是labs命名空间的Pod,但路由规则被自动添加了namespace="monitoring"的匹配条件,导致告警无法匹配该路由。这个额外匹配器是kube-prometheus-stack自动添加的——因为你的AlertmanagerConfig部署在monitoring命名空间,系统默认会给跨命名空间的配置添加命名空间匹配器。
解决方法:在AlertmanagerConfig的route中显式指定匹配目标Pod所在的labs命名空间:
spec: route: receiver: 'email-receiver' groupBy: ['alertname'] groupWait: 30s groupInterval: 5m repeatInterval: 5m matchers: - name: job value: pod-restarts - name: namespace value: labs
如果需要匹配多个命名空间,也可以添加否定匹配排除monitoring:
matchers: - name: job value: pod-restarts - name: namespace value: "!monitoring"
5. AlertManagerConfig是否必须与PrometheusRule及目标Pod处于同一命名空间?
不需要。AlertmanagerConfig可以部署在任意命名空间,只要带有正确的标签(比如你的release: prometheus)让AlertManager能够发现它。但要注意kube-prometheus-stack的自动匹配器逻辑,会给AlertmanagerConfig自动添加所在命名空间的匹配条件,需要手动调整或覆盖。
关于路由树编辑器无法可视化的问题
路由树编辑器需要完整的AlertManager配置文本,直接将UI导出的完整全局配置复制进去即可正常可视化。之前无法显示可能是因为粘贴了部分配置,或者原配置存在格式问题(比如多余空行),清理格式后重新尝试即可。
内容的提问来源于stack exchange,提问作者Golide

