Prometheus告警自定义标签在PagerDuty中未插值,求排查建议
Prometheus告警注解标签未渲染问题排查
问题背景
为特定namespace下的Pod配置CPU使用率Prometheus告警规则,当使用率超过90%时触发告警并推送至PagerDuty,但告警注解中的{{ $labels.product }}、{{ $labels.severity }}未被替换为实际标签值。
告警规则配置
- alert: CpuUtilizationWarning expr: avg by (kubernetes_io_zone) (rate( container_cpu_usage_seconds_total{ pod=~"app.*", container="app", namespace="app"}[5m] )) / on() group_left() avg( kube_pod_container_resource_requests_cpu_cores{ pod=~"app.*", container="app", namespace="apps" } ) * 100 > 90 for: 5m labels: severity: warning service: location product: backend-app annotations: description: 'CPU utilization of {{ $labels.product }} pod is exceeding {{ $labels.severity }} and value is {{ $value }} %' summary: '{{ $labels.severity }} CPU utilization of {{ $labels.product }} pod'
实际告警内容
Labels: - alertname = CpuUtilizationWarning - monitor = prometheus - product = backend-app - service = location - severity = warning - kubernetes_io_zone = us-east-1c Annotations: - description = CPU utilization of pod is exceeding and value is 95.099865567118 % - summary = CPU utilization of pod
预期告警内容
Annotations: - description = CPU utilization of backend-app pod is exceeding warning and value is 95.099865567118 % - summary = warning CPU utilization of backend-app pod
可能的原因
注解模板引号格式问题:注解使用单引号包裹时,部分Prometheus/AlertManager版本无法正确解析内部的Go模板变量。尝试将单引号替换为双引号:
annotations: description: "CPU utilization of {{ $labels.product }} pod is exceeding {{ $labels.severity }} and value is {{ $value }} %" summary: "{{ $labels.severity }} CPU utilization of {{ $labels.product }} pod"AlertManager自定义模板未渲染注解:如果配置了自定义PagerDuty通知模板,可能模板直接输出注解原始字符串,未做Go模板解析。检查AlertManager模板配置,确保注解值经过模板渲染后再发送。
Prometheus静态标签传递异常:需确认Prometheus是否将告警规则中定义的静态标签正确附加到告警实例。在Prometheus的Alerts页面查看该告警实例,核实
product和severity标签是否存在。PagerDuty集成的内容过滤:部分PagerDuty集成方式会对注解内容做额外转义或过滤,导致模板变量被清除。检查AlertManager的PagerDuty配置,确认未开启相关过滤功能。
Prometheus版本兼容性bug:较旧的Prometheus版本可能存在静态标签无法在注解中渲染的问题,尝试升级到最新稳定版本验证。
内容的提问来源于stack exchange,提问作者srethomas
相关产品推荐
相关产品推荐

