You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Prometheus告警规则中$labels.instance解析与配置疑问

解决Prometheus告警中$labels.instance不解析的问题

问题根源

你的告警规则里混用了错误的模板语法:Prometheus的告警注释仅识别{{ 变量 }}格式的模板,直接写$labels.instance会被当作普通字符串输出,所以才显示字面量而非实际实例名称。

$labels.instance的来源

这个变量不需要手动设置,是Prometheus从采集目标的指标标签中自动获取的:

  • 在prometheus.yml的scrape_configs里配置targets(比如localhost:9100、192.168.1.100:9615)后,Prometheus采集指标时会自动给指标加上instance标签,值就是目标的地址
  • 例如up指标的格式为up{instance="localhost:9100", job="node_exporter"} 1,这里的instance标签值就是localhost:9100

修正后的rules.yml示例

groups:
  - name: alert_rules
    rules:
      - alert: InstanceDown
        expr: up == 0
        for: 5m
        labels:
          severity: critical
        annotations:
          summary: "Instance {{ $labels.instance }} down"  # 修正模板语法
          description: "[{{ $labels.instance }}] of job [{{ $labels.job }}] has been down for more than 5 minutes."  # 同步时间描述与for字段

      - alert: HostHighCpuLoad
        expr: 100 - (avg by(instance)(rate(node_cpu_seconds_total{mode="idle"}[2m])) * 100) > 80
        for: 0m
        labels:
          severity: warning
        annotations:
          summary: Host high CPU load (instance {{ $labels.instance }})  # 替换硬编码内容为变量
          description: "CPU load is > 80%
  VALUE = {{ $value }}
  LABELS: {{ $labels }}"

验证步骤

  1. 替换修正后的规则文件,重启Prometheus服务:sudo systemctl restart prometheus
  2. 在Prometheus UI的Alerts页面查看告警,确认变量是否正常解析
  3. 可手动停掉一个采集目标(如node_exporter),等待5分钟触发InstanceDown告警,检查通知内容

内容的提问来源于stack exchange,提问作者user15492347

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 03:24:27