Prometheus磁盘空间告警显示“In host: Unknown”问题排查
问题
配置了如下alert-rules.yml与Alertmanager配置,通过prom2teams将告警推送到MS Teams。内存告警正常显示“In host: node-exporter:9100”,但磁盘空间告警显示“In host: unknown”,请问这是什么原因?
告警规则配置
groups: - name: alert.rules rules: - alert: HostOutOfMemory expr: ((node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100) < 25 for: 5m labels: severity: warning annotations: summary: "Host out of memory (instance {{ $labels.instance }})" description: "Node memory is filling up (< 25% left)\n VALUE = {{ $value }}\n LABELS: {{ $labels }}" - alert: HostOutOfDiskSpace expr: (sum(node_filesystem_free_bytes) / sum(node_filesystem_size_bytes) * 100) < 30 for: 1s labels: severity: warning annotations: summary: "Host out of disk space (instance {{ $labels.instance }})" description: "Disk is almost full (< 30% left)\n VALUE = {{ $value }}\n LABELS: {{ $labels }}"
Alertmanager配置
route: receiver: 'teams' group_wait: 30s group_interval: 5m receivers: - name: 'teams' webhook_configs: - url: "http://prom2teams:8089" send_resolved: true
原因分析与解决方案
核心原因:磁盘告警的
expr使用了无维度限制的sum()聚合,直接把所有实例的磁盘数据求和后,instance标签被聚合消除,导致{{ $labels.instance }}无法获取到对应值,最终显示为unknown。而内存告警的表达式未使用聚合函数,完整保留了instance标签,因此能正常显示主机信息。修复方案:给磁盘指标的
sum()添加by (instance)维度约束,确保聚合后保留instance标签,修改后的磁盘告警规则如下:
- alert: HostOutOfDiskSpace expr: (sum(node_filesystem_free_bytes) by (instance) / sum(node_filesystem_size_bytes) by (instance) * 100) < 30 for: 1s labels: severity: warning annotations: summary: "Host out of disk space (instance {{ $labels.instance }})" description: "Disk is almost full (< 30% left)\n VALUE = {{ $value }}\n LABELS: {{ $labels }}"
- 额外优化建议:若你的磁盘指标带有
mountpoint标签(用于区分根分区、数据分区等),建议同时保留该标签,避免将不同挂载点的磁盘数据错误聚合,此时可将by (instance, mountpoint)添加到sum()后,让告警能精准定位到具体分区。
内容的提问来源于stack exchange,提问作者MsA
相关产品推荐
相关产品推荐

