You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Prometheus-Alertmanager Slack告警内容不完整问题求助

Slack告警内容重启后缺失问题求助

在Docker Compose环境中部署了Prometheus/Alertmanager(版本0.26.0),当前核心问题是Slack告警内容无法稳定完整展示:首次触发告警时能显示全部数据,但执行docker-compose restart后,切换至已知故障实例触发告警时,部分告警数据缺失。已确认告警格式无问题,且通过--log.level=debug开启了调试日志,查阅GitHub相关问题后仍未找到解决方案,现求助如何确保每次Slack通知都能完整展示告警信息。

相关截图

  • Slack告警消息:
    Slack告警消息
  • Prometheus界面:
    Prometheus界面
  • Alertmanager UI:
    Alertmanager UI
  • Alertmanager调试日志:
    Alertmanager调试日志

配置文件

prometheus.yml

global:
  scrape_interval: 10s
  evaluation_interval: 5s

scrape_configs:

  - job_name: My Hosts
    metrics_path: /probe
    params:
      module: [icmp]
    file_sd_configs:
      - files:
        - '/etc/prometheus/targets.yml'  
    relabel_configs:
      - source_labels: [__address__]
        target_label: __param_target
      - source_labels: [__param_target]
        target_label: instance
      - target_label: __address__
        replacement: blackbox:9115

rule_files:
  - "/etc/prometheus/rules/my_rules.yml"

alerting:
  alertmanagers:
    - static_configs:
        - targets:
          - 'alertmanager:9093'

my_targets.yml

- targets: ['10.113.204.3']   
        labels:
          instance_name: '44-base-ws03.dddemo3.com'
          cluster_name: '44-Test'

my_rules.yml

groups:
  - name: Instancess
    rules:
    - alert: InstanceDown
      expr: probe_success == 0
      for: 10s
      labels:
        severity: CRITICAL
        instance: "{{ $labels.instance }}"
        job: "{{ $labels.job }}"
        instance_name: "{{ $labels.instance_name }}"
        cluster_name: "{{ $labels.cluster_name }}"
      annotations:
        description: '{{ $labels.instance }} has been down for more than 5 minutes.'
        summary: 'Instance: {{ $labels.instance }} is down'

alertmanager.yml

global:
  slack_api_url: 'my_slack_url'
route:
  receiver: 'slack-notifications'
  group_by: ['alertname']
  group_interval: 1m
  group_wait: 20s
  repeat_interval: 2m

   
receivers:
- name: 'slack-notifications'
  slack_configs:
  - channel: '#se_goes_alerts_dev'
    send_resolved: true
    title: ':bullhorn-slack: {{ .CommonLabels.severity | toUpper }}:                           {{ .CommonLabels.alertname }} {{ .CommonLabels.instance }} - {{ .CommonLabels.job }}'
    text: |
      Description:  {{ .CommonAnnotations.description }}
      Summary:  {{ .CommonAnnotations.summary }}
      Instance Name: {{ .CommonLabels.instance_name }}
      cluster_name: {{ .CommonLabels.cluster_name }}
      instance: {{ .CommonLabels.instance }}

docker-compose.yml

version: '3'
services:
  prometheus:
    image: prom/prometheus:latest
    container_name: prometheus
    volumes:
      - ./prometheus:/etc/prometheus
    ports:
      - '9090:9090'
    command:
      - '--config.file=/etc/prometheus/prometheus.yml'
      - '--web.enable-remote-write-receiver'
    networks:
      - my-network
      
  alertmanager:
    image: prom/alertmanager:latest
    container_name: alertmanager
    command:
      - '--config.file=/etc/alertmanager/alertmanager.yml'
      - '--storage.path=/alertmanager'
      - '--log.level=debug'
    ports:
      - '9093:9093'
    volumes:
      - ./alertmanager/alertmanager.yml:/etc/alertmanager/alertmanager.yml
      - ./alertmanager:/etc/alertmanager
    networks:
      - my-network
    depends_on:
      - prometheus

解决方案建议

1. 修复Alertmanager存储持久化

当前Alertmanager的/alertmanager目录未做持久化挂载,重启容器会丢失告警状态数据,可能导致标签匹配异常。修改docker-compose.yml添加持久化卷:

alertmanager:
  # 保留原有配置,新增存储卷挂载
  volumes:
    - ./alertmanager/alertmanager.yml:/etc/alertmanager/alertmanager.yml
    - alertmanager-storage:/alertmanager
    - ./alertmanager:/etc/alertmanager

# 在services同级添加卷定义
volumes:
  alertmanager-storage:

2. 调整告警分组逻辑

当前group_by: ['alertname']会将同类型告警合并分组,重启后旧状态未清理可能覆盖新告警标签。建议:

  • 增加分组维度:group_by: ['alertname', 'instance'],确保每个实例告警独立分组
  • 调整group_wait: 0s让告警立即发送,避免分组延迟导致的标签丢失

3. 删除规则中冗余标签定义

my_rules.yml里手动复制的instance、job等标签属于冗余定义,可能引发标签冲突,直接删除:

labels:
  severity: CRITICAL
  # 移除以下四行冗余定义
  # instance: "{{ $labels.instance }}"
  # job: "{{ $labels.job }}"
  # instance_name: "{{ $labels.instance_name }}"
  # cluster_name: "{{ $labels.cluster_name }}"

4. 验证标签传递链路

重启后检查Prometheus的Targets页面,确认instance_name、cluster_name等标签是否正常加载;同时查看Alertmanager调试日志,确认告警事件中是否包含这些标签,定位缺失环节。

内容的提问来源于stack exchange,提问作者Ronald

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 21:14:53