Prometheus告警按severity拆分至不同Slack频道配置失效问题咨询
Absolutely! You can route alerts to separate Slack channels based on their severity label—this is a standard use case fully supported by Alertmanager. Let’s troubleshoot why your modified config isn’t working, then walk through a correct, tested setup.
Common Issues in Your Modified Config
First, let’s cover the most likely reasons alerts aren’t reaching your Slack channels:
- Label Mismatch: Alertmanager uses exact, case-sensitive label matching. If your alert rules set a label like
Severity(capital S) but your config looks forseverity(lowercase), no matches will occur. - Invalid Routing Structure: A misconfigured route tree (e.g., missing
continueflags, incorrect parent/child hierarchy) can cause alerts to fall through the cracks. - Slack Config Errors: Wrong webhook URL, typos in channel names, or missing workspace permissions for the webhook to post to target channels will block alerts.
- Syntax Errors: A missing indent, YAML typo, or invalid template variable will cause Alertmanager to fail loading the config entirely.
Working Configuration Example
Here’s a complete, validated alertmanager.yml that routes critical and warning alerts to separate channels:
Step 1: Define Slack Receivers
Create two receivers—one for each channel. Ensure your Slack webhook URL has permission to post to the target channels (some workspaces require unique webhooks per channel):
global: resolve_timeout: 5m route: group_by: ['alertname'] group_wait: 30s # Wait 30s to group related alerts before sending group_interval: 5m # Wait 5m before sending a new batch for the same alert group repeat_interval: 1h # Repeat unresolved alerts every hour receiver: 'slack-warning' # Default receiver for unlabeled alerts routes: # Route critical alerts to the dedicated critical channel - match: severity: critical receiver: 'slack-critical' continue: false # Stop processing further routes once matched # Route warning alerts to the warning channel - match: severity: warning receiver: 'slack-warning' continue: false receivers: - name: 'slack-warning' slack_configs: - api_url: 'https://hooks.slack.com/services/YOUR_WEBHOOK_URL' channel: '#alerts-warning' text: '*⚠️ Warning Alert:*\nSummary: {{ .CommonAnnotations.summary }}\nDescription: {{ .CommonAnnotations.description }}' send_resolved: true # Send a message when the alert is resolved - name: 'slack-critical' slack_configs: - api_url: 'https://hooks.slack.com/services/YOUR_WEBHOOK_URL' channel: '#alerts-critical' text: '*🚨 Critical Alert:*\nSummary: {{ .CommonAnnotations.summary }}\nDescription: {{ .CommonAnnotations.description }}' send_resolved: true
Step 2: Verify Alert Rules Have Correct Labels
Ensure your Prometheus alert rules explicitly set the severity label (case must match your Alertmanager config):
# Example Prometheus alert rule file groups: - name: node_alerts rules: - alert: HighCPUUsage expr: 100 - (avg by(instance) (irate(node_cpu_seconds_total{mode="idle"}[1m])) * 100) > 90 labels: severity: critical # Matches the Alertmanager match rule annotations: summary: "High CPU usage on {{ $labels.instance }}" description: "CPU usage has been above 90% for 1 minute." - alert: LowDiskSpace expr: (node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}) * 100 < 20 labels: severity: warning # Matches the Alertmanager match rule annotations: summary: "Low disk space on {{ $labels.instance }}" description: "Available disk space is below 20%."
Troubleshooting Steps
If alerts still aren’t coming through:
- Validate Config Syntax: Run this command to check for YAML errors:
Fix any reported errors before restarting Alertmanager.alertmanager --config.check-config --config.file=alertmanager.yml - Check Alertmanager Logs: Look for errors related to config loading or Slack API calls. For systemd-managed services:
You’ll see explicit messages likejournalctl -u alertmanager -ffailed to send slack notificationif the webhook is invalid or the channel doesn’t exist. - Use Alertmanager UI: Access the Alertmanager UI (default port 9093) to verify alerts are being received and routed correctly. The "Alerts" tab shows which receiver each alert is assigned to.
- Test with a Mock Alert: Send a test alert directly to Alertmanager to validate routing:
Check if this appears in your critical Slack channel.curl -XPOST -H "Content-Type: application/json" -d '[ { "labels": { "alertname": "TestCriticalAlert", "severity": "critical" }, "annotations": { "summary": "Test Critical Alert", "description": "This is a test critical alert." } } ]' http://your-alertmanager-ip:9093/api/v2/alerts
内容的提问来源于stack exchange,提问作者Sharmi

