You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWS Kubernetes集群Pod异常状态邮件告警配置需求

Got it, let’s walk through exactly how to set up those critical alerts for non-running Pods and containers using your existing cAdvisor + Prometheus + Alertmanager stack on AWS. I’ve broken this down into actionable steps that play nicely with your current setup:

Step 1: Create a Prometheus Alert Rule for Non-Running Pods

First, we’ll define an alert rule that targets Pods in problematic states. This relies on kube-state-metrics (a standard add-on for Kubernetes metrics—if you don’t have it deployed yet, grab it via Helm or official manifests first).

Create a file named pod-status-alerts.yaml with this rule:

groups:
- name: pod-health-alerts
  rules:
  - alert: PodNotRunning
    expr: |
      # Trigger if Pod phase isn't Running
      kube_pod_status_phase{phase!="Running"} == 1
      OR
      # Trigger if container is in CrashLoopBackOff or Error waiting state
      kube_pod_container_status_waiting_reason{reason=~"CrashLoopBackOff|Error"} == 1
    for: 1m  # Wait 1 minute to avoid false positives from temporary restarts
    labels:
      severity: critical
    annotations:
      summary: "Pod {{ $labels.pod }} ({{ $labels.namespace }}) is not running"
      description: "Pod {{ $labels.pod }} in namespace {{ $labels.namespace }} is in state: {{ $labels.phase | default $labels.reason }}. Check pod logs or events for root cause."

Apply this rule to your Prometheus instance (typically via a ConfigMap or Prometheus Operator’s PrometheusRule resource):

kubectl apply -f pod-status-alerts.yaml -n monitoring
Step 2: Configure Alertmanager to Send Email Alerts

Next, update your Alertmanager config to route these alerts to your email. Create or modify your Alertmanager ConfigMap (alertmanager-config.yaml):

global:
  smtp_smarthost: 'smtp.your-email-provider.com:587'  # e.g., smtp.gmail.com:587 for Gmail
  smtp_from: 'k8s-cluster-alerts@your-domain.com'
  smtp_auth_username: 'k8s-cluster-alerts@your-domain.com'
  smtp_auth_password: 'your-app-specific-password'  # Use an app password, not your main email password
  smtp_require_tls: true

route:
  group_by: ['alertname']
  group_wait: 30s  # Wait 30s to group related alerts
  group_interval: 5m  # Wait 5m before sending a new group of alerts
  repeat_interval: 4h  # Repeat alerts every 4 hours until resolved
  receiver: 'pod-alert-email'

receivers:
- name: 'pod-alert-email'
  email_configs:
  - to: 'your-alert-recipient@your-domain.com'
    subject: '[K8s Critical Alert] {{ .CommonAnnotations.summary }}'
    html: |
      <h2>Kubernetes Pod Alert Triggered</h2>
      <p><strong>Alert Summary:</strong> {{ .CommonAnnotations.summary }}</p>
      <p><strong>Details:</strong> {{ .CommonAnnotations.description }}</p>
      <ul>
        <li>Pod Name: {{ .CommonLabels.pod }}</li>
        <li>Namespace: {{ .CommonLabels.namespace }}</li>
        <li>Alert Started: {{ .StartsAt.Format "2006-01-02 15:04:05 UTC" }}</li>
      </ul>

Apply the config and restart Alertmanager to pick up changes:

kubectl apply -f alertmanager-config.yaml -n monitoring
kubectl rollout restart deployment alertmanager -n monitoring
Step 3: Verify the Setup

Let’s make sure everything works as expected:

  • Test with a faulty Pod: Run a Pod with an invalid image to trigger a CrashLoopBackOff:
    kubectl run test-faulty-pod --image=invalid-image-tag -n default
    
  • Check Prometheus Alerts: Open your Prometheus UI, navigate to the "Alerts" tab, and confirm the PodNotRunning alert shows as "Firing" after 1 minute.
  • Check Email: You should receive an alert email with the test Pod’s name and status within a few minutes.
  • Clean up: Delete the test Pod once you’ve verified:
    kubectl delete pod test-faulty-pod -n default
    
Extra Pro Tips
  • Exclude specific Pods: If you don’t want alerts for system Pods or temporary workloads, add a label (e.g., alert-exclude: true) to those Pods, then update the alert rule’s expr to exclude them:
    expr: |
      (kube_pod_status_phase{phase!="Running"} == 1 OR kube_pod_container_status_waiting_reason{reason=~"CrashLoopBackOff|Error"} == 1)
      AND
      kube_pod_metadata_labels{label="alert-exclude"} != "true"
    
  • Tune alert timing: Adjust the for duration in the alert rule based on your tolerance—shorter for time-sensitive workloads, longer to avoid flapping alerts.
  • Enrich alert details: Add more context to the email (like Pod labels or node info) by expanding the html template with additional {{ .CommonLabels }} fields.

内容的提问来源于stack exchange,提问作者shiv455

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:03:35