AWS Kubernetes集群Pod异常状态邮件告警配置需求
Got it, let’s walk through exactly how to set up those critical alerts for non-running Pods and containers using your existing cAdvisor + Prometheus + Alertmanager stack on AWS. I’ve broken this down into actionable steps that play nicely with your current setup:
First, we’ll define an alert rule that targets Pods in problematic states. This relies on kube-state-metrics (a standard add-on for Kubernetes metrics—if you don’t have it deployed yet, grab it via Helm or official manifests first).
Create a file named pod-status-alerts.yaml with this rule:
groups: - name: pod-health-alerts rules: - alert: PodNotRunning expr: | # Trigger if Pod phase isn't Running kube_pod_status_phase{phase!="Running"} == 1 OR # Trigger if container is in CrashLoopBackOff or Error waiting state kube_pod_container_status_waiting_reason{reason=~"CrashLoopBackOff|Error"} == 1 for: 1m # Wait 1 minute to avoid false positives from temporary restarts labels: severity: critical annotations: summary: "Pod {{ $labels.pod }} ({{ $labels.namespace }}) is not running" description: "Pod {{ $labels.pod }} in namespace {{ $labels.namespace }} is in state: {{ $labels.phase | default $labels.reason }}. Check pod logs or events for root cause."
Apply this rule to your Prometheus instance (typically via a ConfigMap or Prometheus Operator’s PrometheusRule resource):
kubectl apply -f pod-status-alerts.yaml -n monitoring
Next, update your Alertmanager config to route these alerts to your email. Create or modify your Alertmanager ConfigMap (alertmanager-config.yaml):
global: smtp_smarthost: 'smtp.your-email-provider.com:587' # e.g., smtp.gmail.com:587 for Gmail smtp_from: 'k8s-cluster-alerts@your-domain.com' smtp_auth_username: 'k8s-cluster-alerts@your-domain.com' smtp_auth_password: 'your-app-specific-password' # Use an app password, not your main email password smtp_require_tls: true route: group_by: ['alertname'] group_wait: 30s # Wait 30s to group related alerts group_interval: 5m # Wait 5m before sending a new group of alerts repeat_interval: 4h # Repeat alerts every 4 hours until resolved receiver: 'pod-alert-email' receivers: - name: 'pod-alert-email' email_configs: - to: 'your-alert-recipient@your-domain.com' subject: '[K8s Critical Alert] {{ .CommonAnnotations.summary }}' html: | <h2>Kubernetes Pod Alert Triggered</h2> <p><strong>Alert Summary:</strong> {{ .CommonAnnotations.summary }}</p> <p><strong>Details:</strong> {{ .CommonAnnotations.description }}</p> <ul> <li>Pod Name: {{ .CommonLabels.pod }}</li> <li>Namespace: {{ .CommonLabels.namespace }}</li> <li>Alert Started: {{ .StartsAt.Format "2006-01-02 15:04:05 UTC" }}</li> </ul>
Apply the config and restart Alertmanager to pick up changes:
kubectl apply -f alertmanager-config.yaml -n monitoring kubectl rollout restart deployment alertmanager -n monitoring
Let’s make sure everything works as expected:
- Test with a faulty Pod: Run a Pod with an invalid image to trigger a CrashLoopBackOff:
kubectl run test-faulty-pod --image=invalid-image-tag -n default - Check Prometheus Alerts: Open your Prometheus UI, navigate to the "Alerts" tab, and confirm the
PodNotRunningalert shows as "Firing" after 1 minute. - Check Email: You should receive an alert email with the test Pod’s name and status within a few minutes.
- Clean up: Delete the test Pod once you’ve verified:
kubectl delete pod test-faulty-pod -n default
- Exclude specific Pods: If you don’t want alerts for system Pods or temporary workloads, add a label (e.g.,
alert-exclude: true) to those Pods, then update the alert rule’sexprto exclude them:expr: | (kube_pod_status_phase{phase!="Running"} == 1 OR kube_pod_container_status_waiting_reason{reason=~"CrashLoopBackOff|Error"} == 1) AND kube_pod_metadata_labels{label="alert-exclude"} != "true" - Tune alert timing: Adjust the
forduration in the alert rule based on your tolerance—shorter for time-sensitive workloads, longer to avoid flapping alerts. - Enrich alert details: Add more context to the email (like Pod labels or node info) by expanding the
htmltemplate with additional{{ .CommonLabels }}fields.
内容的提问来源于stack exchange,提问作者shiv455

