kube-state-metrics返回值及Prometheus告警规则配置咨询
Let’s walk through your questions and troubleshoot why your alert isn’t firing—this is a super common gotcha with kube-state-metrics!
1. Is your alert rule syntax correct?
Your rule IF kube_pod_container_status_running{pod="pod-name",container="container_name"} > 0 is syntactically valid, but there’s a critical catch depending on what you’re trying to alert on:
- If you want to alert when the container is running, this rule will match, but you’ll only get an alert if that metric exists and has a value >0.
- If your goal is to alert when the container is NOT running, this rule is backwards—you’d need to use
== 0instead (and ensure the metric still exists for the pod/container, which it should as long as the pod hasn’t been deleted).
Also, double-check your label matching:
- Is the pod name exactly
"pod-name"(case-sensitive)? - Did you miss a
namespacelabel? Most Kubernetes clusters have pods in specific namespaces, so addingnamespace="your-namespace"to your label set might be necessary to match the right pod. - Verify the container name is spelled correctly (kube-state-metrics uses the exact container name from the pod spec).
2. Does kube_pod_container_status_running return a boolean or numeric value?
Prometheus doesn’t have native boolean types—all "boolean-like" metrics are numeric gauges. For kube_pod_container_status_running:
- It returns
1when the container is in a running state. - It returns
0when the container is not running (e.g., pending, terminated, crashed).
3. What does kube-state-metrics expose?
kube-state-metrics is a service that scrapes Kubernetes API objects and exposes their state as Prometheus metrics. It covers almost all core Kubernetes resources, with metrics focused on status, availability, and lifecycle events. Here are some key categories:
- Pod metrics: Track phases (
kube_pod_status_phase), container readiness (kube_pod_container_status_ready), restart counts (kube_pod_restarts_total), and termination status. - Deployment metrics: Monitor available replicas (
kube_deployment_status_replicas_available), unavailable replicas (kube_deployment_status_replicas_unavailable), and rollout progress. - Node metrics: Report node readiness (
kube_node_status_ready), allocatable resources (kube_node_allocatable_cpu_cores), and taint status. - Service/Ingress metrics: Track endpoint readiness and service selector matches.
- Persistent Volume metrics: Monitor PV/PVC binding status and usage.
All metrics follow a consistent pattern: they use labels to identify specific resources (e.g., pod, namespace, deployment) and numeric values to represent state or counts.
Quick Troubleshooting Tips for Your Non-Firing Alert
If your rule is syntactically correct but not firing:
- Query the metric directly in Prometheus’s Graph UI: Run
kube_pod_container_status_running{pod="pod-name",container="container_name"}to confirm if any samples exist. If no results come back, your label filters are wrong or kube-state-metrics isn’t scraping that pod. - Check Prometheus’s
Targetspage to ensure kube-state-metrics is being scraped successfully (look for a healthy target with no errors). - Verify your alert rule’s evaluation interval: If your
evaluation_intervalis set to 5m, it might take up to 5 minutes for the rule to trigger after the condition is met.
内容的提问来源于stack exchange,提问作者RV186

