如何通过Prometheus指标以CPU核心数表示Pod的CPU使用率?
Hey there! I’ve dealt with this exact issue before when trying to map Pod CPU usage directly to core counts, so let me share the most reliable solutions I’ve used:
1. 用container_cpu_usage_seconds_total计算实时核心使用量
This is the go-to metric for this use case, and it’s already collected by default if you have cadvisor running (which is part of the kubelet in most Kubernetes clusters).
- 原理:
container_cpu_usage_seconds_totaltracks the cumulative CPU time (in seconds) a container has used since startup. When you use therate()function to calculate the average CPU time consumed per second over a recent window (like 1 minute), the result directly translates to the number of CPU cores being used. For example, a value of 0.7 means the Pod is using 70% of a single core, or splitting usage across multiple cores to total 0.7 cores. - 基础查询(单容器):
rate(container_cpu_usage_seconds_total{namespace="your-namespace", pod=~"your-pod-name-pattern"}[1m]) - 聚合Pod内所有容器的总CPU:
If your Pod has multiple containers, usesum()to get the total core usage across the entire Pod:sum(rate(container_cpu_usage_seconds_total{namespace="your-namespace", pod=~"your-pod-name-pattern"}[1m])) by (pod)
2. (可选)对比Requests/Limits做使用率分析
Since you already have CPU requests and limits as core counts, you can combine them with the usage metric to calculate how close you are to your allocated resources:
sum(rate(container_cpu_usage_seconds_total{namespace="your-namespace", pod=~"your-pod-name-pattern"}[1m])) by (pod) / on(pod) group_left(kube_pod_container_resource_requests_cpu_cores) kube_pod_container_resource_requests_cpu_cores{namespace="your-namespace", pod=~"your-pod-name-pattern"}
This will give you a ratio of used CPU to requested CPU (e.g., 0.5 means you’re using half your requested cores).
3. 确保指标能被采集到
If you don’t see container_cpu_usage_seconds_total in your Prometheus instance, double-check these:
- Verify cadvisor is running: It’s usually embedded in the kubelet, so check if the kubelet exposes the
/metrics/cadvisorendpoint on your nodes. - Confirm your Prometheus scrape config includes the kubelet’s cadvisor metrics. A typical scrape job for this looks like:
- job_name: 'kubelet-cadvisor' kubernetes_sd_configs: - role: node scheme: https tls_config: ca_file: /var/run/secrets/kubernetes.io/serviceaccount/ca.crt bearer_token_file: /var/run/secrets/kubernetes.io/serviceaccount/token relabel_configs: - action: labelmap regex: __meta_kubernetes_node_label_(.+) - target_label: __address__ replacement: kubernetes.default.svc:443 - source_labels: [__meta_kubernetes_node_name] regex: (.+) target_label: __metrics_path__ replacement: /api/v1/nodes/${1}/proxy/metrics/cadvisor
内容的提问来源于stack exchange,提问作者Ramana

