You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWS EKS集群kubelet位置及container_cpu_usage_seconds_total指标获取咨询

Got it, let's break down your questions one by one since this is a common point of confusion with EKS and kubelet metrics:

Why don't you see kubelet Pods in EKS?

Unlike some other Kubernetes distributions, in EKS (running on EC2 nodes), kubelet runs as a systemd service on each underlying EC2 instance—not as a containerized Pod. That's why you won't find it in your kubectl get pods -n kube-system output; it's not hidden, it's just running directly on the node's operating system.

To verify it's running, you can connect to an EC2 node (via SSH or EC2 Instance Connect) and run:

systemctl status kubelet

You'll see it listed as an active system service, responsible for managing Pods on that node and collecting container metrics via cAdvisor.

Can you access kubelet's metrics (including container_cpu_usage_seconds_total)?

Absolutely. Kubelet exposes two key metrics endpoints that include the data you're missing:

  • /metrics: Contains kubelet's own operational metrics
  • /metrics/cadvisor: Contains container-level metrics collected by cAdvisor, including container_cpu_usage_seconds_total, memory usage, disk I/O, etc.

To test this directly, you can run a curl command against a node's private IP (ensure your security group allows access to port 10250):

curl -k https://<node-private-ip>:10250/metrics/cadvisor

The -k flag skips certificate validation (EKS uses self-signed certs for kubelet by default). You'll see the full set of cAdvisor metrics in the output.

How to collect container_cpu_usage_seconds_total in EKS?

Here are the most reliable methods to gather these metrics at scale:

  • Self-hosted Prometheus

    1. Configure Prometheus to discover all EKS nodes via Kubernetes service discovery.
    2. Add a scrape target for port 10250, targeting the /metrics/cadvisor endpoint.
    3. Grant your Prometheus ServiceAccount the required permissions with a ClusterRole:
      apiVersion: rbac.authorization.k8s.io/v1
      kind: ClusterRole
      metadata:
        name: prometheus-kubelet-scrape
      rules:
      - apiGroups: [""]
        resources: ["nodes/proxy"]
        verbs: ["get"]
      
      Bind this role to your Prometheus ServiceAccount to allow it to access kubelet endpoints across the cluster.
  • AWS Distro for OpenTelemetry (ADOT)
    ADOT has pre-built EKS integrations that automatically discover nodes and scrape kubelet metrics. You can deploy the ADOT Collector via Helm or the AWS Console, and configure it to send metrics to Amazon CloudWatch, Amazon Managed Service for Prometheus, or other destinations.

  • Amazon Managed Service for Prometheus (AMP)
    If you're using AMP, set up a scrape configuration that leverages EKS service discovery to target kubelet endpoints. AMP handles the scaling and maintenance of the Prometheus backend, so you just need to define the scrape rules.

Note: The aws-node Pods (VPC CNI) don't expose these container metrics—they're focused on networking, not container monitoring. That's why you won't find container_cpu_usage_seconds_total in their metrics output.

内容的提问来源于stack exchange,提问作者Joost den Boer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:23:24