You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在K8s Sidecar容器中采集主容器的CPU/内存使用率(1s粒度)?

Absolutely, you can build a custom sidecar container to collect your main container's CPU/memory metrics with 1-second granularity—no need for a cluster-wide cAdvisor deployment that wastes resources. Here are the most practical approaches and tools to make this work:

1. Read cgroup Files Directly (Lightest Option)

Kubernetes exposes container resource usage data through the cgroup filesystem on the node. Since your sidecar runs in the same pod as the main container, you can mount the node's /sys/fs/cgroup directory into the sidecar to access the main container's metrics directly.

How to implement this:

  • Mount the cgroup filesystem: Add a hostPath volume to your pod spec that mounts /sys/fs/cgroup from the node into the sidecar.
  • Get the main container's ID: Use Kubernetes' Downward API to inject the main container's ID into the sidecar as an environment variable. This lets you locate the exact cgroup directory for the main container.
  • Write a simple collection script: Use a bash or Python script to read metrics from the cgroup files at 1-second intervals, then send the data to your target endpoint.

Example Pod Spec Snippet:

spec:
  containers:
  - name: main-app
    image: your-main-app-image:latest
  - name: metrics-sidecar
    image: python:3.11-slim
    command: ["python", "/scripts/collect-metrics.py"]
    env:
    - name: MAIN_CONTAINER_ID
      valueFrom:
        fieldRef:
          fieldPath: status.containerStatuses[0].containerID
    - name: POD_UID
      valueFrom:
        fieldRef:
          fieldPath: metadata.uid
    - name: METRICS_ENDPOINT
      value: "https://your-target-endpoint.com/metrics"
    volumeMounts:
    - name: cgroup
      mountPath: /sys/fs/cgroup
    - name: metrics-script
      mountPath: /scripts
  volumes:
  - name: cgroup
    hostPath:
      path: /sys/fs/cgroup
  - name: metrics-script
    configMap:
      name: metrics-collection-script

Example Python Collection Script:

import os
import time
import requests

# Extract container ID from the full URI (e.g., docker://abc123 -> abc123)
main_container_id = os.environ["MAIN_CONTAINER_ID"].split("://")[-1]
# Build cgroup path (adjust based on your K8s QoS class: besteffort, guaranteed, etc.)
cgroup_base_path = f"/sys/fs/cgroup/memory/kubepods.slice/kubepods-besteffort.slice/kubepods-besteffort-pod-{os.environ['POD_UID']}.slice/{main_container_id}"
endpoint = os.environ["METRICS_ENDPOINT"]

# Track previous CPU usage to calculate utilization
prev_cpu_total = 0
prev_time = time.time()

while True:
    # Collect memory usage
    with open(f"{cgroup_base_path}/memory.usage_in_bytes", "r") as f:
        memory_usage = int(f.read().strip())
    
    # Collect CPU usage (calculate delta to get utilization)
    with open(f"{cgroup_base_path}/cpuacct.stat", "r") as f:
        cpu_lines = f.readlines()
        user_cpu = int(cpu_lines[0].split()[1])
        system_cpu = int(cpu_lines[1].split()[1])
        current_cpu_total = user_cpu + system_cpu
    
    current_time = time.time()
    cpu_delta = current_cpu_total - prev_cpu_total
    time_delta = current_time - prev_time
    # Convert to CPU cores (1 CPU tick = 10ms on most systems)
    cpu_utilization = (cpu_delta / time_delta) / 100
    
    # Send metrics to endpoint
    requests.post(
        endpoint,
        json={
            "pod_name": os.environ.get("POD_NAME"),
            "memory_usage_bytes": memory_usage,
            "cpu_utilization_cores": round(cpu_utilization, 2)
        }
    )
    
    prev_cpu_total = current_cpu_total
    prev_time = current_time
    time.sleep(1)

Notes:

  • The cgroup path may vary depending on your Kubernetes QoS class (e.g., kubepods-guaranteed.slice instead of kubepods-besteffort.slice). You can add logic to auto-detect this if needed.
  • Ensure your sidecar has read access to the cgroup files. In GKE, the default service account should have sufficient permissions, but you may need to adjust the pod's securityContext if you run into permission issues.

2. Use Lightweight Metric Collection Tools

If you don't want to write a custom script, there are lightweight tools you can repurpose as sidecars:

  • Custom cAdvisor Build: Fork the official cAdvisor repo, modify the hardcoded 10-second collection interval to 1 second, then build a custom image. Deploy this only in pods where you need metrics—this avoids cluster-wide overhead.
  • Prometheus Node Exporter: Configure it to only scrape the main container's cgroup metrics (using command-line flags to filter paths) and set the scrape interval to 1 second. Mount /sys/fs/cgroup into the exporter sidecar to access the data.
  • Go-Based Lightweight Collectors: Use libraries like github.com/opencontainers/runc/libcontainer/cgroups to build a tiny Go binary that reads cgroup data and sends it to your endpoint. Compile it into a scratch-based image for minimal resource footprint.

3. Key Considerations

  • Resource Overhead: All these approaches have negligible resource usage—scripts or lightweight binaries will use far less CPU/memory than a full cAdvisor instance.
  • Security: Mounting /sys/fs/cgroup gives the sidecar access to node-level cgroup data, but since it's only deployed in specific pods, the risk is limited. Avoid running the sidecar as root unless absolutely necessary.
  • GKE Specifics: GKE's managed cAdvisor can't have its interval changed, but your custom sidecar will work seamlessly alongside it without conflicts.

内容的提问来源于stack exchange,提问作者Jan Matas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 09:16:37