You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

K8s中Java Pod物理机与容器内CPU使用率差异问题排查

问题描述

K8s集群中Java Pod的CPU显示满载,但业务运行正常。
执行kubectl top pods时,目标Pod的显示如下:

xxx-857c847496-6bkj5                           819m         6523Mi
xxx-857c847496-7r2l5                           945m         6356Mi

进入容器执行top命令时,显示CPU已满载:

PID USER      PR  NI    VIRT    RES    SHR S  %CPU  %MEM     TIME+ COMMAND                                                                                                             
  7 root      20   0   26.2g   4.3g  16592 S  125%   3.4 169:22.07 java                                                                                                                

物理机上看到的Pod CPU使用率与容器内显示差异极大,请问这是什么原因?

对应的Deployment配置如下:

---
apiVersion: apps/v1
kind: Deployment
metadata:
  annotations: {}
  labels:
    app: xxx
    version: v1
  name: xxx
  namespace: admin
spec:
  progressDeadlineSeconds: 600
  replicas: 6
  revisionHistoryLimit: 10
  selector:
    matchLabels:
      app: xxx
      version: v1
  strategy:
    rollingUpdate:
      maxSurge: 50%
      maxUnavailable: 50%
    type: RollingUpdate
  template:
    metadata:
      annotations:
        kubectl.kubernetes.io/restartedAt: '2022-10-04T10:15:14+08:00'
      creationTimestamp: null
      labels:
        app: xxx
        version: v1
    spec:
      affinity:
        podAntiAffinity:
          requiredDuringSchedulingIgnoredDuringExecution:
            - labelSelector:
                matchLabels:
                  app: xxx
              topologyKey: kubernetes.io/hostname
      containers:
        - env:
            - name: JAVA_ARGS
              value: >-
                -Xms1G -Xmx4G -XX:NewRatio=1 -XX:+UseContainerSupport
                -XX:MaxRAMPercentage=75.0
                -Dfile.encoding=UTF-8
                -Dlog4j2.formatMsgNoLookups=true
          image: '192.168.10.42/xxx/xxx:03f1006'
          imagePullPolicy: Always
          livenessProbe:
            failureThreshold: 10
            httpGet:
              path: /actuator/health/liveness
              port: 53387
              scheme: HTTP
            initialDelaySeconds: 30
            periodSeconds: 10
            successThreshold: 1
            timeoutSeconds: 2
          name: xxx
          ports:
            - containerPort: 47038
              protocol: TCP
            - containerPort: 53387
              protocol: TCP
          readinessProbe:
            failureThreshold: 10
            httpGet:
              path: /actuator/health/readiness
              port: 53387
              scheme: HTTP
            initialDelaySeconds: 30
            periodSeconds: 10
            successThreshold: 1
            timeoutSeconds: 2
          resources:
            limits:
              cpu: '8'
              memory: 8Gi
            requests:
              cpu: '3'
              memory: 5000Mi
          terminationMessagePath: /dev/termination-log
          terminationMessagePolicy: File
          volumeMounts:
            - mountPath: /opt/logs
              name: xxx-logs-volume
      dnsPolicy: ClusterFirst
      imagePullSecrets:
        - name: harborsecret
      restartPolicy: Always
      schedulerName: default-scheduler
      securityContext: {}
      terminationGracePeriodSeconds: 30
      tolerations:
        - effect: NoSchedule
          key: tag
          operator: Equal
          value: scm-only
      volumes:
        - name: xxx-logs-volume
          nfs:
            path: /opt/nfs_k8s_logs/scm/xxx
            server: xxx.xxx.xxx.xxx
status:
  availableReplicas: 5
  conditions:
    - lastTransitionTime: '2022-09-26T14:30:45Z'
      lastUpdateTime: '2022-09-26T14:30:45Z'
      message: Deployment has minimum availability.
      reason: MinimumReplicasAvailable
      status: 'True'
      type: Available
    - lastTransitionTime: '2022-07-04T09:11:19Z'
      lastUpdateTime: '2022-10-13T11:18:52Z'
      message: ReplicaSet "xxx-766854789f" is progressing.
      reason: ReplicaSetUpdated
      status: 'True'
      type: Progressing
  observedGeneration: 83
  readyReplicas: 6
  replicas: 6
  unavailableReplicas: 1
  updatedReplicas: 6
原因分析

出现这种差异核心是不同工具的CPU使用率计算基准不一致:

  • 容器内top的计算逻辑
    容器内的top默认以**Pod的CPU限制(CPU Limit)**为基准计算百分比。从配置看该Pod的CPU Limit是8核,容器内显示的125%实际是相对于这8核的比例,换算成实际占用核数是8 * 125% = 1核左右,和kubectl top pods显示的819m/945m(约0.8-0.9核)基本匹配。

  • kubectl top pods的计算逻辑
    kubectl top直接统计Pod实际占用的CPU资源,单位为毫核(m),1000m等于1核。它的数值是真实的CPU使用量,没有基于CPU Limit做比例换算,所以看起来数值远低于容器内top的百分比。

  • 物理机上的CPU使用率统计
    物理机统计Pod CPU使用率时,是以物理机总CPU核数为基准。比如物理机有32核,那1核的占用率仅为1/32≈3.125%,自然和容器内基于8核计算的125%差异极大。

业务运行正常也能佐证实际负载不高——容器内的高百分比只是计算基准带来的“假象”,实际占用的CPU资源远低于限制值,不会影响业务运行。


内容的提问来源于stack exchange,提问作者user502735

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 16:10:56