MicroK8s中Pod内存超限未终止问题排查咨询
问题背景
基于MicroK8s部署的服务,Deployment中配置容器内存请求为10Gi、限制为14Gi,但Pod实际占用内存已达20Gi,却未被自动终止重启。尝试过HPA、存活探针等手段均未解决问题。
Deployment配置
apiVersion: apps/v1 kind: Deployment metadata: name: example-development labels: app: example spec: replicas: 1 selector: matchLabels: app: example template: metadata: name: example labels: app: example spec: volumes: - name: appsettings-volume configMap: name: example-appsettings - name: dashboards-volume hostPath: path: /some path/example/api/data/Dashboards - name: reports-volume hostPath: path: /some path/example/api/data/Reports containers: - name: example image: myregistryurl/example:latest volumeMounts: - name: appsettings-volume mountPath: /app/appsettings.json subPath: appsettings.json - name: dashboards-volume mountPath: /app/Dashboards - name: reports-volume mountPath: /app/Reports resources: requests: memory: "10Gi" limits: memory: "14Gi" imagePullPolicy: "IfNotPresent" imagePullSecrets: - name: regcred restartPolicy: Always
Pod描述信息(kubectl describe pod输出)
Name: ****-development-*************** Namespace: default Priority: 0 Service Account: default Node: ****/************* Start Time: Mon, 25 Mar 2024 06:59:37 +0000 Labels: ****=***** pod-template-hash=*************** Annotations: cni.projectcalico.org/containerID: ****** cni.projectcalico.org/podIP: ***** cni.projectcalico.org/podIPs: ***** kubectl.kubernetes.io/restartedAt: 2024-03-15T11:00:35Z Status: Running IP: ***** IPs: IP: ***** Controlled By: ReplicaSet/****-development-*************** Containers: ****: Container ID: ****://************************************ Image: *****************/***/****/api:backup Image ID: ***************/*****/****/api@sha256:****************************** Port: <none> Host Port: <none> State: Running Started: Mon, 25 Mar 2024 06:59:38 +0000 Ready: True Restart Count: 0 Limits: memory: 14Gi Requests: memory: 10Gi Environment: <none> Mounts: /app/Dashboards from dashboards-volume (rw) /app/Reports from reports-volume (rw) /app/appsettings.json from appsettings-volume (rw,path="appsettings.json") /var/run/secrets/kubernetes.io/serviceaccount from kube-api-access-m9mhk (ro) Conditions: Type Status Initialized True Ready True ContainersReady True PodScheduled True Volumes: appsettings-volume: Type: ConfigMap (a volume populated by a ConfigMap) Name: ****-appsettings Optional: false dashboards-volume: Type: HostPath (bare host directory volume) Path: *****/api/data/Dashboards HostPathType: reports-volume: Type: HostPath (bare host directory volume) Path: *****/api/data/Reports HostPathType: kube-api-access-m9mhk: Type: Projected (a volume that contains injected data from multiple sources) TokenExpirationSeconds: 3607 ConfigMapName: kube-root-ca.crt ConfigMapOptional: <nil> DownwardAPI: true QoS Class: Burstable Node-Selectors: <none> Tolerations: node.kubernetes.io/not-ready:NoExecute op=Exists for 300s node.kubernetes.io/unreachable:NoExecute op=Exists for 300s Events: Type Reason Age From Message ---- ------ ---- ---- ------- Warning FailedToRetrieveImagePullSecret 15s (x13 over 12m) kubelet Unable to retrieve some image pull secrets (regcred); attempting to pull the image may not succeed.
可能的故障原因
HostPath挂载目录内存未被计入容器统计
Pod挂载了两个HostPath卷,若容器进程在这些卷中写入大量数据(如生成的报表、仪表盘文件),部分容器运行时(如containerd)默认不会将HostPath卷的内存占用计入容器内存使用量。这会让kubelet误以为容器内存未超限,不会触发OOM终止。kubelet OOM相关配置异常
MicroK8s默认的kubelet配置可能调整了OOM阈值,比如--oom-score-adj设置过低,或者内存可用量计算逻辑忽略了部分占用,导致kubelet未触发Pod驱逐或容器终止。可检查节点上/var/lib/kubelet/config.yaml中的OOM相关参数。容器进程未被系统OOM Killer选中
若容器内进程设置了极低的OOM_SCORE_ADJ值,会被操作系统OOM Killer优先忽略;另外如果节点剩余内存充足,即使Pod超限,kubelet也不会强制终止容器。可查看节点dmesg日志,确认是否有OOM相关记录。ImagePullSecret警告干扰管控逻辑
Pod事件中存在FailedToRetrieveImagePullSecret警告,虽当前Pod已运行,但kubelet处理内存限制时若遇到认证异常,可能导致部分管控逻辑失效。需验证regcredSecret是否存在且配置正确,消除警告后再观察。Burstable QoS等级的特性限制
该Pod属于Burstable QoS等级(requests≠limits),只有当节点内存紧张时,Burstable Pod才会被优先驱逐;若节点内存充足,即使Pod内存超限,kubelet也不会主动终止容器。
内容的提问来源于stack exchange,提问作者Gencay Tekin

