K3S AMD64+ARM64集群Pod持续Pending且无事件求助
问题复现
K3S集群主控节点为线上裸机AMD64服务器,工作节点为本地PI400 ARM64 Debian主机。部署以下DaemonSet后,所有Pod均处于Pending状态:
apiVersion: apps/v1 kind: DaemonSet metadata: name: hello-world labels: app: hello-world spec: selector: matchLabels: app: hello-world template: metadata: labels: app: hello-world spec: containers: - name: hello-world image: nginxdemos/hello
Pod状态
kubectl get pods
NAME READY STATUS RESTARTS AGE hello-world-qsv2d 0/1 Pending 0 7m53s hello-world-6rn5d 0/1 Pending 0 7m53s
Pod详情
kubectl describe pod hello-world-6rn5d
Name: hello-world-6rn5d Namespace: default Priority: 0 Node: <none> Labels: app=hello-world controller-revision-hash=649569d94c pod-template-generation=1 Annotations: <none> Status: Pending IP: IPs: <none> Controlled By: DaemonSet/hello-world Containers: hello-world: Image: hello-world Port: <none> Host Port: <none> Environment: <none> Mounts: /var/run/secrets/kubernetes.io/serviceaccount from kube-api-access-8bh8p (ro) Volumes: kube-api-access-8bh8p: Type: Projected (a volume that contains injected data from multiple sources) TokenExpirationSeconds: 3607 ConfigMapName: kube-root-ca.crt ConfigMapOptional: <nil> DownwardAPI: true QoS Class: BestEffort Node-Selectors: <none> Tolerations: node.kubernetes.io/disk-pressure:NoSchedule op=Exists node.kubernetes.io/memory-pressure:NoSchedule op=Exists node.kubernetes.io/not-ready:NoExecute op=Exists node.kubernetes.io/pid-pressure:NoSchedule op=Exists node.kubernetes.io/unreachable:NoExecute op=Exists node.kubernetes.io/unschedulable:NoSchedule op=Exists Events: <none>
集群节点状态
kubectl get nodes
NAME STATUS ROLES AGE VERSION pi417 Ready <none> 11h v1.24.4+k3s1 pi400 Ready <none> 11h v1.24.4+k3s1
集群版本信息
kubectl version --output=yaml
clientVersion: buildDate: "2022-06-15T14:22:29Z" compiler: gc gitCommit: f66044f4361b9f1f96f0053dd46cb7dce5e990a8 gitTreeState: clean gitVersion: v1.24.2 goVersion: go1.18.3 major: "1" minor: "24" platform: windows/amd64 kustomizeVersion: v4.5.4 serverVersion: buildDate: "2022-08-25T03:45:26Z" compiler: gc gitCommit: c3f830e9b9ed8a4d9d0e2aa663b4591b923a296e gitTreeState: clean gitVersion: v1.24.4+k3s1 goVersion: go1.18.1 major: "1" minor: "24" platform: linux/amd64
工作节点详情
kubectl describe node pi400
Name: pi400 Roles: <none> Labels: adb=true beta.kubernetes.io/arch=arm64 beta.kubernetes.io/instance-type=k3s beta.kubernetes.io/os=linux egress.k3s.io/cluster=true kubernetes.io/arch=arm64 kubernetes.io/hostname=pi400 kubernetes.io/os=linux node.kubernetes.io/instance-type=k3s Annotations: flannel.alpha.coreos.com/backend-data: {"VNI":1,"VtepMAC":"26:a8:bd:f3:1d:fd"} flannel.alpha.coreos.com/backend-type: vxlan flannel.alpha.coreos.com/kube-subnet-manager: true flannel.alpha.coreos.com/public-ip: 192.168.3.25 k3s.io/hostname: pi400 k3s.io/internal-ip: 192.168.3.25 k3s.io/node-args: ["agent"] k3s.io/node-config-hash: CBEQF3QV5PMMQWO2GECMRPJVEIFSCEFARQFZKX4RNV4K5FPB7FGQ==== k3s.io/node-env: {"K3S_DATA_DIR":"/var/lib/rancher/k3s/data/8...2a","K3S_NODE_NAME":"pi400" ...} node.alpha.kubernetes.io/ttl: 0 volumes.kubernetes.io/controller-managed-attach-detach: true CreationTimestamp: Mon, 12 Sep 2022 20:44:50 +0300 Taints: <none> Unschedulable: false Lease: HolderIdentity: pi400 AcquireTime: <unset> RenewTime: Tue, 13 Sep 2022 08:53:29 +0300 Conditions: Type Status LastHeartbeatTime LastTransitionTime Reason Message ---- ------ ----------------- ------------------ ------ ------- MemoryPressure False Tue, 13 Sep 2022 08:51:08 +0300 Mon, 12 Sep 2022 21:33:41 +0300 KubeletHasSufficientMemory kubelet has sufficient memory available DiskPressure False Tue, 13 Sep 2022 08:51:08 +0300 Mon, 12 Sep 2022 21:33:41 +0300 KubeletHasNoDiskPressure kubelet has no disk pressure PIDPressure False Tue, 13 Sep 2022 08:51:08 +0300 Mon, 12 Sep 2022 21:33:41 +0300 KubeletHasSufficientPID kubelet has sufficient PID available Ready True Tue, 13 Sep 2022 08:51:08 +0300 Mon, 12 Sep 2022 21:33:41 +0300 KubeletReady kubelet is posting ready status Addresses: InternalIP: 192.168.3.25 Hostname: pi400 Capacity: cpu: 4 ephemeral-storage: 30473608Ki memory: 3885428Ki pods: 110 Allocatable: cpu: 4 ephemeral-storage: 29644725840 memory: 3885428Ki pods: 110 System Info: Machine ID: d2eb1415b12e45ebac766cc20ce58012 System UUID: d2eb1415b12e45ebac766cc20ce58012 Boot ID: c2531ffa-96b0-4463-9f51-08e0dce6d5c3 Kernel Version: 5.15.61-v8+ OS Image: Debian GNU/Linux 11 (bullseye) Operating System: linux Architecture: arm64 Container Runtime Version: containerd://1.6.6-k3s1 Kubelet Version: v1.24.4+k3s1 Kube-Proxy Version: v1.24.4+k3s1 PodCIDR: 10.42.1.0/24 PodCIDRs: 10.42.1.0/24 ProviderID: k3s://pi400 Non-terminated Pods: (0 in total) Namespace Name CPU Requests CPU Limits Memory Requests Memory Limits Age --------- ---- ------------ ---------- --------------- ------------- --- Allocated resources: (Total limits may be over 100 percent, i.e., overcommitted.) Resource Requests Limits -------- -------- ------ cpu 0 (0%) 0 (0%) memory 0 (0%) 0 (0%) ephemeral-storage 0 (0%) 0 (0%) Events: <none>
怀疑方向
- AMD64与ARM64架构混合兼容性
- 节点间网络连接问题
- K3S集群组件异常
排查与解决方案
1. 镜像架构与配置修正
从Pod详情中发现,实际使用的镜像为hello-world,与部署文件中的nginxdemos/hello不符,这是核心问题之一:
- 重新部署正确的DaemonSet配置,确保镜像名称无误:
kubectl apply -f /path/to/your-daemonset.yaml
- 验证镜像是否支持ARM64架构,在ARM64节点执行:
crictl pull nginxdemos/hello
若拉取失败,更换为明确支持多架构的镜像(如nginx:alpine),修改DaemonSet的image字段后重新部署。
2. 调度器日志排查
主控节点查看kube-scheduler日志,定位调度失败原因:
journalctl -u k3s -f | grep -E 'scheduler|FailedScheduling'
常见错误包括:节点资源不足、镜像架构不匹配、调度约束冲突等,根据日志输出针对性处理。
3. 工作节点Kubelet状态检查
在PI400节点查看K3S Agent日志,确认节点与主控通信正常:
journalctl -u k3s-agent -f
若出现Failed to connect to API server类错误,检查主控节点6443端口是否对外开放、节点间网络是否连通。
4. 集群网络验证
- 检查Flannel网络组件状态:
crictl pods | grep flannel
确认Flannel Pod处于Running状态,同时查看节点VXLAN网络接口:
ip addr show flannel.1
- 验证主控与工作节点端口连通性:
工作节点访问主控6443端口:
nc -zv <主控节点IP> 6443
主控节点访问工作节点10250端口:
nc -zv <PI400节点IP> 10250
5. K3S版本与配置校验
确认主控与工作节点K3S版本一致(当前均为v1.24.4+k3s1,符合要求),若存在版本差异,升级或降级至统一版本。检查工作节点加入集群时的启动参数,确保指定了正确的主控节点地址与token。
内容的提问来源于stack exchange,提问作者Uriel

