无法完成Eclipse Che安装的问题排查求助
问题分析与解决步骤
核心问题定位
从che-operator的日志可以明确看到关键错误:
ERROR Unable determine installation platform {"error": "could not read API groups: Get \"https://10.96.0.1:443/api?timeout=32s\": dial tcp 10.96.0.1:443: i/o timeout"}
这说明che-operator Pod无法访问Kubernetes API Server的ClusterIP(10.96.0.1),属于集群内部网络连通性问题,devworkspace-controller的CrashLoopBackOff状态大概率也是同一原因导致。
排查与解决步骤
1. 验证Pod所在节点的网络连通性
先确定che-operator Pod运行的节点:
kubectl get pod che-operator-69bfb7c98-dg9wc -n eclipse-che -o wide
登录该节点,直接测试与API Server ClusterIP的连通性:
ping 10.96.0.1
如果ping不通,检查以下内容:
- 节点路由配置,确认10.96.0.1属于集群Service CIDR范围(默认是10.96.0.0/12)
- 节点防火墙/iptables规则,是否放行到10.96.0.1:443的流量
- CNI插件(如Calico、Flannel)的节点Pod是否健康,执行
kubectl get pod -n kube-system查看相关组件状态
2. 测试Pod内部到API Server的连通性
若节点能ping通但Pod内部不行,进入che-operator容器(即使处于CrashLoopBackOff状态也可执行):
kubectl exec -it che-operator-69bfb7c98-dg9wc -n eclipse-che -- sh
在容器内执行:
curl -v https://10.96.0.1:443/api?timeout=32s
如果连接超时,检查:
- CNI插件是否正确配置Pod网络,Pod是否获得合法IP
- 集群NetworkPolicy是否阻止eclipse-che命名空间的Pod访问kube-apiserver
- kube-proxy组件是否正常运行,执行
kubectl get pod -n kube-system -l k8s-app=kube-proxy确认状态
3. 检查ServiceAccount与RBAC配置
确认che-operator的ServiceAccount具备访问API Server的权限:
kubectl auth can-i get apigroups --as=system:serviceaccount:eclipse-che:che-operator
若返回no,重新检查RoleBinding配置:
kubectl describe rolebinding che-operator -n eclipse-che
确保RoleBinding的subjects包含system:serviceaccount:eclipse-che:che-operator,且关联的Role拥有apigroups的访问权限。
4. 清理残留资源后重新部署
之前的尝试可能留下冲突资源,建议清理后重新部署:
- 删除相关命名空间:
kubectl delete namespace eclipse-che devworkspace-controller - 删除相关CRDs:
kubectl delete crd checlusters.org.eclipse.che devworkspaces.workspace.devfile.io devworkspacetemplates.workspace.devfile.io - 修正
--domain参数(内部Service域名无需加https://)后重新执行部署:chectl server:deploy --platform=k8s --che-operator-cr-patch-yaml ~/che/patch.yaml --domain kubernetes.default.svc.cluster.local --skip-cert-manager --k8spodreadytimeout=500000 --k8spoderrorrechecktimeout=500000 --installer=operator --debug
5. 确认kube-apiserver可用性
检查kube-apiserver Pod是否正常运行:
kubectl get pod -n kube-system -l component=kube-apiserver
若API Server异常,先修复该问题再重新部署Che。
内容的提问来源于stack exchange,提问作者ScorpioTiger
相关产品推荐
相关产品推荐

