Kubernetes LoadBalancer无法正常运行,请求故障排查
自建Kubernetes集群基于ESXi主机,包含1台Master节点(192.168.0.30)和3台Node节点(192.168.0.31-33),所有节点为Fedora Core虚拟机,配备双网卡,第二网卡位于VLAN6,对应IP为192.168.6.30-33,所有IP均可从工作站ping通,集群ClusterIP段为10.96.0.0/16(ClusterIP为10.96.0.1)。
部署了以下Deployment配置(hello-world.yml):
apiVersion: apps/v1 kind: Deployment metadata: name: hello-world spec: selector: matchLabels: run: load-balancer-example replicas: 3 template: metadata: labels: run: load-balancer-example spec: containers: - name: hello-world image: gcr.io/google-samples/node-hello:1.0 ports: - containerPort: 8080 protocol: TCP
以及Service配置(hello-world-serv.yml):
apiVersion: v1 kind: Service metadata: name: hello-world-service spec: selector: app: hello-world type: LoadBalancer ports: - protocol: TCP port: 8080 targetPort: 8080
应用配置后,hello-world-service始终处于等待External-IP的状态:
NAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGE service/hello-world-service LoadBalancer 10.97.220.48 <pending> 8080:31702/TCP 24s service/kubernetes ClusterIP 10.96.0.1 <none> 443/TCP 2d3h
手动添加externalIPs配置到Service后:
externalIPs: - 192.168.6.31 - 192.168.6.32 - 192.168.6.33
External-IP显示正常,但仍无法通过8080或31702端口访问Web界面:
NAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGE service/hello-world-service LoadBalancer 10.97.220.48 192.168.6.31,192.168.6.32,192.168.6.33 8080:31702/TCP 2m29s service/kubernetes ClusterIP 10.96.0.1 <none> 443/TCP 2d3h
请问遗漏了什么配置?
1. Service与Pod标签选择器不匹配(核心问题)
你的Deployment定义的Pod标签是run: load-balancer-example,但Service的selector设置为app: hello-world,两者完全不匹配,导致Service无法关联到任何Pod,自然无法提供访问。
修改Service的selector,与Deployment的Pod标签保持一致:
apiVersion: v1 kind: Service metadata: name: hello-world-service spec: selector: run: load-balancer-example # 修正为与Pod标签一致的值 type: LoadBalancer ports: - protocol: TCP port: 8080 targetPort: 8080 externalIPs: - 192.168.6.31 - 192.168.6.32 - 192.168.6.33
2. 自建集群缺少LoadBalancer控制器(External-IP pending的根源)
自建K8s集群没有云厂商提供的LoadBalancer控制器,因此type: LoadBalancer的Service会一直处于<pending>状态。你手动添加externalIPs是可行的替代方案,但要确保这些IP属于集群节点,且节点上的kube-proxy能正确转发流量。
3. 检查节点防火墙配置
Fedora默认启用firewalld,需确保节点上的NodePort(31702)和Service端口(8080)允许外部访问:
# 开放NodePort端口 firewall-cmd --add-port=31702/tcp --permanent # 开放Service端口 firewall-cmd --add-port=8080/tcp --permanent # 重新加载防火墙规则 firewall-cmd --reload
4. 验证Pod运行状态
执行以下命令确认Pod处于正常运行状态,无启动异常:
kubectl get pods -l run=load-balancer-example kubectl describe pods <pod-name>
5. 检查kube-proxy工作模式
确认kube-proxy的IPVS或iptables配置正常:
kubectl get configmap kube-proxy -n kube-system -o yaml
如果是iptables模式,需确保节点上的iptables规则生成正确;如果是IPVS模式,需确保节点已安装ipvsadm工具。
内容的提问来源于stack exchange,提问作者Gabrie

