GCE负载均衡器关联K8s服务异常:服务健康性排查与修复咨询
Hey there, let's dig into why your GCE LB is marking your API service as unhealthy even though the service itself runs fine. First, I spotted a critical port mapping issue in your configs that's likely the root cause, then we'll cover deeper troubleshooting steps just in case.
1. Fix the Port Mapping Mismatch (Most Likely Culprit)
Looking at your Service and Deployment configs, there's a clear mismatch here:
Your Service defines an external port that routes port 80 to targetPort: 80:
# Service.yaml snippet ports: - name: external port: 80 targetPort: 80
But your Deployment's container only exposes port 8080:
# Deployment.yaml snippet ports: - name: api containerPort: 8080
The GCE LB uses the Service's ports to perform health checks. When it tries to hit port 80 on your Service, there's no container listening on port 80 to respond—so it marks the service as unhealthy.
Fix this by updating the Service's targetPort to match your container's port. Using the port name (api) is better practice than hardcoding numbers to avoid future mismatches:
ports: - name: http port: 8080 targetPort: api # Matches the container port name in Deployment - name: external port: 80 targetPort: api # Same here
2. Deep Dive Troubleshooting Steps (If Port Fix Doesn't Resolve)
If correcting the ports doesn't fix the health check issue, work through these steps:
Verify Pod Readiness Probe Status
Kubernetes' readiness probe is the foundation for GCE LB health checks—if the probe fails, K8s removes the pod from the Service's endpoint list, and the LB sees it as unhealthy.
- Check pod status:
kubectl get pods -n production—ensure theREADYcolumn shows1/1 - Inspect probe details:
kubectl describe pod <your-api-pod-name> -n production—look for Readiness probe failures in the Events section - Test the probe manually inside the pod:
kubectl exec <your-api-pod-name> -n production -- curl http://localhost:8080/healthz—this should return a 200 OK status code
Quick note: I noticed your liveness probe uses /readinez—is that a typo for /readiness? A typo here would cause liveness failures, but if readiness is working, the pod should still be marked ready. Worth double-checking though!
Check GCE Load Balancer Health Check Config
The GCE Ingress controller auto-creates health checks for your service. Verify these settings:
- Head to the GCE Console's Load Balancing page, find your LB, and check the health check configuration. Confirm the target port matches your Service's NodePort (you can get this with
kubectl get service api -n production) and the path is set to a valid endpoint (like your readiness probe's/healthz). - Test the health check from a cluster node:
curl http://<node-ip>:<node-port>/healthz—this should return 200 OK.
Validate Ingress-Service Association
Make sure your Ingress is correctly linked to the Service:
- Check Ingress status:
kubectl describe ingress api -n production—look for errors like "no endpoints available" which would indicate a misconfiguration. - Confirm the Ingress's
servicePort: 80matches the Service'sexternalport definition (which we fixed earlier).
Check GCE Firewall Rules
GCE health checks come from specific IP ranges—ensure your firewall allows access to the NodePort range:
- In the GCE Console, check firewall rules to confirm access is allowed for
130.211.0.0/22and35.191.0.0/16(GCE health check IPs) on ports 30000-32767 (default NodePort range).
3. Post-Fix Validation
After applying the corrected Service config:
- Run
kubectl apply -f Service.yaml -n production - Wait a few minutes, then check endpoints:
kubectl get endpoints api -n production—you should see your pod's IP listed - Check the GCE LB's health check status to confirm it's now marked healthy
- Test access to
foo.bar.ioto ensure the service responds correctly
内容的提问来源于stack exchange,提问作者Tino

