Google Cloud TCP负载均衡器TCP请求丢包问题技术咨询
Let’s walk through the most likely causes and troubleshooting steps for your TCP request drop issue with Google Cloud’s TCP Load Balancer— I’ve debugged similar setups multiple times, so here’s where to start digging:
First, double-check how your backend service is configured to map frontend ports to backend instance ports. Since your LB exposes a range (100-200) but your backend instances only listen on specific ports (100, 101, 103), you need to ensure:
- You’re using port mapping (not direct port forwarding) in the backend service configuration. Each frontend port should map to the corresponding backend instance’s service port.
- The backend service explicitly includes the ports 100, 101, and 103 in its allowed port list— if it’s set to a wider range but your instances don’t listen on all those ports, requests to unused ports will get dropped.
Unhealthy backend instances get removed from the load balancer’s pool, which will cause requests to be dropped if no healthy instances are available. Here’s what to check:
- Each backend instance has a health check targeting its specific service port (e.g., instance with service on 100 should have a health check for port 100).
- Health check protocol matches your service (TCP, since this is a TCP LB) and the timeout/interval settings are reasonable (avoid overly aggressive timeouts that mark healthy instances as unhealthy).
- In the Google Cloud Console, check the backend service’s health status to confirm all your instances are marked as healthy.
Firewall misconfigurations are one of the top causes of dropped traffic in GCP setups:
- Health check firewall rule: Ensure you have a rule allowing incoming traffic from Google’s health check IP ranges to your backend instances’ service ports (100, 101, 103).
- Backend instance inbound rules: Confirm there’s a rule allowing traffic from your LB’s forwarding rule IP (or the GCP LB IP range) to the three service ports.
- Local instance firewall/iptables: On each backend instance, run
iptables -Lto make sure there are no rules blocking incoming traffic to the service ports.
Even if the LB is configured correctly, if the service on the instance isn’t running or listening, requests will be dropped:
- On each instance, run
ss -tulpn(ornetstat -tulpnifssisn’t available) to verify the service is actively listening on its assigned port (look for entries likeLISTENwith the correct port number). - Use
lsof -i :100(replace 100 with 101/103) to check if any other process is occupying the service port— a port conflict would prevent your service from listening. - Run
tcpdump -i any port 100on the instance to see if traffic from the LB is actually reaching the instance. If no packets show up, the issue is in the network path; if packets arrive but no response is sent, the service itself is the problem.
GCP provides detailed logs and metrics to pinpoint drop reasons:
- In the Cloud Logging console, filter for your LB’s logs and look for entries with
severity=ERRORor keywords likedropped,backend_unhealthy, orno_healthy_backends. These will give you specific error messages explaining why requests are being rejected. - Check the backend service’s metrics in Cloud Monitoring: Look at
backend_healthto confirm instance health, andrequest_dropped_countto see the volume of dropped requests and their associated reasons (e.g., no healthy backends, port unreachable).
Make sure the end-to-end LB chain is correctly linked:
- Your forwarding rule should be configured to use the port range 100-200 and point to the correct TCP target proxy.
- The TCP target proxy should be associated with your backend service (the one linked to your three instances).
Start with health checks and firewall rules— those are the most frequent culprits in these setups. If you find specific error codes in the logs, that’ll help narrow down the issue even faster.
内容的提问来源于stack exchange,提问作者Chandan Kashyap

