Kubernetes部署Envoy作为gRPC代理连接超时问题排查求助
排查Kubernetes中Envoy代理连接超时问题
背景
在Kubernetes环境中部署Envoy作为服务代理,后端应用通过gRPC与客户端通信,已完成以下部署操作:
1. 编写Envoy配置YAML文件
admin: access_log_path: /tmp/admin_access.log address: socket_address: { address: 0.0.0.0, port_value: 9901 } static_resources: listeners: - name: http_listener address: socket_address: { address: 0.0.0.0, port_value: 8080 } filter_chains: - filters: - name: envoy.filters.network.http_connection_manager typed_config: "@type": type.googleapis.com/envoy.extensions.filters.network.http_connection_manager.v3.HttpConnectionManager codec_type: auto stat_prefix: ingress_http route_config: name: local_route virtual_hosts: - name: local_service domains: [ "*" ] routes: - match: { prefix: "/" } route: cluster: my_app_prod_service timeout: 30s max_grpc_timeout: 30s cors: allow_origin_string_match: - safe_regex: { regex: ".*", google_re2: { } } allow_methods: GET, PUT, DELETE, POST, OPTIONS allow_headers: keep-alive,user-agent,cache-control,content-type,content-transfer-encoding,custom-header-1,x-accept-content-transfer-encoding,x-accept-response-streaming,x-user-agent,x-grpc-web,grpc-timeout max_age: "1728000" expose_headers: custom-header-1,grpc-status,grpc-message http_filters: - name: envoy.filters.http.router clusters: - name: my_app_prod_service connect_timeout: 0.5s type: strict_dns http2_protocol_options: {} lb_policy: round_robin load_assignment: cluster_name: my_app_prod_service endpoints: - lb_endpoints: - endpoint: address: socket_address: address: my-app-service-staging port_value: 30015
2. 创建Envoy配置ConfigMap
kubectl create configmap envoy-config-prod \ --from-file=envoy_config_prod.yaml \ -o yaml --dry-run=client | kubectl replace --force -f -
3. 部署Envoy Deployment及Service
apiVersion: apps/v1 kind: Deployment metadata: name: envoy-server-prod spec: replicas: 3 selector: matchLabels: app: envoy-server-prod template: metadata: labels: app: envoy-server-prod spec: containers: - name: envoy-server-prod image: envoyproxy/envoy:v1.18.2 args: - -c - /etc/envoy/envoy_config_prod.yaml - --log-path - /tmp/envoy_info.log ports: - name: http containerPort: 8080 - name: envoy-admin containerPort: 9901 resources: requests: cpu: 5 memory: 5Gi volumeMounts: - mountPath: /etc/envoy name: envoy-config-prod volumes: - name: envoy-config-prod configMap: name: envoy-config-prod --- kind: Service apiVersion: v1 metadata: name: envoy-service-prod labels: app: envoy-service-prod spec: selector: app: envoy-server-prod ports: - name: http protocol: TCP port: 8080 targetPort: 8080 type: ClusterIP externalIPs: - 10.1.4.63
4. 部署后端Headless Service及应用Deployment
apiVersion: v1 kind: Service metadata: name: my-app-service-staging labels: app: my-app-service-staging spec: clusterIP: None ports: - name: grpc port: 30015 targetPort: 30015 protocol: TCP selector: app: my-app-deploy-staging --- apiVersion: apps/v1 kind: Deployment metadata: name: my-app-deploy-staging spec: replicas: 1 selector: matchLabels: app: my-app-deploy-staging template: metadata: labels: app: my-app-deploy-staging spec: containers: - name: my-app-deploy-staging image: $IMAGE_SHA resources: requests: memory: 2G cpu: 1
已确认Envoy Pod中配置文件和日志文件存在,且日志无报错信息。
当前症状
尝试通过ExternalIP访问Envoy服务时出现连接超时:
> curl -v 10.1.4.63:8080 * Trying 10.1.4.63:8080... * TCP_NODELAY set * connect to 10.1.4.63 port 8080 failed: Connection timed out * Failed to connect to 10.1.4.63 port 8080: Connection timed out * Closing connection 0 curl: (28) Failed to connect to 10.1.4.63 port 8080: Connection timed out
查看集群资源状态:
> k get svc my-app-service-staging ClusterIP None <none> 30015/TCP 3h54m envoy-service-prod ClusterIP 10.43.157.121 10.1.4.63 8080/TCP 6h55m > k get deploy my-deploy-deploy-staging 1/1 1 1 29d
排查步骤
1. 验证Envoy Pod的运行状态与内部连通性
- 确认Envoy Pod是否正常运行:
检查kubectl get pods -l app=envoy-server-prodREADY列是否为1/1,STATUS是否为Running。 - 进入Envoy Pod内部,测试自身8080端口的可达性:
如果内部访问正常,说明问题出在Pod外部的网络链路;如果内部也超时,检查Envoy配置是否正确加载或监听是否异常。kubectl exec -it <envoy-pod-name> -- curl localhost:8080 - 测试Envoy Pod到后端Headless Service的连通性:
确认DNS解析是否正常,后端服务端口是否可达。kubectl exec -it <envoy-pod-name> -- nc -zv my-app-service-staging 30015
2. 检查Service的端点关联情况
- 查看Envoy Service的端点信息,确认是否关联到运行中的Envoy Pod:
检查kubectl describe svc envoy-service-prodEndpoints字段是否包含Envoy Pod的IP和8080端口。 - 检查后端Headless Service的端点:
确认是否关联到kubectl describe svc my-app-service-stagingmy-app-deploy-stagingPod的IP和30015端口。
3. 排查网络策略与防火墙限制
- 检查集群内是否存在NetworkPolicy,限制了Envoy Service的ExternalIP访问,或者Envoy Pod与后端Pod之间的通信:
kubectl get networkpolicy kubectl describe networkpolicy <policy-name> - 确认集群节点的防火墙规则,是否允许外部流量访问10.1.4.63的8080端口,以及节点到Pod的网络转发是否开启。
4. 验证Envoy监听与配置生效情况
- 通过Envoy Admin端口查看监听状态:
然后访问kubectl port-forward <envoy-pod-name> 9901:9901http://localhost:9901/listeners,确认8080端口的监听是否正常启动。 - 调整Envoy日志级别(如添加
--log-level info到启动参数),重新部署后查看详细日志,排查是否有隐藏的连接或路由错误。
5. 确认ExternalIP的有效性
- 验证10.1.4.63是否为集群节点的有效IP地址,且该节点上运行有Envoy Pod(或集群网络允许ExternalIP绑定到任意节点)。
- 从集群节点内部访问10.1.4.63:8080,测试连通性:
如果节点内部能访问,说明问题出在外部到节点的网络链路;如果节点内部也超时,检查Service的ExternalIP配置是否正确。curl 10.1.4.63:8080
内容的提问来源于stack exchange,提问作者Tinyden
相关产品推荐
相关产品推荐

