MicroK8s集群中Flask跨Pod调用REST API出现502错误求助
在MicroK8s集群部署了三组件应用:Angular前端、运行在5000端口的Flask后端Backend1、运行在5001端口的Flask后端Backend2。目前前端与Backend1通信正常,但Backend1调用Backend2的特定流程时,返回502 Bad Gateway错误。
背景说明:原本为三个组件配置了Service LoadBalancer,为实现HTTPS关闭了LoadBalancer,启用MicroK8s的Ingress插件,并为每个组件配置了Ingress对象。Backend1基于Gunicorn/Nginx部署,Backend2基于Gunicorn部署,所有组件以Docker容器形式运行在Pod中。
已尝试的操作:
- 在Backend1的nginx.conf中新增5001端口对应的location代理配置
- 调整Backend2的Ingress对象的多种配置组合
- 移除Backend1调用代码中的5001端口,但不确定正确的端口指定方式
- 在技术社区查找类似问题,未找到匹配场景
相关配置文件
Backend1的Ingress配置
apiVersion: networking.k8s.io/v1 kind: Ingress metadata: name: ingress-routes-backend1 namespace: app-namespace annotations: cert-manager.io/cluster-issuer: "letsencrypt-prod" nginx.ingress.kubernetes.io/use-regex: "true" spec: tls: - hosts: - domain-name secretName: tls-secret rules: - host: domain-name http: paths: - path: /api/v1/(.+) pathType: Prefix backend: service: name: backend1 port: number: 5000
Backend2的Ingress配置
apiVersion: networking.k8s.io/v1 kind: Ingress metadata: name: ingress-routes-backend2 namespace: app-namespace annotations: cert-manager.io/cluster-issuer: "letsencrypt-prod" nginx.ingress.kubernetes.io/use-regex: "true" spec: tls: - hosts: - domain-name secretName: tls-secret rules: - host: domain-name http: paths: - path: /api/v1/data/(.*) pathType: Prefix backend: service: name: backend2 port: number: 5001
Backend1的Service配置
apiVersion: v1 kind: Service metadata: labels: app: app-name name: backend1 name: backend1 namespace: app-namespace clusterIP: 10.xxx.yyy.128 clusterIPs: - 10.xxx.yyy.128 externalTrafficPolicy: Cluster ports: - name: http nodePort: 31261 port: 5000 protocol: TCP targetPort: 5000 selector: app: app-name name: backend1 type: NodePort status: loadBalancer: {}
Backend2的Service配置
apiVersion: v1 kind: Service metadata: labels: app: app-name name: backend2 name: backend2 namespace: app-namespace clusterIP: 10.xxx.yyy.128 clusterIPs: - 10.xxx.yyy.128 externalTrafficPolicy: Cluster ports: - name: http nodePort: 31261 port: 5001 protocol: TCP targetPort: 5001 selector: app: app-name name: backend1 type: NodePort status: loadBalancer: {}
Backend1调用Backend2的代码片段
@app.route('/api/v1/process-order/<customer_id>/customer/<order_id>/order/<bill_date>/billdate', methods=['GET']) @jwt_required def process_order(customer_id, order_id, bill_date): data = { "customer_id": customer_id, "order_id": order_id, "bill_date": bill_date, } response1 = requests.post( BASE_URL + ':5001/api/v1/data/load-order', headers=header1, data=json.dumps(data)) result1 = response1.json()
注:BASE_URL为域名
Backend2被调用的接口代码
@app.route('/api/v1/data/load-order', methods=['POST']) def load_order(): pass
Backend1的nginx.conf配置
user www-data; worker_processes auto; pid /run/nginx.pid; events { worker_connections 1024; use epoll; multi_accept on; } http { client_max_body_size 100M; access_log /dev/stdout; error_log /dev/stdout; sendfile on; tcp_nopush on; tcp_nodelay on; keepalive_timeout 65; types_hash_max_size 2048; include /etc/nginx/mime.types; default_type application/octet-stream; index index.html index.htm; server { listen 5000; listen [::]:5000; root /var/www/html; server_name localhost default_server; location / { proxy_read_timeout 360s; proxy_send_timeout 360s; proxy_connect_timeout 360s; uwsgi_read_timeout 360s; uwsgi_send_timeout 360s; uwsgi_connect_timeout 360s; include uwsgi_params; uwsgi_pass unix:/tmp/uwsgi.socket; } } }
Backend1的Dockerfile
FROM python:3.6 COPY code/ app/ WORKDIR /app RUN apt-get clean \ && apt-get -y update \ && apt-get -y install nginx \ && apt-get -y install python3-dev \ && apt-get -y install build-essential \ && pip install -r requirements.txt # Start nginx flask COPY code/nginx.conf /etc/nginx RUN chmod +x ./start.sh CMD ["./start.sh"]
Backend2的Dockerfile
FROM python:3.6 COPY ./ app/ WORKDIR /app RUN pip install -r requirements.txt CMD ["gunicorn", "-w", "3", "-b", ":5001", "-t", "360", "--access-logfile", "-", "--error-logfile", "-", "--reload", "app:app"]
1. 修复Backend2 Service的Selector错误
查看Backend2的Service配置,发现selector字段的name: backend1是明显错误,应指向Backend2的Pod标签,修正后:
selector: app: app-name name: backend2
这个错误会导致Service无法找到Backend2的Pod,是502错误的核心原因之一。
2. 调整Backend1调用Backend2的方式
方式一:使用集群内部Service域名调用(推荐)
在Kubernetes集群内部,直接通过Service的DNS名称访问,无需走Ingress,更高效且避免外部网络问题:
# 替换原BASE_URL调用方式,格式为:<service-name>.<namespace>.svc.cluster.local response1 = requests.post( 'http://backend2.app-namespace.svc.cluster.local:5001/api/v1/data/load-order', headers=header1, data=json.dumps(data) )
方式二:通过Ingress域名调用(保留原域名方式)
如果必须走Ingress,注意Ingress监听的是443端口(HTTPS)或80端口(HTTP),不需要指定5001端口,Ingress会根据路径转发到对应的Service:
# 去掉:5001,使用HTTPS协议 response1 = requests.post( BASE_URL + '/api/v1/data/load-order', headers=header1, data=json.dumps(data), verify=True # 若使用自签名证书可设为False,生产环境不建议 )
3. 检查Ingress路径匹配规则
Backend1的Ingress路径是/api/v1/(.+),Backend2的是/api/v1/data/(.*),需确保路径匹配正确。另外,可查看Ingress Controller日志排查路径转发问题:
microk8s kubectl logs -n ingress -l app.kubernetes.io/name=nginx-ingress
4. 验证Backend2的Pod和Service连通性
在Backend1的Pod中执行以下命令,测试是否能连通Backend2的Service:
# 进入Backend1的Pod microk8s kubectl exec -it <backend1-pod-name> -n app-namespace -- bash # 测试连通Backend2 Service curl http://backend2.app-namespace.svc.cluster.local:5001/api/v1/data/load-order -X POST -d '{"test": "data"}' -H "Content-Type: application/json"
内容的提问来源于stack exchange,提问作者incrediblegiant

