You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

安装Kubeflow时istio-ingressgateway Pod无法就绪的问题求助

问题描述
  • 环境:Ubuntu 22.04,kubectl 1.27.4,minikube v1.31.1 单节点Kubernetes集群,集群启动命令:
    minikube start --driver=none \ 
    --kubernetes-version=v1.27.4 \ 
    --extra-config=apiserver.service-account-signing-key-file=/var/lib/minikube/certs/sa.key \ 
    --extra-config=apiserver.service-account-issuer=kubernetes.default.svc  \
    --extra-config=kubelet.resolv-conf=/run/systemd/resolve/resolv.conf
    
  • 操作:通过kubeflow manifests安装Kubeflow(测试过1.7版本和master分支),执行Istio安装命令后,istio-ingressgateway Pod始终无法就绪:
    # Istio 1.16版本安装命令
    kustomize build common/istio-1-16/istio-install/base | kubectl apply -f -
    # 或Istio 1.17版本安装命令
    kustomize build common/istio-1-17/istio-install/base | kubectl apply -f -
    
  • Pod状态:
    NAME                                   READY   STATUS    RESTARTS   AGE
    istio-ingressgateway-79b665c95-xm454   0/1     Running   0          16s
    istiod-86457659bb-5h58w                1/1     Running   0          16s
    
  • istio-ingressgateway日志报错(DNS解析失败):
    执行命令:
    kubectl logs istio-ingressgateway-79b665c95-xm454 -n istio-system
    
    输出内容:
    warn    sds failed to warm certificate: failed to generate workload certificate: create certificate: rpc error: code = Unavailable desc = connection error: desc = "transport: Error while dialing dial tcp: lookup istiod.istio-system.svc on 10.96.0.10:53: read udp 10.244.0.212:33211->10.96.0.10:53: read: connection refused"
    
  • DNS测试:使用dnsutils执行kubectl exec -i -t dnsutils -- nslookup kubernetes.default超时,提示无法连接DNS服务器。
  • CoreDNS日志显示转发到外部DNS时I/O超时:
    [INFO] 127.0.0.1:51197 - 25941 "HINFO IN 6443204498907807806.1822223731783212999. udp 57 false 512" - - 0 2.000526828s 
    [ERROR] plugin/errors: 2 6443204498907807806.1822223731783212999. HINFO: read udp 10.244.0.227:60893->195.130.130.4:53: i/o timeout 
    [INFO] 127.0.0.1:44564 - 55305 "HINFO IN 6443204498907807806.1822223731783212999. udp 57 false 512" - - 0 2.000507048s 
    [ERROR] plugin/errors: 2 6443204498907807806.1822223731783212999. HINFO: read udp 10.244.0.227:58119->195.130.131.4:53: i/o timeout
    
  • 已尝试操作:修改CoreDNS的ConfigMap,替换转发的resolv.conf路径,添加cluster.local域名转发,问题未解决。
排查与解决方案

1. 验证宿主机DNS可用性

minikube none模式依赖宿主机网络,先确认宿主机DNS是否正常:

# 查看宿主机当前DNS配置
cat /run/systemd/resolve/resolv.conf
# 测试宿主机DNS解析能力
nslookup google.com

如果宿主机DNS本身不可用,先修复宿主机配置(比如将DNS服务器替换为8.8.8.8或1.1.1.1)。

2. 修改CoreDNS配置强制使用公共DNS

直接替换CoreDNS的转发DNS服务器,避免依赖宿主机的问题:

# 编辑CoreDNS ConfigMap
kubectl edit configmap coredns -n kube-system

将配置中的forward . /run/systemd/resolve/resolv.conf替换为:

forward . 8.8.8.8 1.1.1.1

保存后重启CoreDNS Pod:

kubectl rollout restart deployment coredns -n kube-system

3. 验证内部DNS解析

等待CoreDNS重启完成后,测试Istiod服务的DNS解析:

kubectl exec -i -t dnsutils -- nslookup istiod.istio-system.svc

如果解析正常,再查看istio-ingressgateway Pod状态:

kubectl get pods -n istio-system

4. 检查minikube none模式的网络规则

none模式下,确认宿主机iptables规则未阻止DNS流量:

# 查看DNS相关iptables规则
sudo iptables-save | grep dns
# 若规则异常,可重置iptables(谨慎操作:会清除现有规则)
sudo iptables -F && sudo iptables -t nat -F && sudo iptables -t mangle -F && sudo iptables -X

重置后重启minikube集群:

minikube stop
minikube start --driver=none \ 
--kubernetes-version=v1.27.4 \ 
--extra-config=apiserver.service-account-signing-key-file=/var/lib/minikube/certs/sa.key \ 
--extra-config=apiserver.service-account-issuer=kubernetes.default.svc  \
--extra-config=kubelet.resolv-conf=/run/systemd/resolve/resolv.conf

内容的提问来源于stack exchange,提问作者trial and error

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 06:35:09