You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于AKS-engine部署Cilium ClusterMesh无法访问外部端点求助

Cilium ClusterMesh on AKS-engine: Cannot Access Cross-Cluster Endpoints, No External Endpoints Declared in Proxy

Problem Description

I'm deploying Cilium ClusterMesh via AKS-engine, and have completed Cilium installation on two independent clusters. After following the ClusterMesh installation guide:

  • Node lists and statuses appear normal
  • etcd-operator logs show no errors
  • But cannot access cross-cluster endpoints; sample apps only return responses from the current cluster
  • Debugging the proxy reveals no external endpoints are declared

Each cluster has 1 master node and 2 worker nodes. Below are node lists and status details for both clusters. Additional logs are available upon request.

Cluster 1 Node List

kubectl -nkube-system exec -it cilium-vg8sm cilium node list
Name                                    IPv4 Address    Endpoint CIDR   IPv6 Address    Endpoint CIDR
cluster1/k8s-cilium2-29734124-0         172.18.2.5      192.168.1.0/24
cluster1/k8s-cilium2-29734124-1         172.18.2.4      10.4.0.0/16
cluster1/k8s-master-29734124-0          172.18.1.239    10.239.0.0/16
cluster2/k8s-cilium2-14610979-0         172.18.2.6      192.168.2.0/24
cluster2/k8s-cilium2-14610979-1         172.18.2.7      10.7.0.0/16
cluster2/k8s-master-14610979-0          172.18.2.239    10.239.0.0/16

Cluster 1 Status

kubectl -nkube-system exec -it cilium-vg8sm cilium status
KVStore:                Ok  etcd: 1/1 connected: https://cilium-etcd-client.kube-system.svc:2379 - 3.3.11
ContainerRuntime:       Ok  docker daemon: OK
Kubernetes:             Ok 1.15 (v1.15.1) [linux/amd64]
Kubernetes APIs:        ["CustomResourceDefinition", "cilium/v2::CiliumNetworkPolicy", "core/v1::Endpoint", "core/v1::Namespace", "core/v1::Node", "core/v1::Pods", "core/v1::Service", "networking.k8s.io/v1::NetworkPolicy"]
Cilium:                 Ok OK
NodeMonitor:            Disabled
Cilium health daemon:   Ok
IPv4 address pool:      10/65535 allocated from 10.4.0.0/16
Controller Status:      48/48 healthy
Proxy Status:           OK, ip 10.4.0.1, port-range 10000-20000
Cluster health:         6/6 reachable (2019-08-09T10:11:22Z)

Cluster 2 Node List

kubectl -nkube-system exec -it cilium-rl8gt cilium node list
Name                                    IPv4 Address    Endpoint CIDR   IPv6 Address    Endpoint CIDR
cluster1/k8s-cilium2-29734124-0         172.18.2.5      192.168.1.0/24
cluster1/k8s-cilium2-29734124-1         172.18.2.4      10.4.0.0/16
cluster1/k8s-master-29734124-0          172.18.1.239    10.239.0.0/16
cluster2/k8s-cilium2-14610979-0         172.18.2.6      192.168.2.0/24
cluster2/k8s-cilium2-14610979-1         172.18.2.7      10.7.0.0/16
cluster2/k8s-master-14610979-0          172.18.2.239    10.239.0.0/16

Cluster 2 Status

kubectl -nkube-system exec -it cilium-rl8gt cilium status
KVStore:                Ok  etcd: 1/1 connected: https://cilium-etcd-client.kube-system.svc:2379 - 3.3.11
ContainerRuntime:       Ok  docker daemon: OK
Kubernetes:             Ok 1.15 (v1.15.1) [linux/amd64]
Kubernetes APIs:        ["CustomResourceDefinition", "cilium/v2::CiliumNetworkPolicy", "core/v1::Endpoint", "core/v1::Namespace", "core/v1::Node", "core/v1::Pods", "core/v1::Service", "networking.k8s.io/v1::NetworkPolicy"]
Cilium:                 Ok OK
NodeMonitor:            Disabled
Cilium health daemon:   Ok
IPv4 address pool:      10/65535 allocated from 10.7.0.0/16
Controller Status:      48/48 healthy
Proxy Status:           OK, ip 10.7.0.1, port-range 10000-20000
Cluster health:         6/6 reachable (2019-08-09T10:40:39Z)

Alright, let’s dig into why your ClusterMesh isn’t exposing cross-cluster endpoints. The fact that nodes are visible across clusters and etcd looks healthy is a good start—so the issue is likely in service discovery, proxy configuration, or cross-cluster service synchronization. Here are the steps I’d take to debug this:

1. Verify ClusterMesh Core Configuration

First, double-check that each cluster’s Cilium agents are properly configured for ClusterMesh:

  • Check the cilium-config ConfigMap in both clusters:
    kubectl get configmap cilium-config -nkube-system -o yaml
    
    Confirm these settings are present and correct:
    • cluster-name should be unique for each cluster (e.g., cluster1 and cluster2)
    • enable-clustermesh: "true"
    • clustermesh-etcd-config should point to the shared etcd (or each cluster’s etcd if using a mesh setup)
  • Ensure the cilium-agent pods are running with the --cluster-name flag matching the ConfigMap. You can check this with:
    kubectl describe pod cilium-vg8sm -nkube-system | grep Args
    

2. Check Cross-Cluster Service Synchronization

The proxy not seeing external endpoints usually means services from the other cluster aren’t being synced. Let’s verify this:

  • List all synced Cilium services in each cluster:
    kubectl get ciliumservices -A
    
    You should see services from the opposite cluster prefixed with the remote cluster name. If not, check the clustermesh-apiserver logs (if you deployed it) for sync errors:
    kubectl logs -nkube-system deployment/clustermesh-apiserver
    
  • If you’re not using the clustermesh-apiserver, confirm that etcd is properly sharing service data between clusters. Run this in one cluster to check for remote service entries:
    kubectl -nkube-system exec -it cilium-etcd-0 -- etcdctl \
      --endpoints=https://cilium-etcd-client.kube-system.svc:2379 \
      --cacert=/var/lib/etcd-secrets/etcd-client-ca.crt \
      --cert=/var/lib/etcd-secrets/etcd-client.crt \
      --key=/var/lib/etcd-secrets/etcd-client.key \
      get cilium.io/services --prefix
    
    Look for entries containing the remote cluster’s namespace/service names.

3. Inspect Cilium Proxy and Endpoint Data

Since your proxy shows no external endpoints, let’s dig into the proxy’s state:

  • In a Cilium pod on each cluster, list all registered service endpoints:
    cilium service list
    
    Cross-cluster endpoints should be listed with the remote cluster’s node prefix. If missing, check the Cilium agent logs for ClusterMesh-related errors:
    kubectl logs -nkube-system cilium-vg8sm | grep -i clustermesh
    
    Look for messages about failed service sync or endpoint registration.

4. Validate Network Connectivity Between Clusters

Even though nodes show as reachable, let’s confirm deep connectivity:

  • From a worker node in Cluster1, ping the worker nodes in Cluster2 (e.g., 172.18.2.6 and 172.18.2.7). If ping fails, check your AKS-engine network security groups (NSGs) to ensure:
    • Port 4240 (Cilium node-to-node communication) is open between clusters
    • Port 2379 (etcd) is accessible across clusters if using a shared etcd
  • Try directly accessing a pod IP from the opposite cluster. If this works but service access doesn’t, the issue is isolated to service discovery/proxy configuration. If it fails, you have an underlying network connectivity problem.

5. Check Sample App Service Configuration

Finally, confirm your sample app’s Service is configured to work with ClusterMesh:

  • If using a ClusterIP Service, ensure Cilium is set to handle cross-cluster Service routing (this should be enabled by default when enable-clustermesh is true)
  • If using a Headless Service, verify that cluster.local DNS resolution is working across clusters. You can test this by running nslookup <service-name>.<namespace>.cluster.local from a pod in the opposite cluster.

Let me know what you find in these steps—we can narrow it down further based on the results!


内容的提问来源于stack exchange,提问作者Juan Manuel Tirado Martin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:20:45