多Kubernetes集群共享服务部署位置选型咨询
Hey Marius, let's break down your questions and deployment options with real-world Kubernetes operational practices in mind.
Is it a critical issue if monitoring data is stored within the cluster and becomes unavailable when the cluster goes down?
Absolutely—this is a serious problem, especially for production environments. When a cluster crashes, you need historical monitoring data to diagnose why it failed: was it resource exhaustion, a network partition, a control plane component failure, or something else? If your monitoring stack goes down alongside the cluster, you lose the most critical tool for root-cause analysis, which can drastically extend downtime and delay recovery. For staging and testing, this might be less impactful, but it’s still a gap that makes troubleshooting harder.
Is creating a separate cluster for shared services (to avoid impacting applications) a common practice?
Yes, this is a widely adopted pattern in multi-cluster Kubernetes environments. Many teams refer to this as a management cluster or tools cluster. Here’s why it makes sense:
- Isolation: Shared tools (monitoring, logging, CI/CD) don’t compete for resources with your application workloads in production/staging/testing clusters.
- Resilience: If an application cluster goes down, your management cluster remains operational, so you can still access monitoring data, logs, and troubleshooting tools to fix the issue.
- Centralized management: You only need to deploy and maintain one instance of each shared service, rather than three, which reduces operational overhead long-term.
While it adds some cost, the tradeoff for reliability and easier management is almost always worth it for teams running multiple clusters.
Can metrics and logs be easily transferred between Kubernetes clusters in OpenTelekomCloud (based on OpenStack)?
Yes, you can set up smooth cross-cluster data transfer in OpenTelekomCloud. Since it’s built on OpenStack, you have native network tools to enable secure communication between clusters:
- VPC Peering: Connect the virtual private clouds (VPCs) of your application clusters and your management cluster to allow internal, low-latency communication.
- Network Policies: Restrict traffic between clusters to only the necessary ports (e.g., InfluxDB’s 8086 for metrics, Fluentd’s 24224 for logs) to maintain security.
- Encrypted Transfer: Use TLS to encrypt metrics and logs in transit, which is straightforward with tools like InfluxDB, Grafana, or popular log collectors like Fluentd or Loki.
For your pre-installed Heapster, you can reconfigure it to push metrics to the InfluxDB instance in your management cluster instead of a local one—this is a standard configuration change in Heapster’s deployment manifest.
Let’s weigh your three options against the above context:
1. Deploy shared services in each cluster
- Pros: No extra cluster cost, minimal initial network setup.
- Cons:
- Redundant deployment and maintenance (three instances of each tool).
- No centralized view of metrics/logs across environments.
- Critical monitoring data is lost if the cluster goes down.
- Verdict: Only acceptable for short-term testing environments, not suitable for long-term production/staging use.
2. Create an additional management cluster
- Pros:
- Centralized, resilient monitoring/logging/tooling.
- Isolates tooling from application workloads.
- Simplifies long-term maintenance and troubleshooting.
- Cons:
- Adds cloud resource costs for the extra cluster.
- Requires setting up cross-cluster network connectivity.
- Verdict: This is the best practice for your use case, and the long-term benefits far outweigh the initial cost and setup effort.
3. Run shared services on a standalone root-server (Docker containers)
- Pros: Lower initial cost compared to a full Kubernetes cluster.
- Cons:
- Violates Kubernetes orchestration principles—you lose benefits like auto-scaling, self-healing, and service discovery.
- Single point of failure: if the root-server goes down, all your shared services are unavailable.
- Manual container management (or adding Rancher) adds operational overhead that negates any cost savings.
- Verdict: Not recommended for a production-grade multi-cluster environment.
Go with option 2: create a dedicated management cluster for all your shared services (InfluxDB/Grafana, log stack, docker-mailserver, Sentry, GitLab). For your log solution, consider either the ELK Stack (Elasticsearch, Logstash, Kibana) or Loki + Grafana—both integrate seamlessly with Kubernetes and work well with your existing monitoring stack.
In OpenTelekomCloud, set up VPC peering between your management cluster and the three application clusters, configure network policies to secure traffic, and reconfigure Heapster in each application cluster to send metrics to the centralized InfluxDB.
内容的提问来源于stack exchange,提问作者Marius

