You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于kubeadm的单Master节点Kubernetes集群故障恢复咨询

Master Node Recovery Without Disrupting Worker Deployments (Kubeadm Cluster)

Great question—restoring a single master node without interrupting worker-side pods is totally doable, but it depends on proper pre-failure backups and sticking to a process that preserves your cluster's state and identity. Let’s break this down clearly:

First: Yes, You Must Back Up Original Certificates

Your worker nodes’ kubelet, kube-proxy, and other components use certificates signed by the original master’s CA. If you skip restoring these certificates and generate a new CA on the replacement master, workers will reject the new master’s identity—forcing you to re-join workers, which would disrupt running pods. Preserving the original PKI chain is non-negotiable for zero-disruption recovery.

Critical Pre-Failure Backups You Need

Before any failure occurs, make sure you’ve backed up these key components:

  • ETCD Snapshot: This stores your entire cluster state (deployments, pods, configs, etc.). Take a snapshot with:
    ETCDCTL_API=3 etcdctl snapshot save /path/to/cluster-snapshot.db \
      --endpoints=https://127.0.0.1:2379 \
      --cacert=/etc/kubernetes/pki/etcd/ca.crt \
      --cert=/etc/kubernetes/pki/etcd/server.crt \
      --key=/etc/kubernetes/pki/etcd/server.key
    
  • Kubernetes PKI Certificates: Copy the entire /etc/kubernetes/pki/ directory—this includes the root CA, API server certificates, etcd certificates, and more.
  • Kubeconfig Files: Save /etc/kubernetes/admin.conf, /etc/kubernetes/kubelet.conf, /etc/kubernetes/controller-manager.conf, and /etc/kubernetes/scheduler.conf—these are needed for authenticating cluster components.
  • Kubeadm Cluster Config: Export your cluster’s configuration to ensure consistency during recovery:
    kubeadm config view > /path/to/kubeadm-config.yaml
    

Step-by-Step Recovery (No Worker Pod Disruption)

  1. Clean up the failed master (if possible)
    If the old master is still reachable, remove it from the cluster first:

    kubectl delete node <old-master-node-name>
    

    If it’s completely unresponsive, you can skip this step—etcd will sync the node status once the new master is online.

  2. Prepare the replacement master node
    Install the exact same version of kubeadm, kubelet, and kubectl as your existing cluster. Version mismatches will cause compatibility issues.

  3. Restore backed-up files
    Copy your saved PKI certificates, kubeconfig files, and kubeadm config to the new master’s corresponding directories:

    # Restore certificates
    cp -r /path/to/backup/pki/ /etc/kubernetes/
    # Restore kubeconfigs
    cp /path/to/backup/*.conf /etc/kubernetes/
    # Set correct permissions (critical for security)
    chmod -R 600 /etc/kubernetes/pki/
    chown -R root:root /etc/kubernetes/pki/
    
  4. Restore ETCD from your snapshot
    Use kubeadm to restore the etcd cluster state:

    kubeadm init phase etcd restore \
      --config=/path/to/kubeadm-config.yaml \
      --snapshot-path=/path/to/cluster-snapshot.db
    
  5. Re-initialize the master node
    Initialize the new master using your saved cluster config—this ensures it matches the original cluster’s network settings (Pod CIDR, Service CIDR, etc.):

    kubeadm init --config=/path/to/kubeadm-config.yaml \
      --ignore-preflight-errors=DirAvailable--var-lib-etcd
    
  6. Reconnect worker nodes (if needed)

    • If your master uses a static VIP or domain name (highly recommended!), workers will automatically reconnect to the new master without any changes—their kubelet certificates are still trusted by the restored CA, and their kubeconfig points to the stable address.
    • If the master’s IP changed, update the server field in each worker’s /etc/kubernetes/kubelet.conf to the new master’s address, then restart kubelet:
      systemctl restart kubelet
      

    In either case, running pods on workers will keep running—they only lose temporary connectivity to the master, not their runtime state.

Key Tips to Avoid Disruptions

  • Use a VIP or domain name for your master instead of a static IP—this eliminates the need to update worker configs if you replace the master node.
  • Schedule regular automated backups for etcd and PKI files (e.g., with cron jobs) to avoid manual backup errors.
  • Never skip version matching—kubeadm is strict about version consistency between control plane and worker nodes.

内容的提问来源于stack exchange,提问作者fstab

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:12:59