kubeadm init执行失败:基于Vagrantfile的K8s集群部署问题求助
kubeadm init Failure in Your Vagrant Kubernetes Cluster Hey there, let’s work through why your kubeadm init is failing with that Vagrant-based 1-master/2-worker cluster setup. I’ve debugged similar issues with Vagrant Kubernetes clusters before, so here’s a structured approach to get this sorted:
1. Double-Check Kubernetes Prerequisites First
Your script_install_common_software is based on official docs, but let’s confirm all critical prerequisites are fully met on the master node (where kubeadm init runs):
- Swap must be disabled: Kubernetes won’t run with swap enabled. Verify with
sudo swapon --show— if you see output, runsudo swapoff -aand comment out any swap lines in/etc/fstabto make this permanent. - Container runtime is healthy: Ensure Docker, containerd, or your chosen runtime is installed and running. Check status with
systemctl status containerdorsystemctl status docker. For containerd, make sure itsconfig.tomlhas CRI support enabled (older versions might need you to uncomment the CRI plugin section). - IP forwarding is active: Run
sysctl net.ipv4.ip_forward— it should return1. If not, runsudo sysctl -w net.ipv4.ip_forward=1and add this setting to/etc/sysctl.d/k8s.confto persist across reboots. - Required ports are open: On Ubuntu-based VMs, use
sudo ufw allow 6443/tcp(kube-apiserver),sudo ufw allow 2379-2380/tcp(etcd),sudo ufw allow 10250/tcp(kubelet), andsudo ufw allow 10251/tcp(kube-scheduler) to unblock critical Kubernetes ports.
2. Dig Into kubeadm init Logs for Exact Errors
The fastest way to pinpoint the issue is to check the detailed logs:
- Watch the kubelet in real-time while re-running
kubeadm initwithsudo journalctl -u kubelet -f. - If you’ve already attempted init, check the full init log at
/var/log/kubeadm-init.log— this file captures every step of the init process, including failures.
Common log red flags to look for:
- etcd startup failures: Usually happens if the master node’s hostname/IP isn’t resolvable, or etcd ports are blocked.
- CRI runtime unresponsive: The kubelet can’t connect to your container runtime — double-check the runtime is properly configured with CRI support.
- API server issues: Often caused by an incorrect
--apiserver-advertise-addressflag. Make sure you’re using the master node’s private Vagrant network IP (not localhost or a public IP).
3. Validate SSH Key Setup Between Nodes
Your script_copy_key is supposed to enable passwordless SSH from master to workers, but let’s confirm it’s working:
- On the master node, test SSH access to each worker with
ssh worker1(or your worker hostnames). If you get a password prompt, the key copy script didn’t execute correctly. - Fix the key copy step: Ensure the script uses
ssh-copy-idwith the default Vagrant user (usuallyvagrant), and that worker nodes have password auth enabled temporarily during the copy process (Vagrant’s default setup allows password-based SSH for thevagrantuser).
4. Ensure Consistent Kubernetes Versions Across All Nodes
Mismatched versions of kubeadm, kubelet, and kubectl are a common cause of init failures.
- Run
kubeadm versionon the master andkubelet --versionon each worker to confirm all versions match. - If there’s a mismatch, re-run your
script_install_common_softwareon all nodes to sync versions.
5. Reset and Re-Init Cleanly
If you’ve made multiple init attempts, leftover cluster state can cause weird issues:
- On the master, run
sudo kubeadm reset -fto wipe all previous init state. - Delete any leftover kubectl config with
rm -rf $HOME/.kube. - Re-run
kubeadm initwith the correct flags — for example, if you’re using Flannel as your CNI plugin, add--pod-network-cidr=10.244.0.0/16to the command.
Once kubeadm init succeeds, don’t forget to apply your chosen CNI plugin before joining the worker nodes!
内容的提问来源于stack exchange,提问作者Pablo Toledo

