EC2上Kubernetes集群6443端口连接时断时续问题求助
Kubernetes集群6443端口时断时续排查求助
我在EC2实例上搭建1台控制节点+2台工作节点的Kubernetes集群,按照官方指南操作后,初期执行kubectl get nodes能得到如下结果:
NAME STATUS ROLES AGE VERSION ip-172-31-13-48 NotReady control-plane 12m v1.27.2 ip-172-31-25-211 NotReady <none> 11m v1.27.2 ip-172-31-26-181 NotReady <none> 11m v1.27.2
但几分钟后再次执行kubectl get nodes时,出现错误:
The connection to the server 172.31.13.48:6443 was refused - did you specify the right host or port?
通过netstat -tlnp对比6443端口状态:
异常时输出:
Active Internet connections (only servers) Proto Recv-Q Send-Q Local Address Foreign Address State PID/Program name tcp 0 0 127.0.0.1:10259 0.0.0.0:* LISTEN 14044/kube-schedule tcp 0 0 127.0.0.1:10248 0.0.0.0:* LISTEN 6343/kubelet tcp 0 0 127.0.0.1:43423 0.0.0.0:* LISTEN 555/containerd tcp 0 0 127.0.0.53:53 0.0.0.0:* LISTEN 404/systemd-resolve tcp 0 0 0.0.0.0:22 0.0.0.0:* LISTEN 757/sshd: /usr/sbin tcp6 0 0 :::22 :::* LISTEN 757/sshd: /usr/sbin tcp6 0 0 :::10250 :::* LISTEN 6343/kubelet
正常时输出:
(Not all processes could be identified, non-owned process info will not be shown, you would have to be root to see it all.) Active Internet connections (only servers) Proto Recv-Q Send-Q Local Address Foreign Address State PID/Program name tcp 0 0 172.31.13.48:2380 0.0.0.0:* LISTEN - tcp 0 0 172.31.13.48:2379 0.0.0.0:* LISTEN - tcp 0 0 127.0.0.1:10259 0.0.0.0:* LISTEN - tcp 0 0 127.0.0.1:10248 0.0.0.0:* LISTEN - tcp 0 0 127.0.0.1:2379 0.0.0.0:* LISTEN - tcp 0 0 127.0.0.1:2381 0.0.0.0:* LISTEN - tcp 0 0 127.0.0.1:43423 0.0.0.0:* LISTEN - tcp 0 0 127.0.0.53:53 0.0.0.0:* LISTEN - tcp 0 0 0.0.0.0:22 0.0.0.0:* LISTEN - tcp6 0 0 :::22 :::* LISTEN - tcp6 0 0 :::10250 :::* LISTEN - tcp6 0 0 :::6443 :::* LISTEN -
我完全不清楚6443端口为何会时断时续,恳请帮忙排查原因。
排查建议:
- 检查kube-apiserver服务状态与日志:
执行systemctl status kube-apiserver查看服务是否频繁启停,用journalctl -u kube-apiserver -f实时追踪日志,定位崩溃或退出的具体原因。 - 验证EC2实例资源情况:
用top或htop查看CPU、内存占用,kube-apiserver可能因资源不足被OOM Killer终止,可通过dmesg | grep -i oom确认是否存在OOM记录。 - 检查kube-apiserver配置文件:
查看/etc/kubernetes/manifests/kube-apiserver.yaml,确认绑定地址、证书路径、etcd连接参数等配置是否正确,错误配置会导致服务无法稳定运行。 - 确认etcd服务状态:
从端口对比看,异常时etcd的2379/2380端口也消失,说明etcd可能同步故障或停止,执行systemctl status etcd查看状态,用journalctl -u etcd -f排查日志。 - 检查防火墙与安全组规则:
确认EC2安全组是否允许6443端口入站,控制节点的iptables规则是否有变动,执行iptables-save | grep 6443检查相关规则是否正常。
内容的提问来源于stack exchange,提问作者aviv levari
相关产品推荐
相关产品推荐

