Kubernetes服务器宕机数小时,出现INFO: task等报错如何解决?
Hey there, let's dig into your issue—those two errors (docker0 network not ready + task blocked for 120s) are almost certainly linked, so we'll tackle them systematically.
First: Diagnose the Docker0 Network Issue
The IPv6: ADDRCONF(NETDEV_UP): docker0: link is not ready error suggests the docker bridge interface failed to initialize properly. This can break Kubernetes container networking, which in turn causes resource-hungry retries that lead to the task blocking warnings.
Quickly check docker0 status (do this within 15 seconds of logging in before the system freezes):
ip link show docker0Look for the
UPflag in the output. If it's missing, the interface is down.Manually bring up docker0 to test if this resolves the immediate freeze:
ip link set docker0 upWait a minute and see if the system stays responsive. If it does, the network interface initialization is the root cause.
Check Docker's IPv6 configuration:
Open/etc/docker/daemon.jsonand look for settings like"ipv6": false. If IPv6 is disabled but your system or Kubernetes expects it, this can cause the docker0 initialization failure. You can either:- Enable IPv6 in Docker (add
"ipv6": trueand a fixed-cidr-v6 if needed), or - Ensure Kubernetes components are configured to work without IPv6 (e.g., set
--disable-ipv6flags on kubelet if applicable).
- Enable IPv6 in Docker (add
Second: Resolve the Task Blocking Freeze
The INFO: task [TASK]:[PID] blocked for more than 120 seconds message is the kernel warning about a process stuck in an uninterruptible sleep state—usually due to resource deadlocks or failed I/O (often network-related here).
Identify the stuck process immediately after logging in:
Runtoporhtopto spot processes with high CPU/memory usage. Focus on kubelet, docker, or network plugin processes (like flanneld, calico-node). You can also check the kernel log for details:dmesg | grep "blocked for more than 120 seconds"This will show you exactly which task is stuck.
Restart critical services:
If docker is misbehaving, restart it first, then kubelet:systemctl restart docker systemctl restart kubeletThis often clears up stuck network-related processes.
Check for resource exhaustion:
Run these commands to rule out disk or memory issues (common culprits for freezes):df -h # Check disk space, especially /var/lib/docker and /var/lib/kubelet free -m # Check available memoryIf disk is full, clean up unused containers/images with
docker system prune -a(be cautious with this command—only run it if you know you can delete unused resources).
Third: Prevent Future Issues
Once you've fixed the immediate problem, take these steps to avoid recurrence:
Verify Kubernetes network plugin health:
If you're using a CNI plugin like Flannel or Calico, check if their pods are running normally (run this before the system freezes):kubectl get pods -n kube-systemRestart any failed network plugin pods if needed.
Set critical sysctl parameters permanently:
Ensure IP forwarding is enabled (required for Kubernetes networking):sysctl -w net.ipv4.ip_forward=1Then add
net.ipv4.ip_forward=1to/etc/sysctl.confso it persists across reboots.Update Docker/Kubernetes components:
Outdated versions can have known networking bugs. Consider upgrading to stable, supported versions of Docker and Kubernetes if you're running older releases.
内容的提问来源于stack exchange,提问作者Kerren

