如何解决k3d集群中Cilium的"Missed tail calls"丢包问题
问题:k3d集群集成Cilium后出现连接丢包(Missed tail call)
我在k3d集群中集成了Cilium,但系统频繁出现如下丢包日志:
root@k3d-k8s-daemon01-dev-local01-agent-0:/home/cilium# cilium monitor -t drop Listening for events on 16 CPUs with 64x4096 of shared memory Press Ctrl-C to quit level=info msg="Initializing dissection cache..." subsys=monitor xx drop (Missed tail call) flow 0xc75cc925 to endpoint 0, ifindex 3, file bpf_host.c:794, , identity world->unknown: 10.43.0.1:443 -> 10.42.0.92:35152 tcp SYN, ACK xx drop (Missed tail call) flow 0x6933b0fe to endpoint 0, ifindex 3, file bpf_host.c:794, , identity world->unknown: 10.43.0.1:443 -> 10.42.0.92:35152 tcp SYN, ACK xx drop (Missed tail call) flow 0x877a22ae to endpoint 0, ifindex 3, file bpf_host.c:794, , identity world->unknown: 10.43.0.1:443 -> 10.42.0.92:35152 tcp SYN, ACK xx drop (Missed tail call) flow 0xdbbeb5ee to endpoint 0, ifindex 3, file bpf_host.c:794, , identity world->unknown: 10.43.0.1:443 -> 10.42.0.92:35152 tcp SYN, ACK
部分连接能正常工作,但K8s到部分内部Pod的连接无法建立。
使用的Cilium版本:
root@k3d-k8s-daemon01-dev-local01-agent-0:/home/cilium# cilium version Client: 1.14.0-rc.1 53f97a7b 2023-07-17T00:45:13-07:00 go version go1.20.5 linux/amd64 Daemon: 1.14.0-rc.1 53f97a7b 2023-07-17T00:45:13-07:00 go version go1.20.5 linux/amd64
集群部署步骤:
export CLUSTERNAME=k8s-daemon01-dev-local01 k3d cluster create $CLUSTERNAME \ -a 1 \ ... \ --image rancher/k3s:v1.27.3-k3s1 # fixes bpf for k3d: docker exec -it k3d-$CLUSTERNAME-agent-0 mount bpffs /sys/fs/bpf -t bpf docker exec -it k3d-$CLUSTERNAME-agent-0 mount --make-shared /sys/fs/bpf docker exec -it k3d-$CLUSTERNAME-server-0 mount bpffs /sys/fs/bpf -t bpf docker exec -it k3d-$CLUSTERNAME-server-0 mount --make-shared /sys/fs/bpf # this deploys Cilium -- in a quite std way: ansible-playbook site.yml -i inventory/k8s-daemon01-dev-local01/hosts.ini -t cilium # wait till cilium-operator is ready kubectl wait --for=condition=Ready pod -l app.kubernetes.io/name=cilium-operator -n kube-system --timeout=300s # wait a bit more for cilium to be ready sleep 30 # fixes bpf for cilium: kubectl get nodes -o custom-columns=NAME:.metadata.name --no-headers=true | xargs -I {} docker exec {} mount bpffs /sys/fs/bpf -t bpf kubectl get nodes -o custom-columns=NAME:.metadata.name --no-headers=true | xargs -I {} docker exec {} mount --make-shared /sys/fs/bpf kubectl get nodes -o custom-columns=NAME:.metadata.name --no-headers=true | xargs -I {} docker exec {} mount --make-shared /run/cilium/cgroupv2
原因分析与解决方法
核心原因:BPF尾调用缺失(Missed tail call)
日志中的Missed tail call错误,说明Cilium的BPF程序处理流量时无法找到对应尾调用程序,常见触发场景:
- BPF挂载点未正确配置或未持久化共享,容器重启后挂载失效
- k3s默认CNI(flannel)未禁用,与Cilium的BPF程序冲突
- 使用的Cilium候选版本(1.14.0-rc.1)存在与k3s 1.27.3的兼容性bug
具体解决步骤
1. 彻底禁用k3s默认CNI
重新创建k3d集群时必须禁用flannel,确保Cilium作为唯一CNI运行:
export CLUSTERNAME=k8s-daemon01-dev-local01 k3d cluster create $CLUSTERNAME \ -a 1 \ --image rancher/k3s:v1.27.3-k3s1 \ --k3s-arg "--flannel-backend=none@server:*" \ --k3s-arg "--disable-network-policy@server:*"
注:
--flannel-backend=none禁用k3s自带flannel,--disable-network-policy避免与Cilium网络策略冲突
2. 持久化BPF挂载
手动挂载重启后会失效,创建集群时直接映射主机BPF文件系统到容器:
k3d cluster create $CLUSTERNAME \ ... --volume /sys/fs/bpf:/sys/fs/bpf@all
3. 替换为Cilium稳定版
候选版本存在未知风险,降级到与k3s 1.27.x兼容的1.13.x稳定版:
helm repo add cilium https://helm.cilium.io/ helm install cilium cilium/cilium --version 1.13.6 \ --namespace kube-system \ --set kubeProxyReplacement=strict
4. 验证修复效果
部署完成后检查组件状态和BPF加载情况:
# 检查Cilium整体状态 cilium status # 验证BPF程序加载 cilium bpf list # 重新监控丢包事件 cilium monitor -t drop
5. 现有集群修复(无需重建)
若不想重建集群,先清理默认CNI再重启Cilium:
# 删除flannel相关资源 kubectl delete daemonset kube-flannel-ds -n kube-system # 重启Cilium daemonset kubectl rollout restart daemonset cilium -n kube-system # 强制重载BPF程序 kubectl exec -n kube-system -l k8s-app=cilium-agent -- cilium bpf reload
内容的提问来源于stack exchange,提问作者towi
相关产品推荐
相关产品推荐

