如何记录Kubernetes因资源占用杀死进程的事件?含非主进程SIGKILL场景
Great question! Let's walk through how to track and log when Kubernetes (or the node's kernel) kills processes due to resource constraints, including those non-main processes getting SIGKILLed.
Kubernetes explicitly terminates containers when they exceed their configured resources.limits (CPU or memory). Here's how to capture these events:
Check Pod Events: Every time Kubernetes kills a container due to OOM, it logs an event tied to the pod. View this with:
kubectl describe pod <your-pod-name>Look for the Events section—you’ll see a warning entry like:
Warning OOMKill 10m kubelet Container
in pod was terminated due to OOMKilled To watch events in real time across the cluster:
kubectl get events --watch --sort-by='.metadata.creationTimestamp'Inspect Kubelet Logs: The kubelet (the agent running on each node) logs detailed info about OOM kills. On most modern nodes, access these logs via:
# For systemd-based distros (Ubuntu, RHEL, etc.) journalctl -u kubelet -f # Or check the log file directly if applicable tail -f /var/log/kubelet.logYou’ll find entries specifying the pod/container name, terminated process, and the resource limit that was exceeded.
If a non-main process inside the container gets killed by the node’s kernel OOM killer (because the container’s memory usage strains node-wide limits, or the kernel prioritizes freeing up memory), Kubernetes won’t log this as a pod event. Here’s how to track these cases:
Check Node-Level dmesg Logs: The kernel writes OOM kill details to the system’s dmesg buffer. Run this on the node where the pod is running:
dmesg | grep -i oomYou’ll see entries like:
[12345.678901] oom-kill:constraint=CONSTRAINT_MEMCG,nodemask=(null),cpuset=...,mems_allowed=0,oom_memcg=/kubepods/...,task_memcg=/kubepods/...,task=my-non-main-process,pid=1234,uid=0
[12345.678905] Out of memory: Killed process 1234 (my-non-main-process) total-vm:123456kB, anon-rss:78900kB, file-rss:1234kB, shmem-rss:0kBCapture Kernel Logs Inside Containers: If you need to collect these logs directly from within the container (for easier aggregation), mount the
/proc/kmsgfile into the container (requires privileged access). Add this to your pod spec:volumes: - name: kmsg hostPath: path: /proc/kmsg containers: - name: your-container volumeMounts: - name: kmsg mountPath: /proc/kmsg readOnly: trueThen use a tool like
dmesgor a log shipper to read and forward these logs to your central logging system.
For scalable tracking across a cluster, set up a centralized logging stack (like ELK, Loki, or Splunk) to collect:
- Kubernetes Events (use tools like
kube-state-metricsorfluentdto scrape events) - Kubelet logs
- Node dmesg logs
- Container runtime logs
For proactive monitoring and alerting, use Prometheus with these metrics:
kube_pod_container_status_terminated_reason{reason="OOMKilled"}: Tracks how many containers were killed by Kubernetes due to OOMnode_memory_MemAvailable_bytes: Monitors node-level memory to predict potential kernel OOM events
内容的提问来源于stack exchange,提问作者Chris Stryczynski

