如何在Pod内部查看Pod终止原因?(基于Kubernetes 1.15.3)
Great question! In Kubernetes 1.15.3, there isn't a native, direct way for a running Pod to retrieve its own termination reason from inside the container, but there are several indirect approaches you can use to infer or capture this information—especially for OOM (Out of Memory) scenarios. Let's break them down:
1. Capture Termination Signals in Your Container's Entrypoint
When Kubernetes decides to terminate a Pod (like when memory exceeds limits), it first sends a SIGTERM signal to the main process. If the process doesn't exit within the grace period (default 30s), Kubernetes sends a SIGKILL (which can't be caught).
You can modify your container's entrypoint script to trap SIGTERM and log it, giving you visibility into the termination trigger. Here's an example bash script:
#!/bin/bash # Define a handler for SIGTERM handle_termination() { echo "$(date): Received SIGTERM signal. Initiating graceful shutdown." >> /var/log/pod-termination.log # Add any cleanup logic here (e.g., closing connections, saving state) exit 0 } # Trap the SIGTERM signal trap handle_termination SIGTERM # Start your main application process exec your-main-application-command "$@"
If your process gets killed by SIGKILL (resulting in exit code 137), you won't be able to trap this signal, but you can check the exit code of the process in your script (if using a wrapper) to infer it was force-killed—often a sign of OOM or grace period expiration.
2. Monitor Resource Usage to Predict OOM Events
Since OOM is a common termination reason, you can add logic inside your container to monitor memory usage relative to its limits, and log warnings before termination.
You can read the memory limit and current usage from the cgroup filesystem (available inside most containers):
# Get memory limit (in bytes) MEMORY_LIMIT=$(cat /sys/fs/cgroup/memory/memory.limit_in_bytes) # Get current memory usage (in bytes) MEMORY_USAGE=$(cat /sys/fs/cgroup/memory/memory.usage_in_bytes) # Calculate usage percentage USAGE_PERCENT=$(( (MEMORY_USAGE * 100) / MEMORY_LIMIT )) if [ $USAGE_PERCENT -gt 90 ]; then echo "$(date): Memory usage is at ${USAGE_PERCENT}%, approaching limit of ${MEMORY_LIMIT} bytes." >> /var/log/memory-monitor.log fi
You can run this check periodically (e.g., via cron or a background loop in your entrypoint) to log when memory is getting close to the limit—this helps confirm later that termination was likely due to OOM.
3. Infer Termination Reason from Exit Codes
After the Pod has terminated, if you can access the container's logs (even post-termination), you can use the exit code to narrow down the cause:
- Exit code
137: This equals128 + 9, where 9 is the signal number forSIGKILL. This usually means the process was force-killed, either because it didn't respond toSIGTERMin time, or because the node's OOM killer terminated it due to memory exhaustion. - Exit code
143: Equals128 + 15, which isSIGTERM—indicates the process received the graceful termination signal and exited normally.
Keep in mind that while these methods don't give you a direct "termination reason" field inside the Pod, they let you gather enough context to understand why the Pod was terminated.
内容的提问来源于stack exchange,提问作者user11779620

