进程被OOM Killer终止时查看Linux内存使用情况的简便方法
Great question—this is such a common headache with busy build servers. OOM Killer often takes out a process that’s just in the wrong place at the wrong time, not necessarily the actual memory hog. Here are a few proven ways to capture a full process memory snapshot exactly when OOM strikes, so you can pinpoint the real culprit:
If your server runs systemd 245 or newer (most recent distros like RHEL 8+, Ubuntu 20.04+), systemd-oomd is built-in and can be configured to trigger custom actions when OOM is detected.
First, make sure
systemd-oomdis enabled and running:systemctl enable --now systemd-oomdCreate a shell script to capture process data (save it as
/usr/local/bin/capture_oom_snapshot.sh):#!/bin/bash SNAPSHOT_DIR="/var/log/oom_snapshots" mkdir -p "$SNAPSHOT_DIR" TIMESTAMP=$(date +%Y%m%d_%H%M%S) OUTPUT_FILE="$SNAPSHOT_DIR/oom_snapshot_$TIMESTAMP.log" # Capture full process list with complete command lines echo "=== OOM TRIGGERED AT $TIMESTAMP ===" > "$OUTPUT_FILE" echo -e "\n--- FULL PROCESS LIST (ps auxww) ---" >> "$OUTPUT_FILE" ps auxww >> "$OUTPUT_FILE" # Capture detailed memory stats (install smem first if you can: apt install smem / dnf install smem) echo -e "\n--- DETAILED MEMORY USAGE (smem) ---" >> "$OUTPUT_FILE" smem -t -k -u >> "$OUTPUT_FILE" 2>/dev/null || echo "smem not installed - skipping this section" >> "$OUTPUT_FILE" # Grab recent dmesg context for OOM details echo -e "\n--- RECENT DMESG LOGS ---" >> "$OUTPUT_FILE" dmesg | tail -60 >> "$OUTPUT_FILE" echo "OOM snapshot saved to $OUTPUT_FILE"Make the script executable:
chmod +x /usr/local/bin/capture_oom_snapshot.shConfigure
systemd-oomdto run this script when it detects OOM. Create a drop-in config at/etc/systemd/oomd.conf.d/99-capture-snapshot.conf:[OOM] DefaultMemoryPressureLimit=50% DefaultSwapPressureLimit=100% # Run our snapshot script when OOM is triggered OnOOM=capture-oom-snapshot.serviceCreate a systemd service for the script (
/etc/systemd/system/capture-oom-snapshot.service):[Unit] Description=Capture process snapshot on OOM [Service] Type=oneshot ExecStart=/usr/local/bin/capture_oom_snapshot.sh User=rootReload systemd and restart
systemd-oomd:systemctl daemon-reload systemctl restart systemd-oomd
If you’re on an older system without systemd-oomd, you can run a background script that listens for OOM events in dmesg and triggers the snapshot.
Use this script (save as
/usr/local/bin/monitor_oom.sh):#!/bin/bash SNAPSHOT_DIR="/var/log/oom_snapshots" mkdir -p "$SNAPSHOT_DIR" # Listen to real-time dmesg output dmesg -w | while read -r line; do # Detect the OOM Killer trigger line if echo "$line" | grep -q "Out of memory:"; then echo "OOM detected - capturing snapshot..." TIMESTAMP=$(date +%Y%m%d_%H%M%S) OUTPUT_FILE="$SNAPSHOT_DIR/oom_snapshot_$TIMESTAMP.log" # Capture process data echo "=== OOM TRIGGERED AT $TIMESTAMP ===" > "$OUTPUT_FILE" echo -e "\n--- FULL PROCESS LIST (ps auxww) ---" >> "$OUTPUT_FILE" ps auxww >> "$OUTPUT_FILE" echo -e "\n--- TOP MEMORY HOGS (top -b -n1 -o %MEM) ---" >> "$OUTPUT_FILE" top -b -n1 -o %MEM >> "$OUTPUT_FILE" echo -e "\n--- DMESG OOM CONTEXT ---" >> "$OUTPUT_FILE" dmesg | tail -50 >> "$OUTPUT_FILE" echo "Snapshot saved to $OUTPUT_FILE" fi doneMake it executable:
chmod +x /usr/local/bin/monitor_oom.shCreate a systemd service to run it in the background (
/etc/systemd/system/oom-monitor.service):[Unit] Description=Monitor for OOM events and capture snapshots [Service] ExecStart=/usr/local/bin/monitor_oom.sh Restart=always User=root [Install] WantedBy=multi-user.targetEnable and start the service:
systemctl enable --now oom-monitor.service
For a quick, no-extra-tools solution, you can enable a kernel parameter that makes the OOM Killer dump all process memory stats to dmesg when it triggers.
Enable the parameter temporarily:
sysctl -w vm.oom_dump_tasks=1To make it permanent, add this to
/etc/sysctl.conf:vm.oom_dump_tasks=1Then run
sysctl -pto apply.
This will make dmesg include a list of all processes with their memory usage (RSS, VmSize, etc.) when OOM strikes. The downside is it’s less detailed than a custom snapshot (no full command lines by default), but it’s a zero-effort way to get basic data.
- Use
smeminstead ofps:smemcalculates memory usage more accurately by accounting for shared memory (PSS instead of RSS). It’s much better at identifying real memory hogs on systems with shared libraries or containers. - Limit build task memory: Use cgroups, Docker resource limits, or your build tool’s built-in memory constraints (e.g., Maven’s
MAVEN_OPTS, Node.js’s--max-old-space-size) to prevent individual builds from consuming too much memory. - Monitor memory trends: Set up a monitoring tool like Prometheus + Node Exporter to track memory usage over time. This can help you spot memory leaks or gradual memory creep before OOM strikes.
内容的提问来源于stack exchange,提问作者Hedley

