You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

进程被OOM Killer终止时查看Linux内存使用情况的简便方法

Great question—this is such a common headache with busy build servers. OOM Killer often takes out a process that’s just in the wrong place at the wrong time, not necessarily the actual memory hog. Here are a few proven ways to capture a full process memory snapshot exactly when OOM strikes, so you can pinpoint the real culprit:

方案1:用systemd-oomd(推荐,适合现代systemd系统)

If your server runs systemd 245 or newer (most recent distros like RHEL 8+, Ubuntu 20.04+), systemd-oomd is built-in and can be configured to trigger custom actions when OOM is detected.

  1. First, make sure systemd-oomd is enabled and running:

    systemctl enable --now systemd-oomd
    
  2. Create a shell script to capture process data (save it as /usr/local/bin/capture_oom_snapshot.sh):

    #!/bin/bash
    SNAPSHOT_DIR="/var/log/oom_snapshots"
    mkdir -p "$SNAPSHOT_DIR"
    TIMESTAMP=$(date +%Y%m%d_%H%M%S)
    OUTPUT_FILE="$SNAPSHOT_DIR/oom_snapshot_$TIMESTAMP.log"
    
    # Capture full process list with complete command lines
    echo "=== OOM TRIGGERED AT $TIMESTAMP ===" > "$OUTPUT_FILE"
    echo -e "\n--- FULL PROCESS LIST (ps auxww) ---" >> "$OUTPUT_FILE"
    ps auxww >> "$OUTPUT_FILE"
    
    # Capture detailed memory stats (install smem first if you can: apt install smem / dnf install smem)
    echo -e "\n--- DETAILED MEMORY USAGE (smem) ---" >> "$OUTPUT_FILE"
    smem -t -k -u >> "$OUTPUT_FILE" 2>/dev/null || echo "smem not installed - skipping this section" >> "$OUTPUT_FILE"
    
    # Grab recent dmesg context for OOM details
    echo -e "\n--- RECENT DMESG LOGS ---" >> "$OUTPUT_FILE"
    dmesg | tail -60 >> "$OUTPUT_FILE"
    
    echo "OOM snapshot saved to $OUTPUT_FILE"
    
  3. Make the script executable:

    chmod +x /usr/local/bin/capture_oom_snapshot.sh
    
  4. Configure systemd-oomd to run this script when it detects OOM. Create a drop-in config at /etc/systemd/oomd.conf.d/99-capture-snapshot.conf:

    [OOM]
    DefaultMemoryPressureLimit=50%
    DefaultSwapPressureLimit=100%
    # Run our snapshot script when OOM is triggered
    OnOOM=capture-oom-snapshot.service
    
  5. Create a systemd service for the script (/etc/systemd/system/capture-oom-snapshot.service):

    [Unit]
    Description=Capture process snapshot on OOM
    
    [Service]
    Type=oneshot
    ExecStart=/usr/local/bin/capture_oom_snapshot.sh
    User=root
    
  6. Reload systemd and restart systemd-oomd:

    systemctl daemon-reload
    systemctl restart systemd-oomd
    
方案2:自定义内核事件监听脚本

If you’re on an older system without systemd-oomd, you can run a background script that listens for OOM events in dmesg and triggers the snapshot.

  1. Use this script (save as /usr/local/bin/monitor_oom.sh):

    #!/bin/bash
    SNAPSHOT_DIR="/var/log/oom_snapshots"
    mkdir -p "$SNAPSHOT_DIR"
    
    # Listen to real-time dmesg output
    dmesg -w | while read -r line; do
        # Detect the OOM Killer trigger line
        if echo "$line" | grep -q "Out of memory:"; then
            echo "OOM detected - capturing snapshot..."
            TIMESTAMP=$(date +%Y%m%d_%H%M%S)
            OUTPUT_FILE="$SNAPSHOT_DIR/oom_snapshot_$TIMESTAMP.log"
    
            # Capture process data
            echo "=== OOM TRIGGERED AT $TIMESTAMP ===" > "$OUTPUT_FILE"
            echo -e "\n--- FULL PROCESS LIST (ps auxww) ---" >> "$OUTPUT_FILE"
            ps auxww >> "$OUTPUT_FILE"
            echo -e "\n--- TOP MEMORY HOGS (top -b -n1 -o %MEM) ---" >> "$OUTPUT_FILE"
            top -b -n1 -o %MEM >> "$OUTPUT_FILE"
            echo -e "\n--- DMESG OOM CONTEXT ---" >> "$OUTPUT_FILE"
            dmesg | tail -50 >> "$OUTPUT_FILE"
    
            echo "Snapshot saved to $OUTPUT_FILE"
        fi
    done
    
  2. Make it executable:

    chmod +x /usr/local/bin/monitor_oom.sh
    
  3. Create a systemd service to run it in the background (/etc/systemd/system/oom-monitor.service):

    [Unit]
    Description=Monitor for OOM events and capture snapshots
    
    [Service]
    ExecStart=/usr/local/bin/monitor_oom.sh
    Restart=always
    User=root
    
    [Install]
    WantedBy=multi-user.target
    
  4. Enable and start the service:

    systemctl enable --now oom-monitor.service
    
方案3:内核参数自动dump进程信息

For a quick, no-extra-tools solution, you can enable a kernel parameter that makes the OOM Killer dump all process memory stats to dmesg when it triggers.

  1. Enable the parameter temporarily:

    sysctl -w vm.oom_dump_tasks=1
    
  2. To make it permanent, add this to /etc/sysctl.conf:

    vm.oom_dump_tasks=1
    

    Then run sysctl -p to apply.

This will make dmesg include a list of all processes with their memory usage (RSS, VmSize, etc.) when OOM strikes. The downside is it’s less detailed than a custom snapshot (no full command lines by default), but it’s a zero-effort way to get basic data.

额外实用技巧
  • Use smem instead of ps: smem calculates memory usage more accurately by accounting for shared memory (PSS instead of RSS). It’s much better at identifying real memory hogs on systems with shared libraries or containers.
  • Limit build task memory: Use cgroups, Docker resource limits, or your build tool’s built-in memory constraints (e.g., Maven’s MAVEN_OPTS, Node.js’s --max-old-space-size) to prevent individual builds from consuming too much memory.
  • Monitor memory trends: Set up a monitoring tool like Prometheus + Node Exporter to track memory usage over time. This can help you spot memory leaks or gradual memory creep before OOM strikes.

内容的提问来源于stack exchange,提问作者Hedley

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:24:46