You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Go进程被cgroup OOM Killer终止时收集堆内存信息?

解决Go进程被OOM Killer终止时的堆内存分析问题

在cgroup环境下遇到OOM Killer终止Go进程的情况,要定位内存占用Top项,我们可以从外部pprof采集和进程主动dump两个方向来实现,下面是具体方案:

一、利用外部pprof命令捕获堆信息(提前准备+预警触发)

需要注意的是,OOM Killer会直接向进程发送SIGKILL信号,进程无法捕获这个信号来触发pprof采集,因此我们需要提前监控内存使用,在进程被终止前主动触发采集:

  1. 在Go进程中开启pprof的HTTP端点
    在代码里加入以下片段,启动一个HTTP服务暴露pprof调试端点:
    import _ "net/http/pprof"
    import (
        "log"
        "net/http"
    )
    
    func main() {
        // 后台启动pprof服务,端口可自定义为你需要的数值
        go func() {
            if err := http.ListenAndServe("localhost:6060", nil); err != nil {
                log.Printf("pprof server failed: %v", err)
            }
        }()
        // 你的业务逻辑代码...
    }
    
  2. 结合cgroup内存阈值触发采集
    编写一个监控脚本,轮询cgroup的内存使用情况,当内存占用接近限制(比如达到90%)时,主动执行pprof命令导出堆profile:
    # 示例脚本:监控指定cgroup的内存使用,达到阈值时导出堆profile
    CGROUP_DIR="/sys/fs/cgroup/memory/your-target-cgroup"
    MEM_LIMIT=$(cat "${CGROUP_DIR}/memory.limit_in_bytes")
    # 设置预警阈值为内存限制的90%
    THRESHOLD=$((MEM_LIMIT * 90 / 100))
    
    while true; do
        CURRENT_USAGE=$(cat "${CGROUP_DIR}/memory.usage_in_bytes")
        if [ "$CURRENT_USAGE" -ge "$THRESHOLD" ]; then
            # 导出堆profile到带时间戳的文件,方便区分
            go tool pprof -output="heap_profile_$(date +%Y%m%d_%H%M%S).pprof" http://localhost:6060/debug/pprof/heap
            echo "Heap profile exported successfully"
            # 可选择只采集一次后退出,或者持续监控
            break
        fi
        sleep 3
    done
    
    执行这个脚本后,就能在进程被OOM Killer终止前拿到堆内存的Top占用项信息。

二、让Go进程自身主动dump堆内存信息(更直接的方案)

这种方式不需要外部依赖,进程自己就能在内存压力达到预警线时导出堆profile,精准定位内存问题:

1. 基于Go runtime的内存监控与dump

利用Go标准库的runtime和runtime/pprof包,定期检查堆内存使用,达到阈值时自动写入堆profile:

import (
    "log"
    "os"
    "runtime"
    "runtime/pprof"
    "time"
)

func memoryMonitor() {
    ticker := time.NewTicker(3 * time.Second)
    defer ticker.Stop()

    for range ticker.C {
        var memStats runtime.MemStats
        runtime.ReadMemStats(&memStats)
        // 这里可以根据你的cgroup内存限制调整阈值,示例为堆内存达到1GB时触发
        if memStats.HeapAlloc > 1024*1024*1024 {
            dumpFile, err := os.Create("heap_dump_" + time.Now().Format("20060102_150405") + ".pprof")
            if err != nil {
                log.Fatalf("Failed to create heap dump file: %v", err)
            }
            defer dumpFile.Close()
            
            if err := pprof.WriteHeapProfile(dumpFile); err != nil {
                log.Fatalf("Failed to write heap profile: %v", err)
            }
            log.Println("Heap profile dumped successfully")
            // 若只需一次dump,可在此处break退出监控
        }
    }
}

func main() {
    // 启动内存监控goroutine
    go memoryMonitor()
    // 你的业务逻辑代码...
}

2. 结合cgroup内存事件的精准触发

如果想更精准响应cgroup的内存压力,可以监听cgroup的memory.pressure_level或memory.oom_control文件,当收到内存压力通知时触发dump。你可以用inotify机制监控文件变化,或者轮询读取文件状态,根据返回的压力等级(比如high)来触发堆内存导出。

3. 后续分析建议

导出堆profile后,你可以用以下命令查看内存占用Top项:

go tool pprof heap_dump_xxxx.pprof
# 进入pprof交互环境后输入top,即可看到内存占用最高的函数列表
top
# 若想查看具体函数的内存分配细节,可输入list 函数名
list yourHighMemoryFunction

内容的提问来源于stack exchange,提问作者YuFeng Shen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:31:55