You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何记录云端Jupyter Notebook任务的峰值内存使用情况

记录Jupyter Notebook任务峰值内存的方法

方法1:使用memory_profiler库

这是专门用于Python代码内存分析的工具,能直接在Notebook中记录峰值内存。

  • 先安装库:
    pip install memory-profiler
    
  • 在Notebook中加载扩展:
    %load_ext memory_profiler
    
  • 若要分析单个代码块,用%memit魔法命令包裹运行:
    %memit
    # 这里放你的长时间任务代码,示例:
    import numpy as np
    large_array = np.random.rand(10**8)
    
    运行后会直接输出该代码块的峰值内存使用量。
  • 若要分析特定函数,给函数加@profile装饰器,再运行函数:
    @profile
    def long_running_task():
        large_array = np.random.rand(10**8)
        # 其他任务逻辑
        return large_array
    
    long_running_task()
    
    执行后会输出函数每一行的内存使用数据,最后包含峰值统计。

方法2:用psutil手动监控内存

psutil是跨平台的系统监控库,可编写轻量后台监控逻辑,自定义采样频率,性能开销更低。

  • 安装库:
    pip install psutil
    
  • 编写并行监控代码:
    import psutil
    import threading
    import time
    
    peak_memory = 0
    stop_monitor = False
    
    def monitor_memory():
        global peak_memory
        process = psutil.Process()
        while not stop_monitor:
            current_mem = process.memory_info().rss / (1024**3)  # 转换为GB
            if current_mem > peak_memory:
                peak_memory = current_mem
            time.sleep(1)  # 每秒采样一次,可调整间隔
    
    # 启动监控线程
    monitor_thread = threading.Thread(target=monitor_memory)
    monitor_thread.start()
    
    # 执行你的长时间任务
    import numpy as np
    large_array = np.random.rand(10**8)
    # ... 其他任务代码 ...
    
    # 任务结束后停止监控并输出结果
    stop_monitor = True
    monitor_thread.join()
    print(f"峰值内存使用量: {peak_memory:.2f} GB")
    

方法3:系统级命令监控

如果任务可导出为独立Python脚本,可使用系统自带的time命令(Linux/macOS)获取峰值内存,几乎无性能开销。

  • 将Notebook中的任务代码保存为task.py。
  • 在终端运行:
    /usr/bin/time -v python task.py
    
    运行结束后,输出内容里的Maximum resident set size即为峰值内存(单位KB,可自行转换为GB/MB)。
    若要在Notebook中执行,可通过subprocess调用:
    import subprocess
    
    result = subprocess.run(['/usr/bin/time', '-v', 'python', 'task.py'], stderr=subprocess.PIPE, text=True)
    # 从stderr中提取峰值内存数据
    for line in result.stderr.split('\n'):
        if 'Maximum resident set size' in line:
            peak_kb = int(line.split(':')[1].strip())
            peak_gb = peak_kb / (1024**2)
            print(f"峰值内存使用量: {peak_gb:.2f} GB")
    

内容的提问来源于stack exchange,提问作者lpounng

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 11:30:00