You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Jupyter Notebook训练模型时内核频繁崩溃问题求助

解决方案
  • 检查Jupyter的GPU内存分配限制
    Jupyter和PyCharm的GPU显存分配策略可能不一样,直接在训练代码开头加显存控制代码:
    如果你用TensorFlow:

    import tensorflow as tf
    gpus = tf.config.experimental.list_physical_devices('GPU')
    if gpus:
        try:
            for gpu in gpus:
                tf.config.experimental.set_memory_growth(gpu, True)
        except RuntimeError as e:
            print(e)
    

    用PyTorch的话:

    import torch
    torch.cuda.set_per_process_memory_fraction(0.8, 0) # 0.8可按需调整,限制单进程显存占比
    
  • 换方式启动Jupyter
    别用Anaconda自带的Jupyter快捷方式,直接打开你的虚拟环境终端,执行jupyter notebook启动,确保Jupyter完全用虚拟环境里的依赖,避免和全局环境的CUDA库冲突。

  • 调大内核超时时间
    先执行jupyter notebook --generate-config生成配置文件,找到jupyter_notebook_config.py,修改c.MappingKernelManager.timeout的值,把默认30秒改成300秒以上,防止内核因为训练耗时久被误判为无响应终止。

  • 核对Jupyter和PyCharm的CUDA环境变量
    在Jupyter里跑这段代码,对比PyCharm终端的输出,确认环境变量一致:

    import os
    print(os.environ.get('CUDA_PATH'))
    print(os.environ.get('CUDA_VISIBLE_DEVICES'))
    

    如果不一样,启动Jupyter前在虚拟环境终端手动设置:

    set CUDA_PATH=C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v11.7
    set CUDA_VISIBLE_DEVICES=0
    jupyter notebook
    
  • 关闭Jupyter自动保存和冗余输出
    在Jupyter界面的「Settings」里关掉「Autosave Documents」,代码里少用不必要的print,或者执行%config Application.log_level='ERROR'关闭冗余日志,减少资源占用。

内容的提问来源于stack exchange,提问作者gymoon10

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 11:10:42