Jupyter Notebook训练模型时内核频繁崩溃问题求助
解决方案
检查Jupyter的GPU内存分配限制
Jupyter和PyCharm的GPU显存分配策略可能不一样,直接在训练代码开头加显存控制代码:
如果你用TensorFlow:import tensorflow as tf gpus = tf.config.experimental.list_physical_devices('GPU') if gpus: try: for gpu in gpus: tf.config.experimental.set_memory_growth(gpu, True) except RuntimeError as e: print(e)用PyTorch的话:
import torch torch.cuda.set_per_process_memory_fraction(0.8, 0) # 0.8可按需调整,限制单进程显存占比换方式启动Jupyter
别用Anaconda自带的Jupyter快捷方式,直接打开你的虚拟环境终端,执行jupyter notebook启动,确保Jupyter完全用虚拟环境里的依赖,避免和全局环境的CUDA库冲突。调大内核超时时间
先执行jupyter notebook --generate-config生成配置文件,找到jupyter_notebook_config.py,修改c.MappingKernelManager.timeout的值,把默认30秒改成300秒以上,防止内核因为训练耗时久被误判为无响应终止。核对Jupyter和PyCharm的CUDA环境变量
在Jupyter里跑这段代码,对比PyCharm终端的输出,确认环境变量一致:import os print(os.environ.get('CUDA_PATH')) print(os.environ.get('CUDA_VISIBLE_DEVICES'))如果不一样,启动Jupyter前在虚拟环境终端手动设置:
set CUDA_PATH=C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v11.7 set CUDA_VISIBLE_DEVICES=0 jupyter notebook关闭Jupyter自动保存和冗余输出
在Jupyter界面的「Settings」里关掉「Autosave Documents」,代码里少用不必要的print,或者执行%config Application.log_level='ERROR'关闭冗余日志,减少资源占用。
内容的提问来源于stack exchange,提问作者gymoon10
相关产品推荐
相关产品推荐

