Jupyter训练TensorFlow DeepLab模型时内核崩溃问题求助
问题
在Jupyter(VSCode Jupyter、Jupyter Notebook/Lab均出现)中训练基于ResNet50编码器、输入尺寸256的DeepLab模型时,内核在第一个epoch开始前崩溃。其他模型无此问题。
训练时单元输出如下:
Epoch 1/100 2023-01-07 12:22:01.752760: W tensorflow/tsl/platform/profile_utils/cpu_utils.cc:128] Failed to get CPU frequency: 0 Hz 2023-01-07 12:22:05.727903: I tensorflow/core/grappler/optimizers/custom_graph_optimizer_registry.cc:114] Plugin optimizer for device_type GPU is enabled. The Kernel crashed while executing code in the the current cell or a previous cell. Please review the code in the cell(s) to identify a possible cause of the failure. Click here for more info. View Jupyter log for further details. Canceled future for execute_request message before replies were done
设备环境:M1 MacBook Pro,TensorFlow 2.11.0(macos版本),Python 3.10。已尝试的解决方法:重启内核、重新安装TensorFlow、创建新环境、使用nomkl库。
解决建议
针对M1芯片Mac上TensorFlow训练大模型时内核崩溃的问题,可尝试以下方案:
降低模型内存占用:
- 先将输入尺寸减小至128×128验证是否能正常运行,确认内存过载是崩溃原因后,再逐步调回256;
- 启用混合精度训练,添加代码:
from tensorflow.keras.mixed_precision import set_global_policy set_global_policy('mixed_float16') - 使用轻量化的ResNet50V2替代原版ResNet50,或加载预训练权重时设置
include_top=False,自定义后续层以减少冗余参数。
调整TensorFlow环境配置:
- 降级TensorFlow到2.9.x或2.10.x版本,部分M1用户反馈2.11版本存在兼容性问题;
- 确保安装Apple Silicon优化版TensorFlow:
pip uninstall tensorflow pip install tensorflow-macos tensorflow-metal - 开启GPU内存动态增长,避免一次性占用过多显存:
import tensorflow as tf gpus = tf.config.experimental.list_physical_devices('GPU') if gpus: try: for gpu in gpus: tf.config.experimental.set_memory_growth(gpu, True) except RuntimeError as e: print(e)
Jupyter环境优化:
- 关闭Jupyter中其他闲置内核,释放系统内存;
- 直接在终端运行训练脚本,排除Jupyter资源调度问题,验证模型本身是否能正常运行。
内容的提问来源于stack exchange,提问作者neuops
相关产品推荐
相关产品推荐

