Python IDLE运行Unet训练代码时Shell意外重启问题求助
model.fit阶段(Epoch 1/10) 问题描述
我正在使用Python IDLE运行一段用于肺部结节检测的Unet训练代码,但每次执行到模型拟合(model.fit)的Epoch 1/10阶段时,Shell就会意外重启。此前代码运行正常,目前无法继续执行。
相关代码
from __future__ import print_function import numpy as np from keras.models import Model from keras.layers import Input, Conv2D, MaxPooling2D, UpSampling2D from keras.layers import concatenate from keras.optimizers import Adam from keras.optimizers import SGD from keras.callbacks import ModelCheckpoint, LearningRateScheduler from keras import backend as K K.set_image_dim_ordering('th') # Theano dimension ordering in this code img_rows = 512 img_cols = 512 smooth = 1. def dice_coef(y_true, y_pred): y_true_f = K.flatten(y_true) y_pred_f = K.flatten(y_pred) intersection = K.sum(y_true_f * y_pred_f) return (2. * intersection + smooth) / (K.sum(y_true_f) + K.sum(y_pred_f) + smooth) def dice_coef_np(y_true,y_pred): y_true_f = y_true.flatten() y_pred_f = y_pred.flatten() intersection = np.sum(y_true_f * y_pred_f) return (2. * intersection + smooth) / (np.sum(y_true_f) + np.sum(y_pred_f) + smooth) def dice_coef_loss(y_true, y_pred): return -dice_coef(y_true, y_pred) def get_unet(): inputs = Input((1,img_rows, img_cols)) conv1 = Conv2D(32, (3, 3), activation='relu', padding='same')(inputs) conv1 = Conv2D(32, (3, 3), activation='relu', padding='same')(conv1) pool1 = MaxPooling2D(pool_size=(2, 2))(conv1) conv2 = Conv2D(64, (3, 3), activation='relu', padding='same')(pool1) conv2 = Conv2D(64, (3, 3), activation='relu', padding='same')(conv2) pool2 = MaxPooling2D(pool_size=(2, 2))(conv2) conv3 = Conv2D(128, (3, 3), activation='relu', padding='same')(pool2) conv3 = Conv2D(128, (3, 3), activation='relu', padding='same')(conv3) pool3 = MaxPooling2D(pool_size=(2, 2))(conv3) conv4 = Conv2D(256, (3, 3), activation='relu', padding='same')(pool3) conv4 = Conv2D(256, (3, 3), activation='relu', padding='same')(conv4) pool4 = MaxPooling2D(pool_size=(2, 2))(conv4) conv5 = Conv2D(512, (3, 3), activation='relu', padding='same')(pool4) conv5 = Conv2D(512, (3, 3), activation='relu', padding='same')(conv5) up6 = concatenate([UpSampling2D(size=(2, 2))(conv5), conv4], axis=1) conv6 = Conv2D(256, (3, 3), activation='relu', padding='same')(up6) conv6 = Conv2D(256, (3, 3), activation='relu', padding='same')(conv6) up7 = concatenate([UpSampling2D(size=(2, 2))(conv6), conv3], axis=1) conv7 = Conv2D(128, (3, 3), activation='relu', padding='same')(up7) conv7 = Conv2D(128, (3, 3), activation='relu', padding='same')(conv7) up8 = concatenate([UpSampling2D(size=(2, 2))(conv7), conv2], axis=1) conv8 = Conv2D(64, (3, 3), activation='relu', padding='same')(up8) conv8 = Conv2D(64, (3, 3), activation='relu', padding='same')(conv8) up9 = concatenate([UpSampling2D(size=(2, 2))(conv8), conv1], axis=1) conv9 = Conv2D(32, (3, 3), activation='relu', padding='same')(up9) conv9 = Conv2D(32, (3, 3), activation='relu', padding='same')(conv9) conv10 = Conv2D(1, (1, 1), activation='sigmoid')(conv9) model = Model(inputs=inputs, outputs=conv10) model.compile(optimizer=Adam(lr=1.0e-5), loss=dice_coef_loss, metrics=[dice_coef]) return model def train_and_predict(use_existing): print('-'*30) print('Loading and preprocessing train data...') print('-'*30) imgs_train = np.load("C:/Users/hirplk/Desktop/unet/Luna2016-Lung-Nodule-Detection-master_new/DATA_PROCESS/scratch/cse/dual/cs5130287/Luna2016/output_final/"+"trainImages.npy").astype(np.float32) imgs_mask_train = np.load("C:/Users/hirplk/Desktop/unet/Luna2016-Lung-Nodule-Detection-master_new/DATA_PROCESS/scratch/cse/dual/cs5130287/Luna2016/output_final/"+"trainMasks.npy").astype(np.float32) imgs_test = np.load("C:/Users/hirplk/Desktop/unet/Luna2016-Lung-Nodule-Detection-master_new/DATA_PROCESS/scratch/cse/dual/cs5130287/Luna2016/output_final/"+"testImages.npy").astype(np.float32) imgs_mask_test_true = np.load("C:/Users/hirplk/Desktop/unet/Luna2016-Lung-Nodule-Detection-master_new/DATA_PROCESS/scratch/cse/dual/cs5130287/Luna2016/output_final/"+"testMasks.npy").astype(np.float32) mean = np.mean(imgs_train) # mean for data centering std = np.std(imgs_train) # std for data normalization imgs_train -= mean # images should already be standardized, but just in case imgs_train /= std print('-'*30) print('Creating and compiling model...') print('-'*30) model = get_unet() # Saving weights to unet.hdf5 at checkpoints model_checkpoint = ModelCheckpoint('unet.hdf5', monitor='loss', save_best_only=True) if use_existing: model.load_weights('./unet.hdf5') print('-'*30) print('Fitting model...') print('-'*30) model.fit(imgs_train, imgs_mask_train, batch_size=2, epochs=10, verbose=1, shuffle=True, callbacks=[model_checkpoint]) print ('bbbbbbbbbbbbbbbbbbbbbbbbbbbbbb')
IDLE输出信息
RESTART: C:\Users\hirplk\Desktop\unet\DSB3Tutorial-master\tutorial_code\LUNA_train_unet.py Warning (from warnings module): File "C:\Research\Python_installation\lib\site-packages\h5py\__init__.py", line 36 from ._conv import register_converters as _register_converters FutureWarning: Conversion of the second argument of issubdtype from `float` to `np.floating` is deprecated. In future, it will be treated as `np.float64 == np.dtype(float).type`. Using TensorFlow backend. ------------------------------ Loading and preprocessing train data... ------------------------------ ------------------------------ Creating and compiling model... ------------------------------ ------------------------------ Fitting model... ------------------------------ Epoch 1/10 =============================== RESTART: Shell ===============================
请问这是否意味着Python崩溃了?有没有人遇到过类似问题?我是否需要重新安装所有环境?
解答
首先明确:这确实是Python进程意外崩溃的表现,IDLE Shell重启就是系统终止了Python进程的直接反馈。不用急着重装整个环境,先按以下步骤排查:
优先排查内存不足问题:
你的输入是512×512的图像,加上Unet模型的参数量(尤其是conv5层用了512个通道),batch size=2很容易把GPU显存或系统内存占满,触发系统的OOM(内存不足)杀手,直接终止进程。- 先把
batch_size改成1,或者临时把图像尺寸缩小到256×256测试,看是否能正常完成Epoch 1 - 训练前打开任务管理器(Windows)或
top命令(Linux/macOS),观察内存/显存占用,若训练开始后瞬间拉满,那就是内存问题无疑
- 先把
换个运行环境试试:
IDLE对大型计算任务的支持很差,崩溃后不会留下任何详细日志。建议直接用命令行(cmd/终端)运行脚本,或者用VS Code的Python终端、Jupyter Notebook,这样能看到崩溃时的具体报错(比如Segmentation Fault或OOM提示),方便定位问题。降低模型复杂度:
可以先临时修改get_unet()函数里的conv5层通道数,从512改成256,减少模型参数量,测试是否还会崩溃。如果能正常运行,再逐步调回参数,同时优化内存使用(比如用梯度累积代替大batch size)。检查依赖兼容性:
你输出里的h5py警告虽然不是直接崩溃原因,但依赖版本不兼容也可能导致隐性崩溃:- 尝试更新h5py到稳定版:
pip install --upgrade h5py - 确认Keras和TensorFlow版本匹配(比如TensorFlow 1.x对应Keras 2.x的特定版本,若你用的是旧版框架,版本不匹配很容易出问题)
- 尝试更新h5py到稳定版:
验证数据完整性:
加载的npy文件可能损坏或维度异常,在model.fit前添加几行代码检查数据:print(imgs_train.shape, imgs_mask_train.shape) print(np.min(imgs_train), np.max(imgs_train))确认数据维度符合模型输入要求(
(样本数, 1, 512, 512)),且没有异常值。
如果以上步骤都试过还是崩溃,再考虑重新安装环境——但大概率是内存或环境适配问题,针对性调整就能解决。
内容的提问来源于stack exchange,提问作者user2445123

