Keras OCR示例中Theano函数多轮训练后内存分配失败求助
解决Keras OCR训练大轮次后Theano内存分配问题
针对你遇到的训练大轮次后调用K.function触发内存分配错误的问题,我整理了几个除了减小数据集或训练轮次之外的可行方案,按优先级排序:
1. 分离训练模型与推理模型,避免加载冗余结构
你当前保存的是带CTC损失的训练模型,它包含额外的输入层(labels、input_length等),这些在推理时完全没用,会占用额外内存。建议训练时同时定义推理模型(仅输出softmax),训练完成后保存推理模型,而非训练模型:
# 第一步:先定义推理模型(仅用于输出softmax) inputs = Input(name='the_input', shape=x_train.shape[1:], dtype='float32') rnn_encoded = Bidirectional(LSTM(64, return_sequences=True,kernel_initializer=init,bias_initializer=bias),name='bidirectional_1',merge_mode='concat',trainable=trainable)(inputs) birnn_encoded = Bidirectional(LSTM(32, return_sequences=True,kernel_initializer=init,bias_initializer=bias),name='bidirectional_2',merge_mode='concat',trainable=trainable)(rnn_encoded) trirnn_encoded=Bidirectional(LSTM(16,return_sequences=True,kernel_initializer=init,bias_initializer=bias),name='bidirectional_3',merge_mode='concat',trainable=trainable)(birnn_encoded) output = TimeDistributed(Dense(28, name='dense',kernel_initializer=init,bias_initializer=bias))(trirnn_encoded) y_pred = Activation('softmax', name='softmax')(output) inference_model = Model(inputs=inputs, outputs=y_pred) # 第二步:定义训练模型(带CTC损失) labels = Input(name='the_labels', shape=[max_len], dtype='int32') input_length = Input(name='input_length', shape=[1], dtype='int64') label_length = Input(name='label_length', shape=[1], dtype='int64') loss_out = Lambda(ctc_lambda_func, output_shape=(1,), name='ctc')([y_pred, labels, input_length, label_length]) training_model = Model(inputs=[inputs, labels, input_length, label_length], outputs=loss_out) # 编译并训练 opt=RMSprop(lr=0.001,clipnorm=1.) training_model.compile(loss={'ctc': lambda y_true, y_pred: y_pred}, optimizer=opt) hist= training_model.fit_generator(my_generator,epochs=80,steps_per_epoch=100,shuffle=True,use_multiprocessing=False,workers=1) # 保存推理模型 inference_model.save(mfile) # 推理阶段:加载推理模型并创建函数 del training_model # 先删除训练模型释放内存 gc.collect() K.clear_session() gc.collect() from keras.models import load_model inference_model = load_model(mfile) test_func = K.function([inference_model.input], [inference_model.output])
2. 彻底清理训练会话,重启模型上下文
训练大轮次后,Keras/Theano的会话中会积累大量临时变量、梯度缓存等冗余数据,即使删除模型也可能残留。通过K.clear_session()彻底清空会话,再重新加载模型,能有效释放内存:
# 训练完成后 model.save(mfile) # 彻底清理 del model gc.collect() K.clear_session() # 清空Keras会话,释放所有关联内存 gc.collect() # 重新加载模型(从训练模型中提取推理用的输入输出) from keras.models import load_model training_model = load_model(mfile) inference_input = training_model.get_layer('the_input').input inference_output = training_model.get_layer('softmax').output test_func = K.function([inference_input], [inference_output])
3. 调整Theano内存配置,优化内存管理
Theano默认的内存设置可能不够灵活,你可以通过环境变量或代码调整内存回收和分配策略:
方法1:设置环境变量(启动Python前)
export THEANO_FLAGS="mode=FAST_RUN,device=cpu,floatX=float32,allow_gc=True,gc_heap_threshold=0.7"
allow_gc=True:启用自动垃圾回收gc_heap_threshold=0.7:当内存占用超过70%时触发垃圾回收
方法2:在代码中动态设置
import theano theano.config.allow_gc = True theano.config.gc_heap_threshold = 0.7 theano.config.cpu.model = 'fast' # 优化CPU运算效率,间接减少内存占用
4. 推理时分批次处理输入,降低单次内存占用
即使你训练用的批次是12,推理时可以用更小的批次(比如8甚至4),避免一次性为大量数据分配内存:
# 假设test_func已经创建 batch_size_infer = 8 all_predictions = [] for start_idx in range(0, len(x_train), batch_size_infer): end_idx = min(start_idx + batch_size_infer, len(x_train)) batch_x = x_train[start_idx:end_idx] batch_pred = test_func([batch_x])[0] all_predictions.append(batch_pred) # 合并所有批次的结果 y_pred_total = np.concatenate(all_predictions, axis=0)
5. 重启Python进程(终极方案)
如果以上方法都无效,最彻底的方式是训练完成后保存模型,然后重启Python进程,再加载模型进行推理。这样可以完全清除训练阶段积累的所有内存占用,适合内存泄漏比较顽固的情况。
内容的提问来源于stack exchange,提问作者Aayushee
相关产品推荐
相关产品推荐

