You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keras OCR示例中Theano函数多轮训练后内存分配失败求助

解决Keras OCR训练大轮次后Theano内存分配问题

针对你遇到的训练大轮次后调用K.function触发内存分配错误的问题,我整理了几个除了减小数据集或训练轮次之外的可行方案,按优先级排序:

1. 分离训练模型与推理模型,避免加载冗余结构

你当前保存的是带CTC损失的训练模型,它包含额外的输入层(labels、input_length等),这些在推理时完全没用,会占用额外内存。建议训练时同时定义推理模型(仅输出softmax),训练完成后保存推理模型,而非训练模型:

# 第一步:先定义推理模型(仅用于输出softmax)
inputs = Input(name='the_input', shape=x_train.shape[1:], dtype='float32')
rnn_encoded = Bidirectional(LSTM(64, return_sequences=True,kernel_initializer=init,bias_initializer=bias),name='bidirectional_1',merge_mode='concat',trainable=trainable)(inputs)
birnn_encoded = Bidirectional(LSTM(32, return_sequences=True,kernel_initializer=init,bias_initializer=bias),name='bidirectional_2',merge_mode='concat',trainable=trainable)(rnn_encoded)
trirnn_encoded=Bidirectional(LSTM(16,return_sequences=True,kernel_initializer=init,bias_initializer=bias),name='bidirectional_3',merge_mode='concat',trainable=trainable)(birnn_encoded)
output = TimeDistributed(Dense(28, name='dense',kernel_initializer=init,bias_initializer=bias))(trirnn_encoded)
y_pred = Activation('softmax', name='softmax')(output)
inference_model = Model(inputs=inputs, outputs=y_pred)

# 第二步:定义训练模型(带CTC损失)
labels = Input(name='the_labels', shape=[max_len], dtype='int32')
input_length = Input(name='input_length', shape=[1], dtype='int64')
label_length = Input(name='label_length', shape=[1], dtype='int64')
loss_out = Lambda(ctc_lambda_func, output_shape=(1,), name='ctc')([y_pred, labels, input_length, label_length])
training_model = Model(inputs=[inputs, labels, input_length, label_length], outputs=loss_out)

# 编译并训练
opt=RMSprop(lr=0.001,clipnorm=1.)
training_model.compile(loss={'ctc': lambda y_true, y_pred: y_pred}, optimizer=opt)
hist= training_model.fit_generator(my_generator,epochs=80,steps_per_epoch=100,shuffle=True,use_multiprocessing=False,workers=1)

# 保存推理模型
inference_model.save(mfile)

# 推理阶段:加载推理模型并创建函数
del training_model  # 先删除训练模型释放内存
gc.collect()
K.clear_session()
gc.collect()

from keras.models import load_model
inference_model = load_model(mfile)
test_func = K.function([inference_model.input], [inference_model.output])

2. 彻底清理训练会话,重启模型上下文

训练大轮次后,Keras/Theano的会话中会积累大量临时变量、梯度缓存等冗余数据,即使删除模型也可能残留。通过K.clear_session()彻底清空会话,再重新加载模型,能有效释放内存:

# 训练完成后
model.save(mfile)
# 彻底清理
del model
gc.collect()
K.clear_session()  # 清空Keras会话,释放所有关联内存
gc.collect()

# 重新加载模型(从训练模型中提取推理用的输入输出)
from keras.models import load_model
training_model = load_model(mfile)
inference_input = training_model.get_layer('the_input').input
inference_output = training_model.get_layer('softmax').output
test_func = K.function([inference_input], [inference_output])

3. 调整Theano内存配置,优化内存管理

Theano默认的内存设置可能不够灵活,你可以通过环境变量或代码调整内存回收和分配策略:

方法1:设置环境变量(启动Python前)

export THEANO_FLAGS="mode=FAST_RUN,device=cpu,floatX=float32,allow_gc=True,gc_heap_threshold=0.7"
  • allow_gc=True:启用自动垃圾回收
  • gc_heap_threshold=0.7:当内存占用超过70%时触发垃圾回收

方法2:在代码中动态设置

import theano
theano.config.allow_gc = True
theano.config.gc_heap_threshold = 0.7
theano.config.cpu.model = 'fast'  # 优化CPU运算效率,间接减少内存占用

4. 推理时分批次处理输入,降低单次内存占用

即使你训练用的批次是12,推理时可以用更小的批次(比如8甚至4),避免一次性为大量数据分配内存:

# 假设test_func已经创建
batch_size_infer = 8
all_predictions = []

for start_idx in range(0, len(x_train), batch_size_infer):
    end_idx = min(start_idx + batch_size_infer, len(x_train))
    batch_x = x_train[start_idx:end_idx]
    batch_pred = test_func([batch_x])[0]
    all_predictions.append(batch_pred)

# 合并所有批次的结果
y_pred_total = np.concatenate(all_predictions, axis=0)

5. 重启Python进程(终极方案)

如果以上方法都无效,最彻底的方式是训练完成后保存模型,然后重启Python进程,再加载模型进行推理。这样可以完全清除训练阶段积累的所有内存占用,适合内存泄漏比较顽固的情况。

内容的提问来源于stack exchange,提问作者Aayushee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 10:18:11