Optuna优化TensorFlow超参数时Graph execution error问题排查
问题:WSL中Optuna+TensorFlow超参优化时出现CUDNN内存错误
问题背景
此前在WSL环境的Jupyter Notebook(Python 3.x)中,使用Optuna对TensorFlow的GRU模型做超参优化,数百次试验均正常。但添加保存study到.pkl文件的代码后,第三次trial触发CUDNN_STATUS_INTERNAL_ERROR;即使注释掉保存代码和gc_after_trial=True恢复原状态,问题依然存在。对于更密集的神经网络架构,首次trial就出现内存耗尽类错误。
错误日志
常规模型第三次trial错误
E tensorflow/stream_executor/dnn.cc:868] CUDNN_STATUS_INTERNAL_ERROR in tensorflow/stream_executor/cuda/cuda_dnn.cc(2683): 'cudnnRNNForwardTraining( cudnn.handle(), rnn_desc.handle(), model_dims.max_seq_length, input_desc.handles(), input_data.opaque(), input_h_desc.handle(), input_h_data.opaque(), input_c_desc.handle(), input_c_data.opaque(), rnn_desc.params_handle(), params.opaque(), output_desc.handles(), output_data->opaque(), output_h_desc.handle(), output_h_data->opaque(), output_c_desc.handle(), output_c_data->opaque(), workspace.opaque(), workspace.size(), reserve_space.opaque(), reserve_space.size())' W tensorflow/core/framework/op_kernel.cc:1745] OP_REQUIRES failed at cudnn_rnn_ops.cc:1563 : INTERNAL: Failed to call ThenRnnForward with model config: [rnn_mode, rnn_input_mode, rnn_direction_mode]: 2, 0, 0 , [num_layers, input_size, num_units, dir_count, max_seq_length, batch_size, cell_num_units]: [1, 91, 81, 1, 1, 128, 81] Trial 2 failed with parameters: {'units': 81, 'activation': 'softsign', 'dropout': 0.07633939325087957, 'optimizer': 'Adam', 'adam_learning_rate': 0.01799516104446331, 'filters': 91} because of the following error: InternalError(). Traceback (most recent call last): File "/usr/local/lib/python3.8/dist-packages/optuna/study/_optimize.py", line 200, in _run_trial value_or_values = func(trial) File "/tmp/ipykernel_3632372/3676345196.py", line 167, in objective_function self.neural_network.train_model(test_model) File "/tmp/ipykernel_3632372/227766830.py", line 178, in train_model history = model.fit(self.x_train, self.y_train, epochs = epoch_size, batch_size = BATCH_SIZE, callbacks = [early_stop], File "/usr/local/lib/python3.8/dist-packages/keras/utils/traceback_utils.py", line 67, in error_handler raise e.with_traceback(filtered_tb) from None File "/usr/local/lib/python3.8/dist-packages/tensorflow/python/eager/execute.py", line 54, in quick_execute tensors = pywrap_tfe.TFE_Py_Execute(ctx._handle, device_name, op_name, tensorflow.python.framework.errors_impl.InternalError: Graph execution error: Failed to call ThenRnnForward with model config: [rnn_mode, rnn_input_mode, rnn_direction_mode]: 2, 0, 0 , [num_layers, input_size, num_units, dir_count, max_seq_length, batch_size, cell_num_units]: [1, 91, 81, 1, 1, 128, 81] [[{{node CudnnRNN}}]] [[sequential/lstm/PartitionedCall]] [Op:__inference_train_function_121807]
密集模型第一次trial错误
E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:56] Histogram of current allocation: (allocation_size_in_bytes, nb_allocation_of_that_sizes), ...; E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:59] 4, 27 E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:59] 8, 8 E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:59] 272, 3 E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:59] 332, 3 E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:59] 512, 1 E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:59] 544, 6 E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:59] 1028, 1 E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:59] 7968, 4 E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:59] 12288, 1 E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:59] 18496, 6 E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:59] 42496, 1 E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:59] 45152, 6 E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:59] 93908, 1 E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:59] 751264, 1 E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:59] 16819712, 1 E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:90] CU_MEMPOOL_ATTR_RESERVED_MEM_CURRENT: 67108864 E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:92] CU_MEMPOOL_ATTR_USED_MEM_CURRENT: 18140216 E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:93] CU_MEMPOOL_ATTR_RESERVED_MEM_HIGH: 67108864 E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:94] CU_MEMPOOL_ATTR_USED_MEM_HIGH: 34937704 E tensorflow/stream_executor/dnn.cc:868] CUDNN_STATUS_INTERNAL_ERROR in tensorflow/stream_executor/cuda/cuda_dnn.cc(2683): 'cudnnRNNForwardTraining( cudnn.handle(), rnn_desc.handle(), model_dims.max_seq_length, input_desc.handles(), input_data.opaque(), input_h_desc.handle(), input_h_data.opaque(), input_c_desc.handle(), input_c_data.opaque(), rnn_desc.params_handle(), params.opaque(), output_desc.handles(), output_data->opaque(), output_h_desc.handle(), output_h_data->opaque(), output_c_desc.handle(), output_c_data->opaque(), workspace.opaque(), workspace.size(), reserve_space.opaque(), reserve_space.size())' W tensorflow/core/framework/op_kernel.cc:1745] OP_REQUIRES failed at cudnn_rnn_ops.cc:1563 : INTERNAL: Failed to call ThenRnnForward with model config: [rnn_mode, rnn_input_mode, rnn_direction_mode]: 2, 0, 0 , [num_layers, input_size, num_units, dir_count, max_seq_length, batch_size, cell_num_units]: [1, 83, 34, 1, 1, 128, 34] E tensorflow/stream_executor/dnn.cc:868] CUDNN_STATUS_INTERNAL_ERROR in tensorflow/stream_executor/cuda/cuda_dnn.cc(2683): 'cudnnRNNForwardTraining( cudnn.handle(), rnn_desc.handle(), model_dims.max_seq_length, input_desc.handles(), input_data.opaque(), input_h_desc.handle(), input_h_data.opaque(), input_c_desc.handle(), input_c_data.opaque(), rnn_desc.params_handle(), params.opaque(), output_desc.handles(), output_data->opaque(), output_h_desc.handle(), output_h_data->opaque(), output_c_desc.handle(), output_c_data->opaque(), workspace.opaque(), workspace.size(), reserve_space.opaque(), reserve_space.size())' W tensorflow/core/framework/op_kernel.cc:1745] OP_REQUIRES failed at cudnn_rnn_ops.cc:1563 : INTERNAL: Failed to call ThenRnnForward with model config: [rnn_mode, rnn_input_mode, rnn_direction_mode]: 2, 0, 0 , [num_layers, input_size, num_units, dir_count, max_seq_length, batch_size, cell_num_units]: [1, 83, 34, 1, 1, 128, 34] Trial 0 failed with parameters: {'units': 34, 'activation': 'softsign', 'dropout': 0.1391979014457847, 'optimizer': 'Adam', 'adam_learning_rate': 0.07514111264388643, 'filters': 83} because of the following error: InternalError().
相关代码
优化器函数
def optimize(self): best_params, best_values = self.optimize_study() print(f"Best params: {best_params}\n Best value: {best_values}") return self def objective_function(self, trial): units = trial.suggest_int('units', 10, 50) activation = trial.suggest_categorical("activation", ['relu', 'tanh', 'softsign']) dropout = trial.suggest_float('dropout', 0.01, 0.5) test_model = self.build_deep_GRU_model(trial) self.neural_network.train_model(test_model) y_true, y_pred = self.neural_network.predict(test_model) return self.neural_network.evaluate_loss_function(y_true, y_pred)
神经网络函数
def build_GRU_model(self, hidden_neurons, activator, drop_out, OPTIMIZER = 'adam'): keras.backend.clear_session() GRU_layer = keras.layers.GRU(hidden_neurons, dropout = drop_out, activation = activator) gru_model = keras.Sequential(layers = (GRU_layer, keras.layers.Dense(self.output_neurons))) gru_model.reset_states() gru_model.compile(optimizer = OPTIMIZER, loss = self.mae) return gru_model def train_model(self, model, epoch_size = 150, BATCH_SIZE = BATCH_SIZE): early_stop = tf.keras.callbacks.EarlyStopping(monitor = 'val_loss', patience = 30, mode = 'min') history = model.fit(self.x_train, self.y_train, epochs = epoch_size, batch_size = BATCH_SIZE, callbacks = [early_stop], validation_data = (self.x_valid, self.y_valid), shuffle = False) print(model.summary()) return history
原因分析
- GPU内存泄漏:虽然调用了
keras.backend.clear_session(),但Optuna的trial循环中,模型对象、训练历史等资源可能未被完全回收;WSL环境下TensorFlow的GPU内存回收机制存在延迟,导致显存碎片累积。 - CUDNN兼容性bug:特定模型参数组合(如GRU单元数、输入维度)触发了CUDNN与TensorFlow的底层兼容性问题,导致显存分配失败。
- WSL显存管理限制:WSL对GPU显存的管理不如原生Linux稳定,多次试验后显存碎片增多,即使总显存未耗尽,也无法分配连续内存块给CUDNN的RNN操作。
解决方法
- 强制资源回收:在每个trial结束后手动清理模型并触发垃圾回收,添加到
objective_function末尾:del test_model import gc gc.collect() keras.backend.clear_session() - 设置显存按需分配:限制TensorFlow仅按需占用显存,避免预占过多资源:
gpus = tf.config.list_physical_devices('GPU') if gpus: try: for gpu in gpus: tf.config.experimental.set_memory_growth(gpu, True) except RuntimeError as e: print(e) - 降低batch size:针对密集模型,减小
BATCH_SIZE,降低单次训练的显存占用。 - 重启环境:彻底重启Jupyter或WSL,释放所有显存碎片,重新开始试验。
- 检查版本兼容性:确认TensorFlow与CUDNN版本匹配,避免因版本不兼容导致的底层错误。
内容的提问来源于stack exchange,提问作者Kartik
相关产品推荐
相关产品推荐

