Keras RNN训练报错:CudnnRNN OpKernel未注册问题求助
解决DirectML TensorFlow环境下LSTM+GridSearchCV的CudnnRNN报错问题
问题背景
在基于DirectML和TensorFlow 2的tfamdgpu Anaconda环境中,使用Keras构建多层LSTM回归模型,执行GridSearchCV.fit()时触发报错:所有5次拟合失败,底层错误为Graph execution error,提示No OpKernel was registered to support Op 'CudnnRNN',更换AdamW等优化器无效。
核心原因
CudnnRNN是NVIDIA CUDA生态专属算子,DirectML作为AMD GPU的加速框架,不支持该算子;GridSearchCV设置n_jobs=-1会启用多进程并行训练,TensorFlow在多进程模式下会默认尝试调用CuDNN版本的LSTM,与DirectML环境冲突;- 默认的Keras LSTM在TensorFlow中会优先尝试使用CuDNN优化实现,而DirectML环境无法兼容。
解决方案步骤
- 强制LSTM使用非CuDNN兼容实现:在每个LSTM层添加参数
recurrent_activation='sigmoid'(CuDNN LSTM默认用hard_sigmoid,切换为sigmoid会强制使用通用CPU/GPU兼容的LSTM实现),或设置implementation=2; - 关闭GridSearchCV多进程:将
n_jobs从-1改为1,避免多进程导致的TensorFlow Graph模式冲突; - 统一数据类型:确保输入数据
X_train为float32类型,DirectML对数据类型兼容性要求严格。
修改后的代码示例
import numpy as np from tensorflow.keras.models import Sequential from tensorflow.keras.layers import LSTM, Dropout, Dense from tensorflow.keras.wrappers.scikit_learn import KerasRegressor from sklearn.model_selection import GridSearchCV # 确保输入数据为float32类型 X_train = X_train.astype(np.float32) y_train = y_train.astype(np.float32) def build_regressor(optimizer): regressor = Sequential() # 添加recurrent_activation='sigmoid'强制非CuDNN实现 regressor.add(LSTM(units=100, return_sequences=True, input_shape=(X_train.shape[1], 5), recurrent_activation='sigmoid')) regressor.add(Dropout(0.2)) regressor.add(LSTM(units=100, return_sequences=True, recurrent_activation='sigmoid')) regressor.add(Dropout(0.2)) regressor.add(LSTM(units=100, return_sequences=True, recurrent_activation='sigmoid')) regressor.add(Dropout(0.2)) regressor.add(LSTM(units=100, return_sequences=True, recurrent_activation='sigmoid')) regressor.add(Dropout(0.2)) regressor.add(LSTM(units=100, return_sequences=True, recurrent_activation='sigmoid')) regressor.add(Dropout(0.2)) regressor.add(LSTM(units=100, recurrent_activation='sigmoid')) regressor.add(Dropout(0.2)) regressor.add(Dense(units=1)) regressor.compile(optimizer=optimizer, loss="mean_squared_error") return regressor regressor = KerasRegressor(build_fn=build_regressor) parameters = {"batch_size": [8, 16, 32, 64], "epochs": [100, 200, 300], "optimizer": ["adam", "adadelta", "rmsprop"]} # 将n_jobs改为1,关闭多进程 grid_search = GridSearchCV(estimator=regressor, param_grid=parameters, scoring="neg_mean_squared_error", n_jobs=1) grid_search = grid_search.fit(X_train, y_train)
内容的提问来源于stack exchange,提问作者roty2Poulet
相关产品推荐
相关产品推荐

