使用Scikeras KerasRegressor遇'model'参数无效错误,求排查
解决Scikeras KerasRegressor配合GridSearchCV时的"Invalid parameter 'model'"错误
问题场景
使用scikeras.wrappers.KerasRegressor结合GridSearchCV进行超参数调优时,触发如下错误:
ValueError: Invalid parameter 'model' for estimator KerasRegressor( build_fn=<function auto_CreateNeural at 0x71a81c62f600>
相同逻辑在TensorFlow原生的KerasRegressor包装器中可正常运行,相关代码如下:
from scikeras.wrappers import KerasRegressor import numpy as np from sklearn import datasets from sklearn.model_selection import train_test_split, GridSearchCV, KFold from keras.models import Sequential from keras.layers import Dense from keras.optimizers import Adam diabetes = datasets.load_diabetes() X = diabetes.data y = diabetes.target # Split data into training, validation, and test sets X_train_val, X_test, y_train_val, y_test = train_test_split(X, y, test_size=0.2, random_state=1) X_train, X_val, y_train, y_val = train_test_split(X_train_val, y_train_val, test_size=0.125, random_state=1) # 0.125 x 0.8 = 0.1 # Define a model builder function def auto_CreateNeural(optimizer='adam', activation='relu'): regressor = Sequential() regressor.add(Dense(10, input_dim=X.shape[1], activation=activation)) regressor.add(Dense(1)) regressor.compile(loss='mean_squared_error', optimizer=optimizer) return regressor # Wrap the model with KerasRegressor regressor = KerasRegressor(build_fn=auto_CreateNeural, verbose=1) # Define parameters for GridSearchCV param_grid = { 'model__optimizer': ['adam', 'sgd'], 'model__activation': ['relu', 'tanh'], 'model__batch_size': [4, 8], 'model__epochs': [10, 20] } # Setup cross-validation kf = KFold(n_splits=3, shuffle=True, random_state=1) grid = GridSearchCV(estimator=regressor, param_grid=param_grid, cv=kf, scoring='neg_mean_squared_error', return_train_score=True) # Perform Grid Search grid_result = grid.fit(X_train_val, y_train_val) # Evaluate the best model on the test set best_model = grid.best_estimator_ test_loss = best_model.score(X_test, y_test) # Output results print("Best GridSearchCV score: {:.2f}".format(grid_result.best_score_)) print("Best parameters: {}".format(grid_result.best_params_)) print("Test set loss: {:.2f}".format(test_loss)) # Optionally, check how it performs on the validation set if needed validation_loss = best_model.score(X_val, y_val) print("Validation set loss: {:.2f}".format(validation_loss))
错误原因
Scikeras与TensorFlow原生KerasRegressor的参数传递规则存在差异:
- 原生包装器需要用
model__前缀传递模型构建函数的参数,但Scikeras不需要该前缀,直接使用参数名即可 batch_size和epochs是KerasRegressor自身的拟合参数,不属于模型构建函数的输入,不需要添加任何前缀
修正方案
- 移除参数网格中所有
model__前缀,直接使用参数名 - 确保模型构建函数只接收自身需要的参数(
optimizer、activation),batch_size和epochs作为KerasRegressor的参数单独在网格中定义
修正后的完整代码
from scikeras.wrappers import KerasRegressor import numpy as np from sklearn import datasets from sklearn.model_selection import train_test_split, GridSearchCV, KFold from keras.models import Sequential from keras.layers import Dense from keras.optimizers import Adam diabetes = datasets.load_diabetes() X = diabetes.data y = diabetes.target # Split data into training, validation, and test sets X_train_val, X_test, y_train_val, y_test = train_test_split(X, y, test_size=0.2, random_state=1) X_train, X_val, y_train, y_val = train_test_split(X_train_val, y_train_val, test_size=0.125, random_state=1) # 0.125 x 0.8 = 0.1 # Define a model builder function def auto_CreateNeural(optimizer='adam', activation='relu'): regressor = Sequential() regressor.add(Dense(10, input_dim=X.shape[1], activation=activation)) regressor.add(Dense(1)) regressor.compile(loss='mean_squared_error', optimizer=optimizer) return regressor # Wrap the model with KerasRegressor regressor = KerasRegressor(build_fn=auto_CreateNeural, verbose=1) # Define parameters for GridSearchCV - 移除model__前缀 param_grid = { 'optimizer': ['adam', 'sgd'], 'activation': ['relu', 'tanh'], 'batch_size': [4, 8], 'epochs': [10, 20] } # Setup cross-validation kf = KFold(n_splits=3, shuffle=True, random_state=1) grid = GridSearchCV(estimator=regressor, param_grid=param_grid, cv=kf, scoring='neg_mean_squared_error', return_train_score=True) # Perform Grid Search grid_result = grid.fit(X_train_val, y_train_val) # Evaluate the best model on the test set best_model = grid.best_estimator_ test_loss = best_model.score(X_test, y_test) # Output results print("Best GridSearchCV score: {:.2f}".format(grid_result.best_score_)) print("Best parameters: {}".format(grid_result.best_params_)) print("Test set loss: {:.2f}".format(test_loss)) # Optionally, check how it performs on the validation set if needed validation_loss = best_model.score(X_val, y_val) print("Validation set loss: {:.2f}".format(validation_loss))
内容的提问来源于stack exchange,提问作者Deep
相关产品推荐
相关产品推荐

