XGBRegressor使用sklearn的mean_squared_error时样本数不一致问题
XGBoost多输出回归使用sklearn MSE报错问题解决
问题背景
用户尝试训练多输出XGBRegressor,代码如下:
import xgboost as xgb from sklearn.metrics import mean_squared_error def xgboost(): model = xgb.XGBRegressor(n_estimators=200, max_depth=4, subsample=1, min_child_weight=1, objective='reg:squarederror', tree_method='hist', eval_metric=mean_squared_error, # 使用sklearn的MSE early_stopping_rounds=50) return model num_amostras = x_train.shape[0] val_size = 0.2 num_amostras_train = int(num_amostras * (1-val_size)) x_train_xgb = x_train[:num_amostras_train] y_train_xgb = y_train[:num_amostras_train] x_val_xgb = x_train[num_amostras_train:] y_val_xgb = y_train[num_amostras_train:] model_xgb = xgboost() model_xgb.fit(x_train_xgb, y_train_xgb, eval_set=[(x_train_xgb, y_train_xgb), (x_val_xgb, y_val_xgb)]) resultados = model_xgb.evals_result()
数据形状:
x_train: (1458, 55)x_train_xgb: (1166, 55),y_train_xgb: (1166, 24)x_val_xgb: (292, 55),y_val_xgb: (292, 24)
运行时报错:
Traceback (most recent call last): File ~\PeDFurnas\lib\site-packages\spyder_kernels\py3compat.py:356 in compat_exec exec(code, globals, locals) File c:\users\ldsp_\sipredvs\scripts\treinamento_demanda.py:201 model_xgb.fit(x_train_xgb, y_train_xgb, eval_set=[(x_train_xgb, y_train_xgb),(x_val_xgb, y_val_xgb)]) File ~\PeDFurnas\lib\site-packages\xgboost\core.py:729 in inner_f return func(**kwargs) File ~\PeDFurnas\lib\site-packages\xgboost\sklearn.py:1086 in fit self._Booster = train( File ~\PeDFurnas\lib\site-packages\xgboost\core.py:729 in inner_f return func(**kwargs) File ~\PeDFurnas\lib\site-packages\xgboost\training.py:182 in train if cb_container.after_iteration(bst, i, dtrain, evals): File ~\PeDFurnas\lib\site-packages\xgboost\callback.py:238 in after_iteration score: str = model.eval_set(evals, epoch, self.metric, self._output_margin) File ~\PeDFurnas\lib\site-packages\xgboost\core.py:2138 in eval_set feval_ret = feval( File ~\PeDFurnas\lib\site-packages\xgboost\sklearn.py:139 in inner return func.__name__, func(y_true, y_score) File ~\PeDFurnas\lib\site-packages\sklearn\metrics\_regression.py:442 in mean_squared_error y_type, y_true, y_pred, multioutput = _check_reg_targets( File ~\PeDFurnas\lib\site-packages\sklearn\metrics\_regression.py:100 in _check_reg_targets check_consistent_length(y_true, y_pred) File ~\PeDFurnas\lib\site-packages\sklearn\utils\validation.py:397 in check_consistent_length raise ValueError( ValueError: Found input variables with inconsistent numbers of samples: [27984, 1166]
其中27984=1166*24,是y_train_xgb形状的乘积;改用XGBoost默认的'rmse'指标可正常运行。
问题原因
- 形状不匹配:在多输出回归场景下,XGBoost返回的预测值
y_score是扁平化的一维数组(形状为(n_samples * n_outputs,)),而sklearn的mean_squared_error默认接收的y_true是二维数组((n_samples, n_outputs)),两者样本数检查时不匹配,导致报错。 - 内置指标适配性:XGBoost内置的
'rmse'/'mse'指标已经针对多输出场景做了适配,会自动处理形状对齐,因此可以正常运行。
修复方法
方法1:使用XGBoost内置的MSE指标
直接将eval_metric指定为'mse',效果和sklearn的mean_squared_error一致,代码修改如下:
def xgboost(): model = xgb.XGBRegressor(n_estimators=200, max_depth=4, subsample=1, min_child_weight=1, objective='reg:squarederror', tree_method='hist', eval_metric='mse', # 使用XGBoost内置MSE early_stopping_rounds=50) return model
方法2:自定义适配多输出的MSE函数
如果必须使用sklearn的mean_squared_error,可以自定义函数手动还原预测值的形状:
from sklearn.metrics import mean_squared_error import numpy as np def multioutput_mse(y_true, y_pred): # 将扁平化的预测值还原为和真实值一致的形状 y_pred_reshaped = y_pred.reshape(y_true.shape) # 返回指标名称和计算结果 return 'mse', mean_squared_error(y_true, y_pred_reshaped) # 修改模型定义中的eval_metric def xgboost(): model = xgb.XGBRegressor(n_estimators=200, max_depth=4, subsample=1, min_child_weight=1, objective='reg:squarederror', tree_method='hist', eval_metric=multioutput_mse, # 使用自定义函数 early_stopping_rounds=50) return model
内容的提问来源于stack exchange,提问作者Murilo
相关产品推荐
相关产品推荐

