GradientBoostingRegressor网格搜索调参方法及ValueError问题修复求助
修复Gradient Boosting Regressor结合GridSearchCV的参数调优错误并验证代码正确性
我来帮你分析并修复这个问题,同时验证你的代码修改方向是否正确:
错误原因拆解
你遇到的ValueError: n_splits=2 cannot be greater than the number of members in each class.核心问题是用错了交叉验证方法:
StratifiedKFold是专门为分类任务设计的,它会保证每个数据折中的类别分布和原始数据一致,但你的任务是回归,标签是连续数值,根本不存在“类别”的概念,所以这个方法完全不适用。- 你还错误地对回归标签
y使用了LabelEncoder,这个工具是用来把分类任务的离散标签转成整数的,对连续的回归标签完全没必要,反而会破坏数据的连续性。 - 另外还有两个隐藏问题:你用了分类任务的评估指标
neg_log_loss,回归任务不支持这个指标;模块导入行最后多了个逗号,属于语法错误。
一步步修复代码
1. 替换交叉验证方法
把StratifiedKFold换成回归任务通用的KFold,开启shuffle=True保证数据分布随机:
kfold = KFold(n_splits=2, shuffle=True, random_state=0)
2. 移除不必要的标签编码
直接删除label_encoded_y = LabelEncoder().fit_transform(y)这行代码,fit的时候直接传入原始的连续标签y即可。
3. 更换回归任务的评估指标
把scoring="neg_log_loss"换成回归适用的指标,推荐这几个:
neg_mean_squared_error:负均方误差(GridSearchCV会最大化该值,等价于最小化MSE)r2:R²分数(衡量模型拟合优度,越接近1越好)neg_mean_absolute_error:负平均绝对误差
这里我用neg_mean_squared_error作为示例。
4. 修复语法错误
把导入模块的行from sklearn.ensemble import ExtraTreesRegressor,GradientBoostingRegressor,末尾的逗号去掉,改成:
from sklearn.ensemble import ExtraTreesRegressor, GradientBoostingRegressor
5. 补全数据拆分逻辑
你的代码里没看到X和y的提取,需要从Excel数据中拆分特征和标签:
X = Data_ini.iloc[:, :-1] # 假设最后一列是输出标签,前4列是输入特征 y = Data_ini.iloc[:, -1]
修复后的完整代码
# Gradient Boosting Regressor 参数调优(GridSearchCV) from sklearn.model_selection import GridSearchCV, KFold import matplotlib import pandas as pd from sklearn.ensemble import ExtraTreesRegressor, GradientBoostingRegressor from sklearn.feature_selection import SelectFromModel from sklearn import preprocessing from sklearn import model_selection from sklearn.tree import DecisionTreeRegressor from sklearn import ensemble from sklearn.metrics import mean_squared_error, mean_absolute_error from IPython.core.interactiveshell import InteractiveShell matplotlib.use('Agg') from matplotlib import pyplot import numpy as np # 读取数据并拆分特征与标签 Data_ini = pd.read_excel('Data - 1 output -Ra-in - Crossvalidation.xlsx') X = Data_ini.iloc[:, :-1] # 前4列作为输入特征 y = Data_ini.iloc[:, -1] # 最后一列作为回归标签(连续值) # 初始化GBR模型 model = GradientBoostingRegressor(random_state=0) # 定义待搜索的参数网格 n_estimators = [50, 100, 150, 200] max_depth = [2, 4, 6, 8] param_grid = dict(max_depth=max_depth, n_estimators=n_estimators) # 回归任务使用普通KFold交叉验证 kfold = KFold(n_splits=2, shuffle=True, random_state=0) # 配置GridSearchCV,使用回归任务的评估指标 grid_search = GridSearchCV( model, param_grid, scoring="neg_mean_squared_error", n_jobs=-1, cv=kfold, verbose=1 ) # 执行网格搜索 grid_result = grid_search.fit(X, y) # 输出最优结果 print("Best: %f using %s" % (grid_result.best_score_, grid_result.best_params_)) means = grid_result.cv_results_['mean_test_score'] stds = grid_result.cv_results_['std_test_score'] params = grid_result.cv_results_['params'] for mean, stdev, param in zip(means, stds, params): print("%f (%f) with: %r" % (mean, stdev, param)) # 可视化参数搜索结果 scores = np.array(means).reshape(len(max_depth), len(n_estimators)) for i, value in enumerate(max_depth): pyplot.plot(n_estimators, scores[i], label='depth: ' + str(value)) pyplot.legend() pyplot.xlabel('n_estimators') pyplot.ylabel('Negative Mean Squared Error') pyplot.savefig('n_estimators_vs_max_depth.png')
代码修改正确性验证
你的核心修改方向(把XGB分类的GridSearchCV改成GBR回归)是正确的,但之前的几个关键细节没处理好:
- 现在替换了交叉验证方法、评估指标,移除了多余的LabelEncoder,完全适配回归任务。
- 参数网格针对GBR的核心参数
max_depth和n_estimators搜索,调优方向合理。 - 修复了语法错误和数据拆分的缺失,代码可以正常运行。
额外小建议
因为你的样本量只有34个,除了n_splits=2,还可以尝试留一法交叉验证(LeaveOneOut()),更适合小样本场景;另外可以加入更多GBR参数(比如learning_rate、subsample)来扩大搜索范围,进一步提升模型性能。
内容的提问来源于stack exchange,提问作者SH_IQ
相关产品推荐
相关产品推荐

