如何在scikit-learn中为线性回归模型正确使用折叠交叉验证?
线性回归模型的10折交叉验证正确实现方式
问题背景
想要用10折交叉验证训练LinearRegression模型,尝试GridSearchCV时因缺少param_grid报错;用cross_val_score仅得到10折的分数和独立模型,无法得到一个可用的最终优化模型。
解决方案
LinearRegression本身没有可调的超参数(正则化类模型如Ridge/Lasso才有),因此无需网格搜索调参,只需通过交叉验证评估模型性能,同时基于全量数据训练最终模型。以下是两种可行实现方式:
方式一:交叉验证评估 + 全量数据训练
先通过cross_val_score完成10折交叉验证的性能评估,再用全部数据训练最终可用模型:
from sklearn.linear_model import LinearRegression from sklearn.datasets import fetch_california_housing from sklearn.pipeline import Pipeline from sklearn.preprocessing import StandardScaler from sklearn.model_selection import cross_val_score # 加载数据 california = fetch_california_housing() X = california.data Y = california.target # 构建预处理+模型的管道 pipe = Pipeline([ ('scale', StandardScaler()), ('model', LinearRegression()) ]) # 执行10折交叉验证,评估模型性能(这里用负均方误差作为评分指标) cv_scores = cross_val_score(pipe, X, Y, cv=10, scoring='neg_mean_squared_error') # 转换为常规均方误差值 mse_scores = -cv_scores print(f"10折交叉验证各折MSE: {mse_scores.round(4)}") print(f"10折平均MSE: {mse_scores.mean().round(4)}") # 用全量数据训练最终模型 pipe.fit(X, Y) # 查看模型系数 print("最终模型系数:", pipe.named_steps['model'].coef_) # 此时pipe可直接用于预测 # pipe.predict(new_X)
方式二:用GridSearchCV实现交叉验证(适配你提到的流程)
如果想沿用GridSearchCV的流程,由于没有超参数需要搜索,传入空的param_grid即可。GridSearchCV会完成10折交叉验证评估,并自动用全量数据训练最终模型:
from sklearn.linear_model import LinearRegression from sklearn.datasets import fetch_california_housing from sklearn.pipeline import Pipeline from sklearn.preprocessing import StandardScaler from sklearn.model_selection import GridSearchCV # 加载数据 california = fetch_california_housing() X = california.data Y = california.target # 构建管道 pipe = Pipeline([ ('scale', StandardScaler()), ('model', LinearRegression()) ]) # 初始化GridSearchCV,param_grid传空字典表示无超参数搜索 grid = GridSearchCV( estimator=pipe, param_grid={}, cv=10, scoring='neg_mean_squared_error' ) # 执行交叉验证与最终模型训练 grid.fit(X, Y) # 查看交叉验证结果 print(f"10折交叉验证平均MSE: {-grid.best_score_.round(4)}") # 获取最终训练好的模型 final_model = grid.best_estimator_ print("最终模型系数:", final_model.named_steps['model'].coef_) # 用最终模型预测 # final_model.predict(new_X)
关键说明
- LinearRegression没有需要调整的超参数,因此"优化模型"指的是用全量数据训练的模型,交叉验证仅用于评估模型的泛化能力
cross_val_score仅返回各折的性能分数,不会保留最终模型,需额外调用fit- GridSearchCV在无超参数搜索时,会自动用全量数据训练模型并存储在
best_estimator_中
内容的提问来源于stack exchange,提问作者Jenia Be Nice Please
相关产品推荐
相关产品推荐

