You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在scikit-learn中为线性回归模型正确使用折叠交叉验证?

线性回归模型的10折交叉验证正确实现方式

问题背景

想要用10折交叉验证训练LinearRegression模型,尝试GridSearchCV时因缺少param_grid报错;用cross_val_score仅得到10折的分数和独立模型,无法得到一个可用的最终优化模型。

解决方案

LinearRegression本身没有可调的超参数(正则化类模型如Ridge/Lasso才有),因此无需网格搜索调参,只需通过交叉验证评估模型性能,同时基于全量数据训练最终模型。以下是两种可行实现方式:


方式一:交叉验证评估 + 全量数据训练

先通过cross_val_score完成10折交叉验证的性能评估,再用全部数据训练最终可用模型:

from sklearn.linear_model import LinearRegression
from sklearn.datasets import fetch_california_housing
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import cross_val_score

# 加载数据
california = fetch_california_housing()
X = california.data
Y = california.target

# 构建预处理+模型的管道
pipe = Pipeline([
    ('scale', StandardScaler()),
    ('model', LinearRegression())
])

# 执行10折交叉验证,评估模型性能(这里用负均方误差作为评分指标)
cv_scores = cross_val_score(pipe, X, Y, cv=10, scoring='neg_mean_squared_error')
# 转换为常规均方误差值
mse_scores = -cv_scores
print(f"10折交叉验证各折MSE: {mse_scores.round(4)}")
print(f"10折平均MSE: {mse_scores.mean().round(4)}")

# 用全量数据训练最终模型
pipe.fit(X, Y)
# 查看模型系数
print("最终模型系数:", pipe.named_steps['model'].coef_)
# 此时pipe可直接用于预测
# pipe.predict(new_X)

方式二:用GridSearchCV实现交叉验证(适配你提到的流程)

如果想沿用GridSearchCV的流程,由于没有超参数需要搜索,传入空的param_grid即可。GridSearchCV会完成10折交叉验证评估,并自动用全量数据训练最终模型:

from sklearn.linear_model import LinearRegression
from sklearn.datasets import fetch_california_housing
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import GridSearchCV

# 加载数据
california = fetch_california_housing()
X = california.data
Y = california.target

# 构建管道
pipe = Pipeline([
    ('scale', StandardScaler()),
    ('model', LinearRegression())
])

# 初始化GridSearchCV,param_grid传空字典表示无超参数搜索
grid = GridSearchCV(
    estimator=pipe,
    param_grid={},
    cv=10,
    scoring='neg_mean_squared_error'
)

# 执行交叉验证与最终模型训练
grid.fit(X, Y)

# 查看交叉验证结果
print(f"10折交叉验证平均MSE: {-grid.best_score_.round(4)}")
# 获取最终训练好的模型
final_model = grid.best_estimator_
print("最终模型系数:", final_model.named_steps['model'].coef_)
# 用最终模型预测
# final_model.predict(new_X)

关键说明

  • LinearRegression没有需要调整的超参数,因此"优化模型"指的是用全量数据训练的模型,交叉验证仅用于评估模型的泛化能力
  • cross_val_score仅返回各折的性能分数,不会保留最终模型,需额外调用fit
  • GridSearchCV在无超参数搜索时,会自动用全量数据训练模型并存储在best_estimator_中

内容的提问来源于stack exchange,提问作者Jenia Be Nice Please

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 04:15:33