使用sklearn LassoCV遇AxisError及模型使用相关技术问询
关于sklearn LassoCV的使用疑问与错误排查
以下是我首次尝试使用sklearn的LassoCV的代码:
import numpy as np import pandas as pd import seaborn as sns from sklearn.pipeline import Pipeline from sklearn.impute import SimpleImputer from sklearn.preprocessing import MinMaxScaler, OneHotEncoder from sklearn.compose import ColumnTransformer from sklearn.linear_model import LassoCV from sklearn.model_selection import GridSearchCV, KFold numeric_features = ['AGE_2019', 'Inhabitants'] categorical_features = ['familty_type','studying','Job_42','sex','DEGREE', 'Activity_type', 'Nom de la commune', 'city_type', 'DEP', 'INSEE', 'Nom du département', 'reg', 'Nom de la région'] numeric_transformer = Pipeline(steps=[ ('imputer', SimpleImputer(strategy='median')), ('scaler', MinMaxScaler()) # 数据中心化 ]) categorical_transformer = Pipeline(steps=[ ('imputer', SimpleImputer(strategy='constant', fill_value='missing')), ('encoder', OneHotEncoder(handle_unknown='ignore')) # 为分类变量创建二元变量 ]) preprocessor = ColumnTransformer( transformers=[ ('numeric', numeric_transformer, numeric_features), ('categorical', categorical_transformer, categorical_features) ]) # 创建管道 lassocv_piped = Pipeline([ ('preprocessor', preprocessor), ('model', LassoCV()) ]) # 创建参数网格 dt_params = {'model__alphas': np.array([0.5])} cv_folds = KFold(n_splits=5, shuffle=True, random_state=0) lassocv_grid_piped = GridSearchCV(lassocv_piped, dt_params, cv=cv_folds, n_jobs=-1, scoring=['neg_mean_squared_error', 'r2'], refit='r2') # 拟合模型 lassocv_grid_piped.fit(df_X_train, df_Y_train.values.ravel()) # 获取指标与预测结果 Y_pred_lassocv = lassocv_grid_piped.predict(df_X_test) metrics_lassocv = lassocv_grid_piped.cv_results_ best_lassocv_parameters = lassocv_grid_piped.best_params_ print('基准模型的最佳测试负MSE : ', max(metrics_lassocv['mean_test_neg_mean_squared_error'])) print('基准模型的最佳测试R² : ', max(metrics_lassocv['mean_test_r2'])) print('基准模型的最佳参数 : ', best_lassocv_parameters) # 图形展示 results = pd.DataFrame(dt_params) for k in range(5): results = pd.concat([results, pd.DataFrame(lassocv_grid_piped.cv_results_['split'+str(k)+'_test_neg_mean_squared_error'])],axis=1) sns.relplot(data=results.melt('model__alphas', value_name='neg_mean_squared_error'), x='model__alphas', y='neg_mean_squared_error', kind='line')
我的疑问:
- 是否有必要像我这样在估计器外部使用cv_fold?
- 是否需要搭建GridSearchCV来测试不同的alpha值?
- 如何从模型中提取R²?
遇到的错误:
AxisError: axis -1 is out of bounds for array of dimension 0
解答:
1. 外部cv_fold的必要性
可以用,但非必须。LassoCV本身内置交叉验证逻辑,默认使用KFold(n_splits=5)。如果需要自定义交叉验证策略(比如分层抽样、固定随机种子的打乱),像你这样传入自定义cv_folds是合理的,能让交叉验证更适配你的数据特性。
2. 是否需要用GridSearchCV测试alpha值
完全不需要。LassoCV的核心功能就是自动在给定alpha范围内做交叉验证,选出最优alpha。你现在用GridSearchCV只传了一个alpha值,完全浪费了LassoCV的自动调参能力。正确做法是直接给LassoCV的alphas参数传入一组候选值(比如np.logspace(-4, 2, 100)),或者让它自动生成(默认会生成),无需嵌套GridSearchCV。
3. 提取R²的方法
- 直接用
LassoCV拟合:训练完成后,调用model.score(X_test, y_test)即可得到测试集的R²,这是sklearn模型的通用方法。 - 若坚持用GridSearchCV(虽无必要),可从
cv_results_中取mean_test_r2的最大值,或用lassocv_grid_piped.best_estimator_.score(df_X_test, df_Y_test)获取最优模型在测试集上的R²。
错误排查:AxisError: axis -1 is out of bounds for array of dimension 0
错误出在图形展示代码中,原因是:
- 你传入
dt_params的model__alphas是仅含一个元素的数组np.array([0.5]),转换为DataFrame后results只有一行一列。 - 循环中添加的
splitk_test_neg_mean_squared_error是标量(仅一个参数组合),转换为DataFrame后是一维数组,concat时轴维度不匹配。
解决方法:
- 移除多余的GridSearchCV,直接使用LassoCV:
lassocv_piped = Pipeline([ ('preprocessor', preprocessor), ('model', LassoCV(alphas=np.logspace(-4, 2, 50), cv=cv_folds, random_state=0)) ]) lassocv_piped.fit(df_X_train, df_Y_train.values.ravel())
- 绘制不同alpha的交叉验证结果,直接从LassoCV属性提取数据:
# 获取LassoCV的交叉验证结果 cv_mse = lassocv_piped.named_steps['model'].mse_path_ alphas = lassocv_piped.named_steps['model'].alphas_ # 整理成DataFrame results = pd.DataFrame(cv_mse.T, columns=[f"fold_{i}" for i in range(5)]) results['alpha'] = alphas results = results.melt(id_vars='alpha', value_name='mse') # 绘图 sns.relplot(data=results, x='alpha', y='mse', kind='line')
若非要保留GridSearchCV,需给dt_params传入多个alpha值(比如np.logspace(-4,2,10)),这样cv_results_里的每个split结果都是一维数组,concat时维度匹配,不会报错。
内容的提问来源于stack exchange,提问作者Gaspard_Boyer
相关产品推荐
相关产品推荐

