Python Ridge正则化线性回归中KeyError问题排查求助
正则化代码报错排查(KeyError: None of [Index([] are in the [index])
问题背景
执行岭回归正则化时触发KeyError,数据集列名如下:
Index(['Year', 'Life_expectancy', 'Adult_Mortality', 'infant_deaths',
'Alcohol', 'percentage_expenditure', 'Hepatitis_B', 'Measles', 'BMI',
'under_five_deaths', 'Polio', 'Total_expenditure', 'Diphtheria',
'HIV_or_AIDS', 'GDP', 'Population', 'thinness_1_19_years',
'thinness_5_9_years', 'Income_composition_of_resources', 'Schooling'],
dtype='object')
运行代码:
X = StandardScaler().fit_transform(x_train[features]) y = x_train[label] X_test = StandardScaler().fit_transform(y_test[features]) y_test_1 = y_test[label] for myalpha in [0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 1, 1.5, 2]: # for myalpha in [0.01, 0.05, 0.1]: ridge_reg = Ridge(alpha=myalpha) ridge_reg.fit(X,y) ## Model's training error y_pred = ridge_reg.predict(X) residuals = y - y_pred MSE = np.mean(residuals**2) RMSE = np.sqrt(MSE) MAE = np.mean(abs(residuals)) ## Model's test set error y_test_pred = ridge_reg.predict(X_test) residuals = y_test_1 - y_test_pred MSE = np.mean(residuals**2) RMSE_test = np.sqrt(MSE) # print('MSE:',MSE) print('alpha:', myalpha,', RMSE train:', RMSE, ', RMSE test:', RMSE_test) # print('MAE:',MAE)
报错信息:
KeyError: None of [Index([] are in the [index]
错误原因
- 核心问题:
features变量未定义或为空列表,导致x_train[features]和y_test[features]无法匹配任何列,触发KeyError。 - 次要问题:测试集标准化错误,不应重新
fit_transform,会导致数据泄露;同时代码中误用y_test(标签数据集)取特征列,本身就不存在对应列。
解决步骤
定义有效特征与标签变量
假设预测目标为Life_expectancy,可按以下方式定义:label = 'Life_expectancy' # 自动生成除标签外的所有特征列 features = [col for col in x_train.columns if col != label]或手动指定特征列,确保
features为包含有效列名的非空列表。修正测试集标准化逻辑
复用训练集拟合的Scaler转换测试集,避免数据泄露:scaler = StandardScaler() X = scaler.fit_transform(x_train[features]) # 使用训练集的Scaler转换测试集,且测试集特征数据应为x_test而非y_test X_test = scaler.transform(x_test[features])修正测试集变量名
代码中y_test[features]需改为x_test[features],y_test仅存储标签数据,没有特征列。
内容的提问来源于stack exchange,提问作者Amelia Putri Damayanti
相关产品推荐
相关产品推荐

