机器学习项目ColumnTransformer报错ValueError:特征索引越界求助
机器学习预处理阶段索引越界问题排查
问题场景
在信用评分预测项目中,使用ColumnTransformer对特征做标准化和独热编码时,触发索引越界错误,无法完成数据转换。
错误代码及报错信息
数据预处理代码
data_df_cleaned.info() Data columns (total 23 columns): # Column Non-Null Count Dtype --- ------ -------------- ----- 0 id 59884 non-null int64 1 age 59884 non-null float64 2 occupation 59884 non-null object 3 annual_income 59884 non-null float64 4 monthly_inhand_salary 59884 non-null float64 5 num_bank_accounts 59884 non-null float64 6 num_credit_card 59884 non-null float64 7 interest_rate 59884 non-null float64 8 num_of_loan 59884 non-null float64 9 type_of_loan 59884 non-null object 10 delay_from_due_date 59884 non-null float64 11 num_of_delayed_payment 59884 non-null float64 12 changed_credit_limit 59884 non-null float64 13 num_credit_inquiries 59884 non-null float64 14 credit_mix 59884 non-null object 15 outstanding_debt 59884 non-null float64 16 credit_utilization_ratio 59884 non-null float64 17 credit_history_age 59884 non-null float64 18 payment_of_min_amount 59884 non-null object 19 amount_invested_monthly 59884 non-null float64 ... 21 monthly_balance 59884 non-null float64 22 credit_score 59884 non-null object dtypes: float64(16), int64(1), object(6) memory usage: 11.0+ MB from sklearn.preprocessing import LabelEncoder data_df_cleaned['occupation'] = LabelEncoder().fit_transform(data_df_cleaned['occupation']) data_df_cleaned['type_of_loan'] = LabelEncoder().fit_transform(data_df_cleaned['type_of_loan']) data_df_cleaned['credit_mix'] = LabelEncoder().fit_transform(data_df_cleaned['credit_mix']) data_df_cleaned['payment_behaviour'] = LabelEncoder().fit_transform(data_df_cleaned['payment_behaviour']) data_df_cleaned['payment_of_min_amount'] = LabelEncoder().fit_transform(data_df_cleaned['payment_of_min_amount']) data_df_cleaned.info() Data columns (total 23 columns): # Column Non-Null Count Dtype --- ------ -------------- ----- 0 id 59884 non-null int64 1 age 59884 non-null float64 2 occupation 59884 non-null int32 3 annual_income 59884 non-null float64 4 monthly_inhand_salary 59884 non-null float64 5 num_bank_accounts 59884 non-null float64 6 num_credit_card 59884 non-null float64 7 interest_rate 59884 non-null float64 8 num_of_loan 59884 non-null float64 9 type_of_loan 59884 non-null int32 10 delay_from_due_date 59884 non-null float64 11 num_of_delayed_payment 59884 non-null float64 12 changed_credit_limit 59884 non-null float64 13 num_credit_inquiries 59884 non-null float64 14 credit_mix 59884 non-null int32 15 outstanding_debt 59884 non-null float64 16 credit_utilization_ratio 59884 non-null float64 17 credit_history_age 59884 non-null float64 18 payment_of_min_amount 59884 non-null int32 19 amount_invested_monthly 59884 non-null float64 ... 21 monthly_balance 59884 non-null float64 22 credit_score 59884 non-null object dtypes: float64(16), int32(5), int64(1), object(1) memory usage: 9.8+ MB from sklearn.compose import ColumnTransformer from sklearn.preprocessing import StandardScaler, OneHotEncoder pipeline=ColumnTransformer([ ('num',StandardScaler(),[1,3,4,5,6,7,8,10,11,12,13,15,16,17,19,21]), # Encoding numerical variables ('cat',OneHotEncoder(handle_unknown='ignore'),[2,9,14,18,20]), # Encoding categorical variables ]) # Now that the pipeline is set and ready we can use it to transform our data X = data_df_cleaned.iloc[:, [1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21]].values #independent variables y = data_df_cleaned.iloc[:, 22].values #dependent variables from sklearn.model_selection import train_test_split X_train, X_test, y_train, y_test = train_test_split(X, y, test_size = 0.25, random_state = 0) X_train = pipeline.fit_transform(X_train) X_test = pipeline.transform(X_test)
报错信息
IndexError: index 21 is out of bounds for axis 0 with size 21
后续触发的连锁错误:
ValueError Traceback (most recent call last) <ipython-input-388-64cb46e4a9e0> in <module> ----> 1 X_train = pipeline.fit_transform(X_train) 2 X_test = pipeline.transform(X_test) ... --> 367 raise ValueError( 368 'all features must be in [0, {}] or [-{}, 0]' 369 .format(n_columns - 1, n_columns) ValueError: all features must be in [0, 20] or [-21, 0]
错误原因
核心问题是索引不匹配:
ColumnTransformer中指定的列索引(如[1,3,4,...21])是原数据集data_df_cleaned的列索引;- 但
X是从原数据集中筛选出的子集(共21列),此时X的列索引范围是0~20,不再对应原数据的索引; - 当
pipeline尝试访问X_train的索引21时,自然触发越界错误。
解决方案
推荐两种修复方式,优先选择第一种(更健壮,不易出错):
方法1:使用列名指定特征(推荐)
直接用列名字符串代替索引,避免子集筛选后的索引混乱:
# 定义数值型和分类型特征的列名 num_features = ['age', 'annual_income', 'monthly_inhand_salary', 'num_bank_accounts', 'num_credit_card', 'interest_rate', 'num_of_loan', 'delay_from_due_date', 'num_of_delayed_payment', 'changed_credit_limit', 'num_credit_inquiries', 'outstanding_debt', 'credit_utilization_ratio', 'credit_history_age', 'amount_invested_monthly', 'monthly_balance'] cat_features = ['occupation', 'type_of_loan', 'credit_mix', 'payment_of_min_amount', 'payment_behaviour'] # 构建ColumnTransformer时传入列名 pipeline = ColumnTransformer([ ('num', StandardScaler(), num_features), ('cat', OneHotEncoder(handle_unknown='ignore'), cat_features) ]) # 直接用DataFrame作为输入(不要转成values,否则列名失效) X = data_df_cleaned[num_features + cat_features] y = data_df_cleaned['credit_score'] X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=0) # 执行转换 X_train = pipeline.fit_transform(X_train) X_test = pipeline.transform(X_test)
方法2:调整为子集的列索引
如果坚持用索引,需要将ColumnTransformer中的索引替换为X子集对应的位置:
修改后的pipeline:
pipeline=ColumnTransformer([ # 对应X子集的索引:0,2,3,4,5,6,7,9,10,11,12,14,15,16,18,20 ('num',StandardScaler(),[0,2,3,4,5,6,7,9,10,11,12,14,15,16,18,20]), # 对应X子集的索引:1,8,13,17,19 ('cat',OneHotEncoder(handle_unknown='ignore'),[1,8,13,17,19]) ]) # 后续代码不变 X = data_df_cleaned.iloc[:, [1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21]].values y = data_df_cleaned.iloc[:, 22].values X_train, X_test, y_train, y_test = train_test_split(X, y, test_size = 0.25, random_state = 0) X_train = pipeline.fit_transform(X_train) X_test = pipeline.transform(X_test)
内容的提问来源于stack exchange,提问作者demetrio
相关产品推荐
相关产品推荐

