You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

机器学习项目ColumnTransformer报错ValueError:特征索引越界求助

机器学习预处理阶段索引越界问题排查

问题场景

在信用评分预测项目中,使用ColumnTransformer对特征做标准化和独热编码时,触发索引越界错误,无法完成数据转换。

错误代码及报错信息

数据预处理代码

data_df_cleaned.info()

Data columns (total 23 columns):
 #   Column                    Non-Null Count  Dtype  
---  ------                    --------------  -----  
 0   id                        59884 non-null  int64  
 1   age                       59884 non-null  float64
 2   occupation                59884 non-null  object 
 3   annual_income             59884 non-null  float64
 4   monthly_inhand_salary     59884 non-null  float64
 5   num_bank_accounts         59884 non-null  float64
 6   num_credit_card           59884 non-null  float64
 7   interest_rate             59884 non-null  float64
 8   num_of_loan               59884 non-null  float64
 9   type_of_loan              59884 non-null  object 
 10  delay_from_due_date       59884 non-null  float64
 11  num_of_delayed_payment    59884 non-null  float64
 12  changed_credit_limit      59884 non-null  float64
 13  num_credit_inquiries      59884 non-null  float64
 14  credit_mix                59884 non-null  object 
 15  outstanding_debt          59884 non-null  float64
 16  credit_utilization_ratio  59884 non-null  float64
 17  credit_history_age        59884 non-null  float64
 18  payment_of_min_amount     59884 non-null  object 
 19  amount_invested_monthly   59884 non-null  float64
...
 21  monthly_balance           59884 non-null  float64
 22  credit_score              59884 non-null  object 
dtypes: float64(16), int64(1), object(6)
memory usage: 11.0+ MB


from sklearn.preprocessing import LabelEncoder
data_df_cleaned['occupation'] = LabelEncoder().fit_transform(data_df_cleaned['occupation'])
data_df_cleaned['type_of_loan'] = LabelEncoder().fit_transform(data_df_cleaned['type_of_loan'])
data_df_cleaned['credit_mix'] = LabelEncoder().fit_transform(data_df_cleaned['credit_mix'])
data_df_cleaned['payment_behaviour'] = LabelEncoder().fit_transform(data_df_cleaned['payment_behaviour'])
data_df_cleaned['payment_of_min_amount'] = LabelEncoder().fit_transform(data_df_cleaned['payment_of_min_amount'])

data_df_cleaned.info()

Data columns (total 23 columns):
 #   Column                    Non-Null Count  Dtype  
---  ------                    --------------  -----  
 0   id                        59884 non-null  int64  
 1   age                       59884 non-null  float64
 2   occupation                59884 non-null  int32  
 3   annual_income             59884 non-null  float64
 4   monthly_inhand_salary     59884 non-null  float64
 5   num_bank_accounts         59884 non-null  float64
 6   num_credit_card           59884 non-null  float64
 7   interest_rate             59884 non-null  float64
 8   num_of_loan               59884 non-null  float64
 9   type_of_loan              59884 non-null  int32  
 10  delay_from_due_date       59884 non-null  float64
 11  num_of_delayed_payment    59884 non-null  float64
 12  changed_credit_limit      59884 non-null  float64
 13  num_credit_inquiries      59884 non-null  float64
 14  credit_mix                59884 non-null  int32  
 15  outstanding_debt          59884 non-null  float64
 16  credit_utilization_ratio  59884 non-null  float64
 17  credit_history_age        59884 non-null  float64
 18  payment_of_min_amount     59884 non-null  int32  
 19  amount_invested_monthly   59884 non-null  float64
...
 21  monthly_balance           59884 non-null  float64
 22  credit_score              59884 non-null  object 
dtypes: float64(16), int32(5), int64(1), object(1)
memory usage: 9.8+ MB

from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import StandardScaler, OneHotEncoder

pipeline=ColumnTransformer([
    ('num',StandardScaler(),[1,3,4,5,6,7,8,10,11,12,13,15,16,17,19,21]), # Encoding numerical variables
    ('cat',OneHotEncoder(handle_unknown='ignore'),[2,9,14,18,20]), # Encoding categorical variables
])

# Now that the pipeline is set and ready we can use it to transform our data

X = data_df_cleaned.iloc[:, [1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21]].values #independent variables
y = data_df_cleaned.iloc[:, 22].values #dependent variables

from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size = 0.25, random_state = 0)

X_train = pipeline.fit_transform(X_train)
X_test = pipeline.transform(X_test)

报错信息

IndexError: index 21 is out of bounds for axis 0 with size 21

后续触发的连锁错误:

ValueError                                Traceback (most recent call last)
<ipython-input-388-64cb46e4a9e0> in <module>
----> 1 X_train = pipeline.fit_transform(X_train)
      2 X_test = pipeline.transform(X_test)
...
--> 367             raise ValueError(
    368                 'all features must be in [0, {}] or [-{}, 0]'
    369                 .format(n_columns - 1, n_columns)

ValueError: all features must be in [0, 20] or [-21, 0]

错误原因

核心问题是索引不匹配:

  • ColumnTransformer中指定的列索引(如[1,3,4,...21])是原数据集data_df_cleaned的列索引;
  • 但X是从原数据集中筛选出的子集(共21列),此时X的列索引范围是0~20,不再对应原数据的索引;
  • 当pipeline尝试访问X_train的索引21时,自然触发越界错误。

解决方案

推荐两种修复方式,优先选择第一种(更健壮,不易出错):

方法1:使用列名指定特征(推荐)

直接用列名字符串代替索引,避免子集筛选后的索引混乱:

# 定义数值型和分类型特征的列名
num_features = ['age', 'annual_income', 'monthly_inhand_salary', 'num_bank_accounts', 
                'num_credit_card', 'interest_rate', 'num_of_loan', 'delay_from_due_date', 
                'num_of_delayed_payment', 'changed_credit_limit', 'num_credit_inquiries', 
                'outstanding_debt', 'credit_utilization_ratio', 'credit_history_age', 
                'amount_invested_monthly', 'monthly_balance']
cat_features = ['occupation', 'type_of_loan', 'credit_mix', 'payment_of_min_amount', 'payment_behaviour']

# 构建ColumnTransformer时传入列名
pipeline = ColumnTransformer([
    ('num', StandardScaler(), num_features),
    ('cat', OneHotEncoder(handle_unknown='ignore'), cat_features)
])

# 直接用DataFrame作为输入(不要转成values,否则列名失效)
X = data_df_cleaned[num_features + cat_features]
y = data_df_cleaned['credit_score']

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=0)

# 执行转换
X_train = pipeline.fit_transform(X_train)
X_test = pipeline.transform(X_test)

方法2:调整为子集的列索引

如果坚持用索引,需要将ColumnTransformer中的索引替换为X子集对应的位置:
修改后的pipeline:

pipeline=ColumnTransformer([
    # 对应X子集的索引:0,2,3,4,5,6,7,9,10,11,12,14,15,16,18,20
    ('num',StandardScaler(),[0,2,3,4,5,6,7,9,10,11,12,14,15,16,18,20]),
    # 对应X子集的索引:1,8,13,17,19
    ('cat',OneHotEncoder(handle_unknown='ignore'),[1,8,13,17,19])
])

# 后续代码不变
X = data_df_cleaned.iloc[:, [1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21]].values
y = data_df_cleaned.iloc[:, 22].values

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size = 0.25, random_state = 0)

X_train = pipeline.fit_transform(X_train)
X_test = pipeline.transform(X_test)

内容的提问来源于stack exchange,提问作者demetrio

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 05:01:42