使用Stratified K-Fold时出现KeyError:索引不在列中求助
问题:Stratified K-Fold分层划分时报KeyError错误
错误原因
直接使用df_train[train_index]时,pandas会默认将传入的索引值当作列名去匹配,但这些数值并非数据集的列名,因此触发KeyError。正确的做法是使用.iloc[]方法按行位置索引提取数据行。
另外,skfold.split()的第一个参数应传入特征集x_train而非完整数据集df_train,这虽不是报错的直接原因,但更符合逻辑规范。
修正后的代码
import pandas as pd import numpy as np from sklearn.model_selection import StratifiedKFold from imblearn.over_sampling import SMOTE # 读取数据集 df_train = pd.read_csv("train_numeric_shuffled_50000_cleaned_90.csv") # 划分特征与标签 x_train = df_train.drop(['Id','Response'], axis=1) y_train = df_train['Response'] # 初始化4折分层交叉验证器 skfold = StratifiedKFold(n_splits=4) # 遍历分层划分结果 for train_index, test_index in skfold.split(x_train, y_train): # 通过.iloc按行位置索引提取数据 x_train_skf = x_train.iloc[train_index] x_test_skf = x_train.iloc[test_index] y_train_skf = y_train.iloc[train_index] y_test_skf = y_train.iloc[test_index]
补充说明
.iloc[]是pandas中基于整数位置访问行/列的方法,完美匹配StratifiedKFold.split()返回的位置索引值。- 若数据集使用自定义非连续行索引,也可使用
.loc[],但针对当前场景,.iloc[]是更准确的选择。 - 将
skfold.split()的第一个参数设为特征集x_train,符合交叉验证基于特征数据划分的逻辑,标签仅用于维持分层比例。
内容的提问来源于stack exchange,提问作者Martina Pascucci
相关产品推荐
相关产品推荐

