You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Stratified K-Fold时出现KeyError:索引不在列中求助

问题:Stratified K-Fold分层划分时报KeyError错误

错误原因

直接使用df_train[train_index]时,pandas会默认将传入的索引值当作列名去匹配,但这些数值并非数据集的列名,因此触发KeyError。正确的做法是使用.iloc[]方法按行位置索引提取数据行。

另外,skfold.split()的第一个参数应传入特征集x_train而非完整数据集df_train,这虽不是报错的直接原因,但更符合逻辑规范。

修正后的代码

import pandas as pd
import numpy as np
from sklearn.model_selection import StratifiedKFold
from imblearn.over_sampling import SMOTE

# 读取数据集
df_train = pd.read_csv("train_numeric_shuffled_50000_cleaned_90.csv")
# 划分特征与标签
x_train = df_train.drop(['Id','Response'], axis=1)
y_train = df_train['Response']
# 初始化4折分层交叉验证器
skfold = StratifiedKFold(n_splits=4)

# 遍历分层划分结果
for train_index, test_index in skfold.split(x_train, y_train):
    # 通过.iloc按行位置索引提取数据
    x_train_skf = x_train.iloc[train_index]
    x_test_skf = x_train.iloc[test_index]
    y_train_skf = y_train.iloc[train_index]
    y_test_skf = y_train.iloc[test_index]

补充说明

  • .iloc[]是pandas中基于整数位置访问行/列的方法,完美匹配StratifiedKFold.split()返回的位置索引值。
  • 若数据集使用自定义非连续行索引,也可使用.loc[],但针对当前场景,.iloc[]是更准确的选择。
  • 将skfold.split()的第一个参数设为特征集x_train,符合交叉验证基于特征数据划分的逻辑,标签仅用于维持分层比例。

内容的提问来源于stack exchange,提问作者Martina Pascucci

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 19:42:48