You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用numpy split函数分割DataFrame时触发KeyError:0问题求助

Magic04数据集划分时KeyError:0问题解决

问题描述

处理magic04.data数据集时,执行以下操作:

  • 定义列名列表cols并读取数据集
  • 将class列转为数值类型
  • 打乱数据后尝试用np.split按6:2:2比例划分train、valid、test子集

执行代码:

cols =["fLength","fWidth","fSize","fConc","fConcl","fAsym","fM3Long","fAlpha","fDist","class"]
df = pd.read_csv("magic04.data",names = cols)
df['class'] = (df['class']=='g').astype(int)

train, valid, test = np.split(df.sample(frac=1), [int(0.6*len(df)) , int(0.8*len(df)), ])

运行后触发KeyError:0,错误栈信息:

KeyError                                  Traceback (most recent call last)
/usr/local/lib/python3.9/dist-packages/pandas/core/indexes/base.py in get_loc(self, key, method, tolerance)
   3628             try:
-> 3629                 return self._engine.get_loc(casted_key)
   3630             except KeyError as err:

17 frames
KeyError: 0

The above exception was the direct cause of the following exception:

KeyError                                  Traceback (most recent call last)
/usr/local/lib/python3.9/dist-packages/pandas/core/indexes/base.py in get_loc(self, key, method, tolerance)
   3629                 return self._engine.get_loc(casted_key)
   3630             except KeyError as err:
-> 3631                 raise KeyError(key) from err
   3632             except TypeError:
   3633                 # If we have a listlike key, _check_indexing_error will raise

错误原因

np.split基于连续位置索引对数组分割,但df.sample(frac=1)仅打乱数据行顺序,未重置原DataFrame的行索引(索引仍为原数据集的非连续整数)。np.split尝试按位置0、1...访问非连续索引时,触发KeyError。

解决方法

方法1:打乱后重置索引

打乱数据后调用reset_index(drop=True)丢弃原索引,生成连续的0起始索引,再用np.split分割:

cols =["fLength","fWidth","fSize","fConc","fConcl","fAsym","fM3Long","fAlpha","fDist","class"]
df = pd.read_csv("magic04.data",names = cols)
df['class'] = (df['class']=='g').astype(int)

# 打乱数据并重置索引
shuffled_df = df.sample(frac=1).reset_index(drop=True)
# 按6:2:2比例分割
train, valid, test = np.split(shuffled_df, [int(0.6*len(shuffled_df)), int(0.8*len(shuffled_df))])

方法2:直接使用iloc切片

利用pandas的iloc按位置提取数据,无需依赖索引连续性:

cols =["fLength","fWidth","fSize","fConc","fConcl","fAsym","fM3Long","fAlpha","fDist","class"]
df = pd.read_csv("magic04.data",names = cols)
df['class'] = (df['class']=='g').astype(int)

shuffled_df = df.sample(frac=1)
train = shuffled_df.iloc[:int(0.6*len(df))]
valid = shuffled_df.iloc[int(0.6*len(df)):int(0.8*len(df))]
test = shuffled_df.iloc[int(0.8*len(df)):]

推荐方案:使用sklearn分层划分

若需保持数据集类别分布(避免子集类别失衡),推荐使用sklearn.model_selection.train_test_split:

from sklearn.model_selection import train_test_split

cols =["fLength","fWidth","fSize","fConc","fConcl","fAsym","fM3Long","fAlpha","fDist","class"]
df = pd.read_csv("magic04.data",names = cols)
df['class'] = (df['class']=='g').astype(int)

# 先划分80%训练集和20%临时集,再将临时集拆分为10%验证集和10%测试集
train, temp = train_test_split(df, test_size=0.4, random_state=42, stratify=df['class'])
valid, test = train_test_split(temp, test_size=0.5, random_state=42, stratify=temp['class'])

内容的提问来源于stack exchange,提问作者Guneet Singh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 13:45:32