使用numpy split函数分割DataFrame时触发KeyError:0问题求助
Magic04数据集划分时KeyError:0问题解决
问题描述
处理magic04.data数据集时,执行以下操作:
- 定义列名列表
cols并读取数据集 - 将
class列转为数值类型 - 打乱数据后尝试用
np.split按6:2:2比例划分train、valid、test子集
执行代码:
cols =["fLength","fWidth","fSize","fConc","fConcl","fAsym","fM3Long","fAlpha","fDist","class"] df = pd.read_csv("magic04.data",names = cols) df['class'] = (df['class']=='g').astype(int) train, valid, test = np.split(df.sample(frac=1), [int(0.6*len(df)) , int(0.8*len(df)), ])
运行后触发KeyError:0,错误栈信息:
KeyError Traceback (most recent call last) /usr/local/lib/python3.9/dist-packages/pandas/core/indexes/base.py in get_loc(self, key, method, tolerance) 3628 try: -> 3629 return self._engine.get_loc(casted_key) 3630 except KeyError as err: 17 frames KeyError: 0 The above exception was the direct cause of the following exception: KeyError Traceback (most recent call last) /usr/local/lib/python3.9/dist-packages/pandas/core/indexes/base.py in get_loc(self, key, method, tolerance) 3629 return self._engine.get_loc(casted_key) 3630 except KeyError as err: -> 3631 raise KeyError(key) from err 3632 except TypeError: 3633 # If we have a listlike key, _check_indexing_error will raise
错误原因
np.split基于连续位置索引对数组分割,但df.sample(frac=1)仅打乱数据行顺序,未重置原DataFrame的行索引(索引仍为原数据集的非连续整数)。np.split尝试按位置0、1...访问非连续索引时,触发KeyError。
解决方法
方法1:打乱后重置索引
打乱数据后调用reset_index(drop=True)丢弃原索引,生成连续的0起始索引,再用np.split分割:
cols =["fLength","fWidth","fSize","fConc","fConcl","fAsym","fM3Long","fAlpha","fDist","class"] df = pd.read_csv("magic04.data",names = cols) df['class'] = (df['class']=='g').astype(int) # 打乱数据并重置索引 shuffled_df = df.sample(frac=1).reset_index(drop=True) # 按6:2:2比例分割 train, valid, test = np.split(shuffled_df, [int(0.6*len(shuffled_df)), int(0.8*len(shuffled_df))])
方法2:直接使用iloc切片
利用pandas的iloc按位置提取数据,无需依赖索引连续性:
cols =["fLength","fWidth","fSize","fConc","fConcl","fAsym","fM3Long","fAlpha","fDist","class"] df = pd.read_csv("magic04.data",names = cols) df['class'] = (df['class']=='g').astype(int) shuffled_df = df.sample(frac=1) train = shuffled_df.iloc[:int(0.6*len(df))] valid = shuffled_df.iloc[int(0.6*len(df)):int(0.8*len(df))] test = shuffled_df.iloc[int(0.8*len(df)):]
推荐方案:使用sklearn分层划分
若需保持数据集类别分布(避免子集类别失衡),推荐使用sklearn.model_selection.train_test_split:
from sklearn.model_selection import train_test_split cols =["fLength","fWidth","fSize","fConc","fConcl","fAsym","fM3Long","fAlpha","fDist","class"] df = pd.read_csv("magic04.data",names = cols) df['class'] = (df['class']=='g').astype(int) # 先划分80%训练集和20%临时集,再将临时集拆分为10%验证集和10%测试集 train, temp = train_test_split(df, test_size=0.4, random_state=42, stratify=df['class']) valid, test = train_test_split(temp, test_size=0.5, random_state=42, stratify=temp['class'])
内容的提问来源于stack exchange,提问作者Guneet Singh
相关产品推荐
相关产品推荐

