You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas拆分数据时触发KeyError:['label']不在轴中

问题

执行数据拆分代码X = tester.drop('label', axis=1)和y = tester['label']时,触发KeyError: "['label'] not found in axis"错误。

代码片段

import pandas as pd 

tester = pd.read_table('C:\Users\42the\Downloads\PA_2\PA2_Data\PA2_Data\test.txt')

X = tester.drop('label', axis=1)

报错信息

Traceback (most recent call last):
  File "c:/Users/42the/Desktop/Project2-SVM.py", line 19, in <module>
    X = tester.drop('label', axis=1, inplace = True)
  File "C:\Users\42the\AppData\Local\Programs\Python\Python37\lib\site-packages\pandas\util\_decorators.py", line 311, in wrapper
    return func(*args, **kwargs)
  File "C:\Users\42the\AppData\Local\Programs\Python\Python37\lib\site-packages\pandas\core\frame.py", line 4913, in drop
    errors=errors,
  File "C:\Users\42the\AppData\Local\Programs\Python\Python37\lib\site-packages\pandas\core\generic.py", line 4150, in drop
    obj = obj._drop_axis(labels, axis, level=level, errors=errors)
  File "C:\Users\42the\AppData\Local\Programs\Python\Python37\lib\site-packages\pandas\core\generic.py", line 4185, in _drop_axis
    new_axis = axis.drop(labels, errors=errors)
  File "C:\Users\42the\AppData\Local\Programs\Python\Python37\lib\site-packages\pandas\core\indexes\base.py", line 6017, in drop
    raise KeyError(f"{labels[mask]} not found in axis")
KeyError: "['label'] not found in axis"

补充信息

  • 已尝试转换文件格式(txt转csv再转回txt),无效。
  • 测试文件前两个数据点(每行第一个值是标签,后续为特征):
9 -1 -1 -1 -1 -1 -0.948 -0.561 0.148 0.384 0.904 0.29 -0.782 -1 -1 -1 -1 -1 -1 -1 -1 -0.748 0.588 1 1 0.991 0.915 1 0.931 -0.476 -1 -1 -1 -1 -1 -1 -0.787 0.794 1 0.727 -0.178 -0.693 -0.786 -0.624 0.834 0.756 -0.822 -1 -1 -1 -1 -0.922 0.81 1 0.01 -0.928 -1 -1 -1 -1 -0.39 1 0.271 -1 -1 -1 -1 0.012 1 0.248 -1 -1 -1 -1 -1 -0.402 0.326 1 0.801 -0.998 -1 -1 -0.981 0.645 1 -0.687 -1 -1 -1 -1 -0.792 0.976 1 1 0.413 -0.976 -1 -1 -0.993 0.834 0.897 -0.951 -1 -1 -1 -0.831 0.14 1 1 0.302 -0.889 -1 -1 -1 -1 0.356 0.794 -0.836 -1 -0.445 0.074 0.833 1 1 0.696 -0.881 -1 -1 -1 -1 -1 -0.368 0.955 1 1 1 1 0.905 1 1 -0.262 -1 -1 -1 -1 -1 -1 -1 -0.507 0.451 0.692 0.692 -0.007 -0.237 1 0.882 -0.795 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 0.155 1 0.436 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -0.991 0.703 1 -0.025 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -0.833 0.959 1 -0.629 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -0.6 0.998 0.841 -0.932 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -0.424 1 0.732 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -0.908 0.43 0.622 -0.973 -1 -1 -1 -1 -1

6 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -0.783 -0.973 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -0.364 0.789 -0.371 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -0.467 0.963 0.609 -0.986 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -0.875 0.605 0.96 -0.351 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 0.05 1 0.096 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -0.582 0.97 0.532 -0.922 -1 -1 -1 -1 -1 -1 -0.602 0.307 0.718 0.718 -0.373 -0.998 0.723 1 -0.431 -1 -1 -1 -1 -1 -0.817 -0.136 0.808 1 1 1 0.697 -0.67 0.965 0.659 -1 -1 -1 -1 -1 -0.512 0.738 1 0.839 -0.336 -0.977 0.433 0.878 0.161 1 -0.102 -1 -1 -1 -1 -0.643 0.87 0.97 0.264 -0.971 -1 -0.399 1 0.117 0.835 0.968 -0.701 -1 -1 -1 -1 0.198 1 0.052 -1 -1 -0.291 0.876 0.79 -0.819 0.392 0.962 -0.461 -1 -1 -1 -0.948 0.82 1 -0.168 -0.475 0.28 0.968 0.88 -0.613 -1 -0.551 0.854 0.98 0.498 0.324 0.324 0.328 0.998 1 0.97 0.995 0.976 0.25 -0.642 -1 -1 -1 -0.64 0.661 0.971 1 1 1 0.95 0.774 0.774 0.302 -0.522 -1 -1 -1 -1 -1 -1 -1 -0.663 -0.606 -0.606 -0.606 -0.688 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 -1

解决方案

问题核心:pd.read_table默认将文件第一行作为表头,但你的测试文件无表头行,第一列是标签值,后续为特征,导致DataFrame的列名是0,1,2,...,不存在名为label的列。同时,数据是空格分隔,read_table默认用制表符分隔,若不指定分隔符会导致数据读入异常。

方式1:读取时指定列名

先统计特征数量(比如第一个数据点除标签外的数值个数),手动生成列名列表,将第一列命名为label:

import pandas as pd 

# 假设特征数为255,可根据实际数据调整
column_names = ['label'] + [f'feature_{i}' for i in range(255)]
# 指定分隔符为任意数量空白字符,适配连续空格的情况
tester = pd.read_table('C:\Users\42the\Downloads\PA_2\PA2_Data\PA2_Data\test.txt', 
                       names=column_names, sep='\s+')

# 正常拆分数据
X = tester.drop('label', axis=1)
y = tester['label']

方式2:按列索引直接拆分

无需依赖列名,利用索引定位标签列和特征列:

import pandas as pd 

# 指定分隔符为任意数量空白字符
tester = pd.read_table('C:\Users\42the\Downloads\PA_2\PA2_Data\PA2_Data\test.txt', sep='\s+')

# 第一列(索引0)为标签,其余列为特征
y = tester.iloc[:, 0]
X = tester.iloc[:, 1:]

内容的提问来源于stack exchange,提问作者Wesley Thompson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 20:34:54