You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python垃圾邮件分类项目中DataFrame访问列出现KeyError:8如何解决

问题根因

你遇到的KeyError核心原因是数据集切分后行索引不连续:

  • 调用train_test_split切分得到的xtrain是原始数据集的随机子集,它保留了原始数据的行索引,并没有重置为从0开始的连续整数索引
  • 你在text_pre_processing函数中用for i in range(len(data['MESSAGE']))生成连续整数i,再用data['MESSAGE'][i]按索引取数的时候,就会出现子集里不存在这个索引值的情况,也就是你遇到的索引8不存在的报错。

解决建议

1. 修改预处理函数的遍历方式(最稳妥方案)

不要按位置索引取数,直接遍历MESSAGE列的元素即可,无需关心行索引:

def text_pre_processing(data):
  corpus = []
  # 直接遍历MESSAGE列的每一个元素,跳过索引匹配步骤
  for token in data['MESSAGE']:
    token = token.replace('\n',' ')
    token = token.replace('\t',' ')
    token = token.replace('©',' ')
    token = token.replace('/b',' ')
    token = re.sub('(https?:\/\/)([\w]+.)*',' ',token) # 移除url
    token = re.sub('(www.)([\w]+.)*',' ',token)
    token = re.sub(remove_tags, ' ', token)#移除html标签
    token = "".join([word for word in token if word not in punct])
    token = re.sub('([\d])*','',token) #移除数字
    token = token.lower()
    token = word_tokenize(token)
    token = " ".join([wordnet_lemmatizer.lemmatize(word) for word in token if not word in set(stopWords)])
    corpus.append(token)
  
  return corpus

2. 可选:切分后重置索引

如果你要保留原来的按索引取数的写法,可以在切分数据集后重置索引,删除原来的索引值:

xtrain,xtest,ytrain,ytest = train_test_split(X,y, test_size=0.3, random_state=42)
xtrain = xtrain.reset_index(drop=True)
xtest = xtest.reset_index(drop=True)

3. 额外适配Pipeline的输入要求

TfidfVectorizer要求输入是一维的文本序列,你当前的X = data.iloc[:,1:2]得到的是单列DataFrame,如果后续运行出现维度报错,可以修改为获取Series格式的输入:

X = data.iloc[:,1]

内容的提问来源于stack exchange,提问作者Archana Jalaja Surendran

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 04:36:03