Python垃圾邮件分类项目中DataFrame访问列出现KeyError:8如何解决
问题根因
你遇到的KeyError核心原因是数据集切分后行索引不连续:
- 调用
train_test_split切分得到的xtrain是原始数据集的随机子集,它保留了原始数据的行索引,并没有重置为从0开始的连续整数索引 - 你在
text_pre_processing函数中用for i in range(len(data['MESSAGE']))生成连续整数i,再用data['MESSAGE'][i]按索引取数的时候,就会出现子集里不存在这个索引值的情况,也就是你遇到的索引8不存在的报错。
解决建议
1. 修改预处理函数的遍历方式(最稳妥方案)
不要按位置索引取数,直接遍历MESSAGE列的元素即可,无需关心行索引:
def text_pre_processing(data): corpus = [] # 直接遍历MESSAGE列的每一个元素,跳过索引匹配步骤 for token in data['MESSAGE']: token = token.replace('\n',' ') token = token.replace('\t',' ') token = token.replace('©',' ') token = token.replace('/b',' ') token = re.sub('(https?:\/\/)([\w]+.)*',' ',token) # 移除url token = re.sub('(www.)([\w]+.)*',' ',token) token = re.sub(remove_tags, ' ', token)#移除html标签 token = "".join([word for word in token if word not in punct]) token = re.sub('([\d])*','',token) #移除数字 token = token.lower() token = word_tokenize(token) token = " ".join([wordnet_lemmatizer.lemmatize(word) for word in token if not word in set(stopWords)]) corpus.append(token) return corpus
2. 可选:切分后重置索引
如果你要保留原来的按索引取数的写法,可以在切分数据集后重置索引,删除原来的索引值:
xtrain,xtest,ytrain,ytest = train_test_split(X,y, test_size=0.3, random_state=42) xtrain = xtrain.reset_index(drop=True) xtest = xtest.reset_index(drop=True)
3. 额外适配Pipeline的输入要求
TfidfVectorizer要求输入是一维的文本序列,你当前的X = data.iloc[:,1:2]得到的是单列DataFrame,如果后续运行出现维度报错,可以修改为获取Series格式的输入:
X = data.iloc[:,1]
内容的提问来源于stack exchange,提问作者Archana Jalaja Surendran
相关产品推荐
相关产品推荐

