如何从Pandas DataFrame中移除停用词及指定词汇?
问题解决:移除DataFrame指定列中的停用词与自定义词汇
错误原因分析
1. 初始循环代码的问题
你写的循环里判断条件用的是if j not in b,而b只是自定义词汇列表,根本没用到合并了停用词的c,所以自然不会移除英文停用词。另外这个循环把所有符合条件的元素都放到了一个一维列表里,没保留原DataFrame的每行结构。
2. 列表推导式的错误
- 分割逻辑错了:原数据是用逗号分割短语,你写的
x.split(' , ')是按「空格+逗号+空格」分割,根本匹配不上原数据的格式,导致分割失败。 - 遍历逻辑错了:
for j in i是把每个短语拆成单个字符(比如把"ryzen"拆成r,y,z,e,n),然后判断单个字符是否在停用词列表里——停用词都是完整单词,单个字符肯定不在,但这样会把短语里的单词拆碎,最后拼接出来的就是残缺的字符,这就是你看到"ryzen"变成"rzen"的原因。
正确实现代码
首先确保你已经下载了NLTK的停用词资源,如果没下载,先运行:
import nltk nltk.download('stopwords')
然后用以下代码处理:
import pandas as pd from nltk.corpus import stopwords # 初始化数据 df = pd.DataFrame({'a' : ['ryzen cpu,ryzen 5 5600x,best,amd ryzen,sale', 'cpu,ryzen 9 7800x,available,computer for ryzen,new']}) # 加载停用词并合并自定义词汇 stop_words = set(stopwords.words('english')) custom_words = {'best', 'sale', 'new', 'available'} all_filter_words = stop_words.union(custom_words) def filter_text(text): # 按逗号分割成单个短语 phrases = text.split(',') processed_phrases = [] for phrase in phrases: # 把短语拆成单词,过滤掉需要移除的词 words = phrase.strip().split() filtered_words = [word for word in words if word.lower() not in all_filter_words] # 如果过滤后还有单词,就重新拼接成短语 if filtered_words: processed_phrases.append(' '.join(filtered_words)) # 把处理后的短语用逗号连接 return ','.join(processed_phrases) # 应用处理函数到指定列 df['a'] = df['a'].apply(filter_text) print(df)
输出结果
| a |
|---|
| ryzen cpu,ryzen 5 5600x,amd ryzen |
| cpu,ryzen 9 7800x,computer ryzen |
完全符合预期输出。
内容的提问来源于stack exchange,提问作者Popeye
相关产品推荐
相关产品推荐

