如何用Python移除仅含停用词的Pandas单元格内容
处理Pandas数据框中仅含停用词的单元格
你原来的代码会移除单元格里所有停用词,但现在要实现的是只清空那些完全由停用词组成的单元格——比如单元格内容是"the and"就删掉,要是"the apple"就保留原内容。下面是修改后的代码,每一步都加了注释,方便你理解:
import nltk from nltk.corpus import stopwords import pandas as pd # 下载英文停用词库(只需要运行一次,之后可以注释掉) nltk.download('stopwords') # 把停用词转成集合,查询速度更快 stop_words = set(stopwords.words('english')) # 读取csv文件,和你原来的操作一致 df = pd.read_csv('my-file.csv', sep=',', engine='python') # 定义一个函数,用来判断单元格是否全是停用词 def filter_stopword_only_cells(text): # 先处理空值情况,避免报错 if pd.isna(text): return text # 把文本按空格拆分成单词列表,同时过滤掉空字符串(比如多个空格产生的空元素) words = [word.strip() for word in text.split() if word.strip()] # 如果拆分后没有单词,直接返回原内容(或者空值,看你需求) if not words: return text # 检查所有单词是不是都在停用词集合里(转小写是为了匹配停用词库的格式) all_stopwords = all(word.lower() in stop_words for word in words) # 如果全是停用词,返回空字符串;否则返回原文本 return "" if all_stopwords else text # 把函数应用到目标列上 df["English term"] = df["English term"].apply(filter_stopword_only_cells)
关键说明:
- 把
stop_words转成集合set()是因为集合的查询速度比列表快很多,数据量大的时候更高效 - 函数里先处理了空值和空字符串的情况,避免运行时报错
- 用
all()函数判断所有单词是否都是停用词,只要有一个不是就保留原内容 - 最后返回空字符串,你也可以改成
pd.NA,后续方便用dropna()删除这些空行
内容的提问来源于stack exchange,提问作者Watheq Alshowaiter
相关产品推荐
相关产品推荐

