如何移除DataFrame中含指定停用词的整行数据
需求实现:移除包含停用词的DataFrame行
给定数据
停用词列表
stop_w = ["in", "&", "the", "|", "and", "is", "of", "a", "an", "as", "for", "was"]
原始DataFrame
| words | frequency |
|---|---|
| the company | 10 |
| green energy | 9 |
| founded in | 8 |
| gases for | 8 |
| electricity | 5 |
需求
移除words列中包含任意一个指定停用词的整行数据。
预期输出
| words | frequency |
|---|---|
| green energy | 9 |
| electricity | 5 |
解决方案代码
import pandas as pd # 构造原始DataFrame df = pd.DataFrame({ 'words': ["the company", "green energy", "founded in", "gases for", "electricity"], 'frequency': [10, 9, 8, 8, 5] }) stop_w = ["in", "&", "the", "|", "and", "is", "of", "a", "an", "as", "for", "was"] # 生成匹配完整停用词的正则表达式,避免部分匹配 pattern = r'\b(' + '|'.join(stop_w) + r')\b' # 筛选不包含任何停用词的行 filtered_df = df[~df['words'].str.contains(pattern, case=False)] print(filtered_df)
代码说明
- 正则中的
\b用于匹配完整停用词,防止类似"interesting"里的"in"被误判; ~符号实现取反逻辑,保留不含停用词的行;case=False支持忽略大小写匹配,不需要可直接删除该参数。
内容的提问来源于stack exchange,提问作者Kas
相关产品推荐
相关产品推荐

