如何基于列中子串匹配过滤Pandas DataFrame并保留指定列
Pandas筛选包含指定子串的行
给定如下Pandas DataFrame:
import pandas as pd d = { "item": ["a", "b", "c", "d"], "report": [ "john rode the subway through new york", "sally says she no longer wanted any fish, but", "was not submitted", "the doctor proceeded to call washington and new york", ], } df = pd.DataFrame(data=d)
术语列表:
terms = ["new york", "fish"]
需要筛选出report列包含terms中任意子串的行,保留item和report列,最终得到目标结果。
方法1:使用str.contains结合正则表达式
将术语列表用|拼接成正则匹配模式,一次性匹配任意目标子串,是最高效的实现方式:
# 拼接术语为正则匹配模式 pattern = '|'.join(terms) # 筛选符合条件的行 filtered_df = df[df['report'].str.contains(pattern)] print(filtered_df)
输出结果:
item report 0 a john rode the subway through new york 1 b sally says she no longer wanted any fish, but 3 d the doctor proceeded to call washington and new york
方法2:使用apply自定义匹配逻辑
如果需要更灵活的匹配规则(比如添加额外判断),可以用apply遍历每行检查:
def has_target_term(text): return any(term in text for term in terms) filtered_df = df[df['report'].apply(has_target_term)] print(filtered_df)
此方法结果与方法1一致,适合需要自定义匹配逻辑的场景。
注意事项
- 若术语包含正则特殊字符(如
.,*),需用re.escape()转义,避免匹配错误:import re pattern = '|'.join(re.escape(term) for term in terms) filtered_df = df[df['report'].str.contains(pattern)] str.contains默认区分大小写,如需忽略大小写,添加case=False参数:df['report'].str.contains(pattern, case=False)
内容的提问来源于stack exchange,提问作者Benjamin
相关产品推荐
相关产品推荐

