You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于列中子串匹配过滤Pandas DataFrame并保留指定列

Pandas筛选包含指定子串的行

给定如下Pandas DataFrame:

import pandas as pd

d = {
    "item": ["a", "b", "c", "d"],
    "report": [
        "john rode the subway through new york",
        "sally says she no longer wanted any fish, but",
        "was not submitted",
        "the doctor proceeded to call washington and new york",
    ],
}
df = pd.DataFrame(data=d)

术语列表:

terms = ["new york", "fish"]

需要筛选出report列包含terms中任意子串的行,保留item和report列,最终得到目标结果。


方法1:使用str.contains结合正则表达式

将术语列表用|拼接成正则匹配模式,一次性匹配任意目标子串,是最高效的实现方式:

# 拼接术语为正则匹配模式
pattern = '|'.join(terms)
# 筛选符合条件的行
filtered_df = df[df['report'].str.contains(pattern)]
print(filtered_df)

输出结果:

item                                             report
0    a        john rode the subway through new york
1    b  sally says she no longer wanted any fish, but
3    d  the doctor proceeded to call washington and new york

方法2:使用apply自定义匹配逻辑

如果需要更灵活的匹配规则(比如添加额外判断),可以用apply遍历每行检查:

def has_target_term(text):
    return any(term in text for term in terms)

filtered_df = df[df['report'].apply(has_target_term)]
print(filtered_df)

此方法结果与方法1一致,适合需要自定义匹配逻辑的场景。


注意事项

  • 若术语包含正则特殊字符(如., *),需用re.escape()转义,避免匹配错误:
    import re
    pattern = '|'.join(re.escape(term) for term in terms)
    filtered_df = df[df['report'].str.contains(pattern)]
    
  • str.contains默认区分大小写,如需忽略大小写,添加case=False参数:df['report'].str.contains(pattern, case=False)

内容的提问来源于stack exchange,提问作者Benjamin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 02:55:39